OpenAI legacy search model
GPT-4o Search Preview
A retired GPT-4o-based search model that combined higher-capability web-grounded answering with built-in Chat Completions search and source citations.
- Context window
- 128K
- tokens
- Max output
- 16.4K
- tokens
- Input
- $2.50
- per 1M tokens
- Output
- $10.00
- per 1M tokens
- Shutdown
- Jul 23
- 2026
01 / Overview
What GPT-4o Search Preview Was
GPT-4o Search Preview was OpenAI's higher-capability specialized search model for the Chat Completions API, built to understand web-search queries and answer with information retrieved from the web.
Full GPT-4o Economics Applied to Search
OpenAI launched gpt-4o-search-preview in March 2025 alongside GPT-4o Mini Search Preview and the Responses API web-search tool. Both preview models exposed search directly through Chat Completions, but the full GPT-4o version used the more expensive GPT-4o pricing tier.
The model was designed for applications where search answer quality mattered more than minimizing per-token cost. Typical workloads included current-events assistants, research features, product and market research, travel planning, shopping research, and other requests that depended on information newer than the model's October 2023 knowledge cutoff.
Unlike standard GPT-4o, this search-specialized model was text-only. It did not accept image input and did not support function calling or fine-tuning.
The model is now retired. OpenAI shut down the preview search models on July 23, 2026.
- Specialized for web search through Chat Completions.
- Higher-cost alternative to GPT-4o Mini Search Preview.
- Text input and text output only.
- Retired on July 23, 2026.
- Provider
- OpenAI
- Family
- GPT-4o Search
- Model ID
- gpt-4o-search-preview
- Knowledge cutoff
- Oct 1, 2023
- Input modality
- Text
- Output modality
- Text
- Lifecycle
- Shut down
02 / Search quality
Why GPT-4o Search Preview Existed Alongside the Mini Version
The full search preview traded much higher model cost for a stronger search-answering baseline than GPT-4o Mini Search Preview.
Search Quality Was the Main Reason to Pay More
At launch, OpenAI reported 90% SimpleQA accuracy for GPT-4o Search Preview and 88% for GPT-4o Mini Search Preview. SimpleQA measures short factual question answering, so the difference did not prove that the full model would win every production search task, but it illustrated the intended positioning.
This made the full model relevant when incorrect synthesis, missed context, or weak interpretation of search results had a higher business cost than the model bill. A research assistant for analysts, for example, may tolerate higher inference cost if it improves the reliability of web-grounded answers.
The Mini version was the better economic fit when query volume dominated the decision and its answer quality met the application's threshold. The full version was the candidate to test when the same queries exposed quality gaps.
For historical evaluations, this distinction is useful: compare accepted-answer rate and citation quality, not only token prices.
- 01
Run the same queries
Use one representative search dataset for both full and Mini search variants.
- 02
Score factual accuracy
Measure whether the answer correctly reflects the retrieved evidence and the user's question.
- 03
Inspect citations
Check whether important claims are backed by relevant and useful sources.
- 04
Calculate cost per accepted answer
A more expensive search model can still be economical if it materially reduces failures or retries.
03 / Search behavior
How GPT-4o Search Preview Used the Web
The legacy Chat Completions search model always searched before responding, making web retrieval part of the model path rather than an optional tool decision.
A Dedicated Search Model Simplified One Integration Pattern
Developers selected gpt-4o-search-preview, passed the user's text request, and received a response based on web results. The model handled the search-oriented flow without requiring the application to orchestrate a separate retrieval tool.
That simplicity came with a tradeoff. Because the model always followed the search path, applications had less control than they do with the modern Responses API web_search tool. A request that did not actually require the web still used the dedicated search model path.
The current tool architecture is more composable. Search can be exposed as one capability among others, and newer controls can govern domains, live web access, and returned-token budgets.
For legacy systems, the old architecture is worth documenting because migration changes more than a model ID. It changes how search is invoked and controlled.
- 01
Receive a text query
The application submitted the user's request to the dedicated Chat Completions search model.
- 02
Search automatically
The specialized model searched the web before generating its answer.
- 03
Synthesize evidence
Retrieved information was combined into a natural-language response.
- 04
Return source annotations
Citation metadata allowed the client to expose the sources behind web-grounded claims.
04 / Citations
Source Citations Were Part of the Search Experience
GPT-4o Search Preview returned web-grounded answers with source information so applications could show users where current claims came from.
Search Quality Includes Source Quality
A fluent answer can still be a poor search result if the underlying sources are weak, irrelevant, or difficult for the user to inspect.
OpenAI's web-search responses include inline citations and URL citation annotations. Those annotations provide source metadata such as page URLs and titles, allowing the client interface to render clear links next to the claims they support.
OpenAI's current web-search guidance requires citations to be visible and clickable when search-derived information is shown to end users. That requirement should survive a migration from the retired preview model.
For evaluation, score citations separately from prose quality. A useful test can ask whether each important current claim is supported, whether the cited page is relevant, and whether the user can actually open the source.
- 01
Generate a grounded answer
Use retrieved web evidence rather than relying only on static model knowledge.
- 02
Attach URL annotations
Associate relevant answer spans with the web pages used to support them.
- 03
Render visible citations
Expose source links clearly in the product instead of hiding provenance in raw API data.
- 04
Audit source relevance
Evaluate whether cited pages actually substantiate the important claims in the response.
05 / Pricing
GPT-4o Search Preview Pricing
GPT-4o Search Preview used $2.50 per million input tokens and $10.00 per million output tokens, with a separate charge for web-search calls.
The Full Search Variant Was Much More Expensive Than Mini
The token rates matched the standard GPT-4o tier rather than GPT-4o Mini. That made the full search model appropriate only when its quality advantage justified a much higher inference budget.
At the March 2025 search launch, OpenAI said GPT-4o search pricing started at $30 per 1,000 queries, compared with $25 per 1,000 queries for GPT-4o Mini search. Search pricing and tool-call accounting are separate from the model's text-token rates, so historical cost analysis should include both components.
For a real product, the relevant metric is cost per useful search result. A higher-priced model may still make sense if it reduces failed answers, follow-up calls, or manual verification.
- $2.50 per 1M input tokens.
- $10.00 per 1M output tokens.
- Separate web-search call fees applied.
- Launch search pricing started at $30 per 1,000 queries.
1M tokens Β· USD
- Input
- $2.50
- Output
- $10.00
Example: 10K input + 2K output
- Input token cost
- $0.0250
- Output token cost
- $0.0200
- Token subtotal
- $0.0450
06 / Context
GPT-4o Search Preview Had a 128K Context Window
The search preview supported 128,000 tokens of total context and up to 16,384 output tokens.
Retrieved Web Evidence Uses Context Too
In a search workflow, the context budget includes more than user messages and conversation history. Retrieved web evidence also has to be passed to the model so it can synthesize a grounded answer.
The legacy search interface exposed search_context_size, which controlled how much search-result context was made available before answer generation. Increasing the search context could expose more evidence, while decreasing it could reduce latency and token usage.
This means search quality and context efficiency were linked. A difficult query might benefit from richer evidence, while a narrow factual query could be answered with less retrieved material.
Modern Responses API web search keeps search context as a separate design concern and adds newer controls unavailable to the legacy Chat Completions preview models.
- Context window: 128,000 tokens.
- Maximum output: 16,384 tokens.
- Retrieved web evidence consumes part of the effective working context.
- Search context size influenced quality, latency, and cost.
Context window
128,000
Max output
16,384
Search-grounded context can include instructions, conversation history, retrieved web content, citation metadata, and the generated answer.
07 / Capabilities
GPT-4o Search Preview API Capabilities
The model was deliberately specialized for text-based web search: streaming and Structured Outputs were supported, while image input, function calling, fine-tuning, and predicted outputs were not.
Search Specialization Removed Several Standard GPT-4o Features
The gpt-4o base model accepted image input and supported function calling and fine-tuning. GPT-4o Search Preview did not inherit that full feature set.
This distinction matters when migrating a legacy application. A product that needs both web search and arbitrary application tools should not recreate the old specialized-model architecture if a current Responses API model can combine search with other tools more naturally.
Structured Outputs were supported, which allowed search answers to be constrained into compatible machine-readable structures. Streaming was also available for progressive rendering.
- Supported
Text
Accepted text input and generated text output.
- Supported
Web search
Specialized to understand and execute web-search queries in the Chat Completions search path.
- Supported
Streaming
Supported progressive response delivery.
- Supported
Structured outputs
Supported schema-constrained output for compatible workflows.
- Not listed
Image input
The search preview model was text-only, unlike standard GPT-4o.
- Not listed
Function calling
Application-defined function calling was not supported.
- Not listed
Fine-tuning
Fine-tuning was not supported for this specialized search model.
08 / Migration
Migrating From GPT-4o Search Preview
OpenAI shut down GPT-4o Search Preview on July 23, 2026, so any remaining production integration must use a supported search path.
Move From a Dedicated Search Model to a Search Tool
OpenAI recommends migrating gpt-4o-search-preview to the Responses API web_search tool. If an application must stay on Chat Completions, OpenAI points to gpt-5-search-api.
The Responses approach is more flexible because search is a hosted tool rather than the identity of the model itself. It supports newer controls such as domain filters, external web-access settings, and returned-token budgets. Search can also be combined with broader tool-oriented workflows.
Migration testing should preserve the behavior users actually depend on: current-information accuracy, citation visibility, source relevance, response latency, search frequency, structured-output validity, and total cost.
Do not validate migration with a handful of generic prompts. Use the real search queries and difficult cases that justified the full GPT-4o search model in the first place.
- 01
Preserve the legacy evaluation set
Save representative search queries, expected facts, source requirements, and historical quality metrics.
- 02
Move to Responses web_search
Prefer a current model with the hosted web_search tool when the application can use the Responses API.
- 03
Use the Chat Completions alternative if necessary
Use gpt-5-search-api when the integration must remain on Chat Completions.
- 04
Compare quality before cost
Validate factual accuracy, citations, source quality, and failure rate before optimizing the economics of the replacement.
09 / Evaluation
GPT-4o Search Preview Strengths and Limitations
Historically, GPT-4o Search Preview was the higher-quality option in OpenAI's first dedicated Chat Completions search pair. Today, its main value is as a benchmark for legacy search behavior and migration quality.
Historical strengths
Stronger launch search baseline
OpenAI reported a higher SimpleQA score for GPT-4o Search Preview than for the Mini search variant at launch.
Web-grounded answers
The model combined search retrieval with GPT-4o-class synthesis for current-information questions.
Source citations
Responses exposed web citations that applications could render as visible source links.
Simple Chat Completions integration
A dedicated search model reduced the orchestration required to add web-grounded answers in the legacy API path.
What to consider
Already shut down
OpenAI removed the preview search models from service on July 23, 2026.
Much higher model cost than Mini
The $2.50 input / $10 output token rates were dramatically higher than GPT-4o Mini Search Preview.
No function calling or image input
Search specialization removed important capabilities available in standard GPT-4o.
Search was always on
The dedicated Chat Completions search path searched before responding instead of exposing web search as an optional tool.
Preserve search quality while modernizing the stack
Compare your legacy GPT-4o search workload with current web search
Reuse the queries, citation requirements, freshness checks, and answer-quality evaluations from your old GPT-4o Search Preview integration, then measure a current web_search workflow before migrating.
Start FreeThe preview model is no longer available. Treat it as a historical search baseline and migrate production integrations to supported search paths.
Common Questions
What was GPT-4o Search Preview?
GPT-4o Search Preview was a specialized OpenAI model for web search through the Chat Completions API. It used GPT-4o-class search synthesis to answer text queries with current web information and source citations.
Is GPT-4o Search Preview still available?
No. OpenAI shut down GPT-4o Search Preview and GPT-4o Mini Search Preview on July 23, 2026.
What should replace GPT-4o Search Preview?
OpenAI recommends the Responses API web_search tool. If an application must stay on Chat Completions, OpenAI recommends gpt-5-search-api.
How much did GPT-4o Search Preview cost?
The model used $2.50 per 1M input tokens and $10.00 per 1M output tokens, plus a separate web-search query or tool-call fee. OpenAI's March 2025 launch materials said GPT-4o search pricing started at $30 per 1,000 queries.
What was the GPT-4o Search Preview context window?
The model supported a 128,000-token context window and up to 16,384 output tokens.
Did GPT-4o Search Preview always search the web?
Yes. OpenAI's legacy Chat Completions search models searched before responding. Current Responses API web_search instead exposes search as a tool.
Did GPT-4o Search Preview return citations?
Yes. Web-grounded responses included inline citations and URL citation annotations so applications could display visible, clickable sources.
Did GPT-4o Search Preview support image input?
No. The specialized search model accepted text input and generated text output. Image, audio, and video input were not supported.
Did GPT-4o Search Preview support function calling?
No. Function calling was not supported by this specialized search model.
Did GPT-4o Search Preview support Structured Outputs?
Yes. OpenAI lists Structured Outputs and streaming as supported, while fine-tuning and predicted outputs were not supported.
How was GPT-4o Search Preview different from GPT-4o Mini Search Preview?
Both were text-only dedicated search models with 128K context and legacy Chat Completions search behavior. GPT-4o Search Preview used much higher GPT-4o token pricing and was positioned as the stronger search-quality tier, while Mini prioritized lower cost.
Model information
Last updated
Specifications, historical pricing, search behavior, benchmark context, citations, feature support, and lifecycle information on this page are based on official OpenAI documentation for GPT-4o Search Preview, the March 2025 web-search launch, the current web-search guide, and OpenAI deprecation notices.