OpenAI legacy search model
GPT-4o Mini Search Preview
A retired low-cost search-specialized model that combined GPT-4o Mini economics with built-in web search for Chat Completions, returning web-grounded answers with source citations.
- Context window
- 128K
- tokens
- Max output
- 16.4K
- tokens
- Input
- $0.15
- per 1M tokens
- Output
- $0.60
- per 1M tokens
- Shutdown
- Jul 23
- 2026
01 / Overview
What GPT-4o Mini Search Preview Was
GPT-4o Mini Search Preview was a specialized OpenAI model trained to understand and execute web search queries through the Chat Completions API.
A Search Model, Not Just GPT-4o Mini With Another Name
The important distinction was behavioral. Standard GPT-4o Mini is a general-purpose small multimodal model. GPT-4o Mini Search Preview was optimized around web search and current-information retrieval.
In the legacy Chat Completions search path, the model searched the web before responding. That made it useful for questions where the answer depended on information outside the model's October 2023 knowledge cutoff: recent news, current facts, changing product information, travel research, market context, and other time-sensitive topics.
OpenAI released the search preview models in March 2025 as direct Chat Completions access to the same family of web-search capabilities being introduced through the Responses API.
The model is now retired. OpenAI shut down the preview search models on July 23, 2026, so this page is most useful for understanding historical integrations and planning or documenting migrations.
- Specialized for web search through Chat Completions.
- Text-only input and text output.
- 128K context window and 16,384-token maximum output.
- Shut down on July 23, 2026.
- Provider
- OpenAI
- Family
- GPT-4o Mini Search
- Model ID
- gpt-4o-mini-search-preview
- Knowledge cutoff
- Oct 1, 2023
- Input modality
- Text
- Output modality
- Text
- Lifecycle
- Shut down
02 / Search behavior
How GPT-4o Mini Search Preview Used Web Search
Unlike a normal model call that could answer entirely from model knowledge, the legacy Chat Completions search model always performed web search before producing its answer.
Search Was Part of the Model Path
This architecture made the integration simple: developers selected gpt-4o-mini-search-preview, supplied the user's request, and received a response grounded in retrieved web results.
That convenience also made the model less flexible than the modern Responses API approach. Search was not an optional tool the model could decide to use only when needed. The specialized Chat Completions model followed the search path for the request.
For workloads where every request genuinely required current web information, that behavior was straightforward. For mixed workloads—some current, some answerable without search—it could mean unnecessary search calls and less control over orchestration.
Modern Responses API web search separates the model from the search tool. That design allows applications to combine search with other tools and use newer controls instead of selecting a dedicated legacy search model.
- 01
Send a text query
The application sent a user request to the specialized Chat Completions search model.
- 02
Search the web
The search model executed web search as part of the response path.
- 03
Ground the answer
Retrieved web information was used to produce a current, source-backed response.
- 04
Return citations
The response included source annotations that applications could display as clickable citations.
03 / Citations
Web-Grounded Answers and URL Citations
A defining feature of OpenAI's search API was not merely retrieving fresh information, but returning source-linked answers that applications could expose to users.
Citations Were Part of the Product Contract
Search responses included inline citations and URL citation annotations. Those annotations exposed source metadata such as the cited URL and page title, allowing a client application to render citations alongside the generated answer.
This matters for product design. A search-backed answer should not look identical to an unsupported model response when sources are available. Users need a way to inspect where time-sensitive claims came from.
OpenAI's web search guidance explicitly requires citations to be clearly visible and clickable when web-derived information is shown to end users.
For a legacy migration, citation rendering should therefore be treated as a functional requirement rather than decorative UI. A replacement should preserve source visibility while taking advantage of current search controls.
- 01
Answer text
The model generated a natural-language response using retrieved web information.
- 02
Inline citations
Source references appeared in the response where web-backed claims were used.
- 03
URL annotations
Structured citation metadata exposed source URLs and titles to the application.
- 04
Clickable source UI
Applications were expected to make web citations visible and clickable for end users.
04 / Pricing
GPT-4o Mini Search Preview Pricing
The model's text-token pricing was $0.15 per million input tokens and $0.60 per million output tokens, with additional fees for web search tool calls.
Search Cost Was More Than Token Cost
The headline token rates matched the low-cost GPT-4o Mini tier, which made the model attractive for search-heavy products that did not need the more expensive full GPT-4o Search Preview.
However, web search introduced another cost dimension. OpenAI charged search queries in addition to model token usage, so estimating a production workflow required accounting for both the generated-token bill and search-call fees.
This is especially important when comparing old and current search architectures. A search assistant may issue many short requests where search-call frequency matters more than output-token volume.
For historical cost analysis, record the number of searches, input tokens, output tokens, and accepted answers rather than comparing only the model's token price.
- $0.15 per 1M input tokens.
- $0.60 per 1M output tokens.
- Web search queries had an additional per-tool-call fee.
- Historical launch pricing for GPT-4o Mini Search started at $25 per 1,000 search queries.
1M tokens · USD
- Input
- $0.15
- Output
- $0.60
Example: 10K input + 2K output
- Input token cost
- $0.0015
- Output token cost
- $0.0012
- Token subtotal
- $0.0027
05 / Context
A 128K Context Window for Search-Grounded Answers
GPT-4o Mini Search Preview supported a 128,000-token context window and up to 16,384 output tokens.
Search Results Compete for Context Too
In a search application, context is not only the user's message and conversation history. Retrieved web material also needs to fit into the model's working context.
The legacy search interface exposed search_context_size, allowing developers to influence how much search-result context was provided to the model before answer generation.
A larger search context can expose more evidence, but it can also increase token usage and give the model more material to reconcile. A smaller search context can reduce cost and latency but may omit useful supporting information.
Current Responses API web search provides newer controls beyond the old specialized search-model path, including features such as filters, live-access control, and returned-token budgeting.
- Context window: 128,000 tokens.
- Maximum output: 16,384 tokens.
- Search-result material consumes part of the effective context.
- Search context size was an important quality, latency, and cost variable.
Context window
128,000
Max output
16,384
Search-grounded requests combine user instructions, conversation context, retrieved web evidence, citation data, and generated output within the model workflow.
06 / Capabilities
GPT-4o Mini Search Preview Capabilities
The model was intentionally specialized: it supported text search workflows and streaming, but did not inherit the full capability set of standard GPT-4o Mini.
Specialization Came With Important Constraints
The model accepted text input and produced text output. Image, audio, and video input were not supported.
Function calling was not supported, which meant the legacy search model was a poor fit for workflows that needed to search the web and then directly choose arbitrary application-defined tools inside the same model call.
Structured Outputs were supported, giving applications a path to constrained machine-readable responses. Fine-tuning and predicted outputs were not supported.
These constraints help explain why the modern tool-based architecture is more flexible. A current Responses API model can combine web_search with broader orchestration rather than relying on a dedicated search-only model.
- Supported
Text
Text input and text output.
- Supported
Web search
Specialized to understand and execute web-search queries through the legacy search model path.
- Supported
Streaming
Supported progressive response delivery.
- Supported
Structured outputs
Supported schema-constrained output for compatible structured workflows.
- Not listed
Function calling
Application-defined function calling was not supported by this specialized search model.
- Not listed
Image input
Unlike standard GPT-4o Mini, the search preview model was text-only.
- Not listed
Fine-tuning
Fine-tuning was not supported.
07 / Migration
Migrating From GPT-4o Mini Search Preview
OpenAI shut down the GPT-4o preview search models on July 23, 2026. Existing integrations need to move to a current search architecture.
Prefer Responses API web_search for New Search Workflows
OpenAI's current migration guidance recommends the Responses API web_search tool for legacy GPT-4o search integrations.
This changes the architecture in a useful way. Instead of selecting a dedicated model that always searches, the application uses a current model with search as a hosted tool. The current search tool supports controls that the old Chat Completions preview models did not provide, including domain filters, external web-access control, and returned-token budgeting.
If an application must remain on Chat Completions, OpenAI points developers to gpt-5-search-api.
The migration should preserve more than answer text. Re-test citation rendering, current-information accuracy, search frequency, latency, query cost, and any structured-output contract the old integration relied on.
- 01
Capture legacy search cases
Preserve representative user queries, expected facts, citation requirements, latency targets, and cost measurements.
- 02
Move to Responses web_search
Use a current model with the hosted web_search tool when the application can migrate away from the legacy Chat Completions search path.
- 03
Use the Chat Completions alternative if required
Choose gpt-5-search-api when the integration must remain on Chat Completions.
- 04
Re-test source behavior
Validate citation visibility, source quality, freshness, search controls, latency, and end-to-end cost before release.
08 / Evaluation
GPT-4o Mini Search Preview Strengths and Limitations
Historically, the model offered a simple and inexpensive path to web-grounded answers. Today, its value is primarily as a reference for legacy search behavior and migration testing.
Historical strengths
Low-cost web grounding
GPT-4o Mini token pricing made search-backed answers economical compared with the full GPT-4o Search Preview.
Simple Chat Completions integration
Developers could select a dedicated search model instead of orchestrating a separate search tool.
Inline citations
Search-backed responses included source-linked citation data that applications could expose to users.
128K context
The model had substantial capacity for prompts, search evidence, conversation history, and generated answers.
What to consider
Already shut down
OpenAI removed the preview search models from service on July 23, 2026.
Search was always on
The Chat Completions search model searched before responding instead of exposing search as an optional hosted tool.
No function calling
The model could not combine its search specialization with arbitrary application-defined function calls.
Text-only modality
Image input available in standard GPT-4o Mini was not supported by GPT-4o Mini Search Preview.
Move legacy search integrations forward
Re-evaluate your search workflow with current models
Use the queries and answer-quality checks from your old GPT-4o Mini Search Preview integration as a baseline, then compare a current web_search workflow for freshness, citations, controls, latency, and cost.
Start FreeThe legacy preview model has been shut down. New search integrations should use current OpenAI web-search paths rather than depend on gpt-4o-mini-search-preview.
Common Questions
What was GPT-4o Mini Search Preview?
GPT-4o Mini Search Preview was a specialized OpenAI model for web search through the Chat Completions API. It was designed to understand search-oriented queries and generate answers grounded in web results.
Is GPT-4o Mini Search Preview still available?
No. OpenAI shut down the GPT-4o preview search models on July 23, 2026.
What should replace GPT-4o Mini Search Preview?
OpenAI recommends migrating to the Responses API web_search tool. If an application must stay on Chat Completions, OpenAI recommends gpt-5-search-api.
How much did GPT-4o Mini Search Preview cost?
Its text-token rates were $0.15 per 1M input tokens and $0.60 per 1M output tokens. Web search also carried a separate query or tool-call charge.
What was the GPT-4o Mini Search Preview context window?
The model supported a 128,000-token context window and up to 16,384 output tokens.
Did GPT-4o Mini Search Preview always search the web?
Yes. OpenAI's legacy Chat Completions search models searched before responding. Current Responses API search instead exposes web_search as a tool.
Did GPT-4o Mini Search Preview return citations?
Yes. Search-backed responses included inline citations and URL citation annotations that applications could render as visible, clickable sources.
Did GPT-4o Mini Search Preview support images?
No. The specialized search preview model supported text input and text output. Image, audio, and video input were not supported.
Did GPT-4o Mini Search Preview support function calling?
No. Function calling was not supported by this search-specialized model.
Did GPT-4o Mini Search Preview support Structured Outputs?
Yes. OpenAI lists Structured Outputs as supported, while fine-tuning and predicted outputs were not supported.
How was GPT-4o Mini Search Preview different from GPT-4o Mini?
GPT-4o Mini was a general-purpose low-cost multimodal model with image input, function calling, fine-tuning, and other application features. GPT-4o Mini Search Preview was a text-only specialized search model that automatically searched the web before answering.
Model information
Last updated
Specifications, historical pricing, search behavior, feature support, citations, and lifecycle information on this page are based on the official OpenAI GPT-4o Mini Search Preview model page, web search guide, March 2025 launch materials, and deprecation documentation.