OpenAI legacy search model

GPT-4o Mini Search Preview

A retired low-cost search-specialized model that combined GPT-4o Mini economics with built-in web search for Chat Completions, returning web-grounded answers with source citations.

Context window
128K
tokens
Max output
16.4K
tokens
Input
$0.15
per 1M tokens
Output
$0.60
per 1M tokens
Shutdown
Jul 23
2026

01 / Overview

What GPT-4o Mini Search Preview Was

GPT-4o Mini Search Preview was a specialized OpenAI model trained to understand and execute web search queries through the Chat Completions API.

A Search Model, Not Just GPT-4o Mini With Another Name

The important distinction was behavioral. Standard GPT-4o Mini is a general-purpose small multimodal model. GPT-4o Mini Search Preview was optimized around web search and current-information retrieval.

In the legacy Chat Completions search path, the model searched the web before responding. That made it useful for questions where the answer depended on information outside the model's October 2023 knowledge cutoff: recent news, current facts, changing product information, travel research, market context, and other time-sensitive topics.

OpenAI released the search preview models in March 2025 as direct Chat Completions access to the same family of web-search capabilities being introduced through the Responses API.

The model is now retired. OpenAI shut down the preview search models on July 23, 2026, so this page is most useful for understanding historical integrations and planning or documenting migrations.

  • Specialized for web search through Chat Completions.
  • Text-only input and text output.
  • 128K context window and 16,384-token maximum output.
  • Shut down on July 23, 2026.
Search model profile
Provider
OpenAI
Family
GPT-4o Mini Search
Model ID
gpt-4o-mini-search-preview
Knowledge cutoff
Oct 1, 2023
Input modality
Text
Output modality
Text
Lifecycle
Shut down

03 / Citations

Web-Grounded Answers and URL Citations

A defining feature of OpenAI's search API was not merely retrieving fresh information, but returning source-linked answers that applications could expose to users.

Citations Were Part of the Product Contract

Search responses included inline citations and URL citation annotations. Those annotations exposed source metadata such as the cited URL and page title, allowing a client application to render citations alongside the generated answer.

This matters for product design. A search-backed answer should not look identical to an unsupported model response when sources are available. Users need a way to inspect where time-sensitive claims came from.

OpenAI's web search guidance explicitly requires citations to be clearly visible and clickable when web-derived information is shown to end users.

For a legacy migration, citation rendering should therefore be treated as a functional requirement rather than decorative UI. A replacement should preserve source visibility while taking advantage of current search controls.

Search response data
  1. 01

    Answer text

    The model generated a natural-language response using retrieved web information.

  2. 02

    Inline citations

    Source references appeared in the response where web-backed claims were used.

  3. 03

    URL annotations

    Structured citation metadata exposed source URLs and titles to the application.

  4. 04

    Clickable source UI

    Applications were expected to make web citations visible and clickable for end users.

04 / Pricing

GPT-4o Mini Search Preview Pricing

The model's text-token pricing was $0.15 per million input tokens and $0.60 per million output tokens, with additional fees for web search tool calls.

Search Cost Was More Than Token Cost

The headline token rates matched the low-cost GPT-4o Mini tier, which made the model attractive for search-heavy products that did not need the more expensive full GPT-4o Search Preview.

However, web search introduced another cost dimension. OpenAI charged search queries in addition to model token usage, so estimating a production workflow required accounting for both the generated-token bill and search-call fees.

This is especially important when comparing old and current search architectures. A search assistant may issue many short requests where search-call frequency matters more than output-token volume.

For historical cost analysis, record the number of searches, input tokens, output tokens, and accepted answers rather than comparing only the model's token price.

  • $0.15 per 1M input tokens.
  • $0.60 per 1M output tokens.
  • Web search queries had an additional per-tool-call fee.
  • Historical launch pricing for GPT-4o Mini Search started at $25 per 1,000 search queries.
Historical pricing

1M tokens · USD

Input
$0.15
Output
$0.60

Example: 10K input + 2K output

Input token cost
$0.0015
Output token cost
$0.0012
Token subtotal
$0.0027

05 / Context

A 128K Context Window for Search-Grounded Answers

GPT-4o Mini Search Preview supported a 128,000-token context window and up to 16,384 output tokens.

Search Results Compete for Context Too

In a search application, context is not only the user's message and conversation history. Retrieved web material also needs to fit into the model's working context.

The legacy search interface exposed search_context_size, allowing developers to influence how much search-result context was provided to the model before answer generation.

A larger search context can expose more evidence, but it can also increase token usage and give the model more material to reconcile. A smaller search context can reduce cost and latency but may omit useful supporting information.

Current Responses API web search provides newer controls beyond the old specialized search-model path, including features such as filters, live-access control, and returned-token budgeting.

  • Context window: 128,000 tokens.
  • Maximum output: 16,384 tokens.
  • Search-result material consumes part of the effective context.
  • Search context size was an important quality, latency, and cost variable.
Context capacity

Context window

128,000

Max output

16,384

Prompt + search contextOutput limit

Search-grounded requests combine user instructions, conversation context, retrieved web evidence, citation data, and generated output within the model workflow.

06 / Capabilities

GPT-4o Mini Search Preview Capabilities

The model was intentionally specialized: it supported text search workflows and streaming, but did not inherit the full capability set of standard GPT-4o Mini.

Specialization Came With Important Constraints

The model accepted text input and produced text output. Image, audio, and video input were not supported.

Function calling was not supported, which meant the legacy search model was a poor fit for workflows that needed to search the web and then directly choose arbitrary application-defined tools inside the same model call.

Structured Outputs were supported, giving applications a path to constrained machine-readable responses. Fine-tuning and predicted outputs were not supported.

These constraints help explain why the modern tool-based architecture is more flexible. A current Responses API model can combine web_search with broader orchestration rather than relying on a dedicated search-only model.

Legacy capabilities
  • Text

    Text input and text output.

    Supported
  • Web search

    Specialized to understand and execute web-search queries through the legacy search model path.

    Supported
  • Streaming

    Supported progressive response delivery.

    Supported
  • Structured outputs

    Supported schema-constrained output for compatible structured workflows.

    Supported
  • Function calling

    Application-defined function calling was not supported by this specialized search model.

    Not listed
  • Image input

    Unlike standard GPT-4o Mini, the search preview model was text-only.

    Not listed
  • Fine-tuning

    Fine-tuning was not supported.

    Not listed

07 / Migration

Migrating From GPT-4o Mini Search Preview

OpenAI shut down the GPT-4o preview search models on July 23, 2026. Existing integrations need to move to a current search architecture.

Prefer Responses API web_search for New Search Workflows

OpenAI's current migration guidance recommends the Responses API web_search tool for legacy GPT-4o search integrations.

This changes the architecture in a useful way. Instead of selecting a dedicated model that always searches, the application uses a current model with search as a hosted tool. The current search tool supports controls that the old Chat Completions preview models did not provide, including domain filters, external web-access control, and returned-token budgeting.

If an application must remain on Chat Completions, OpenAI points developers to gpt-5-search-api.

The migration should preserve more than answer text. Re-test citation rendering, current-information accuracy, search frequency, latency, query cost, and any structured-output contract the old integration relied on.

Recommended migration
  1. 01

    Capture legacy search cases

    Preserve representative user queries, expected facts, citation requirements, latency targets, and cost measurements.

  2. 02

    Move to Responses web_search

    Use a current model with the hosted web_search tool when the application can migrate away from the legacy Chat Completions search path.

  3. 03

    Use the Chat Completions alternative if required

    Choose gpt-5-search-api when the integration must remain on Chat Completions.

  4. 04

    Re-test source behavior

    Validate citation visibility, source quality, freshness, search controls, latency, and end-to-end cost before release.

08 / Evaluation

GPT-4o Mini Search Preview Strengths and Limitations

Historically, the model offered a simple and inexpensive path to web-grounded answers. Today, its value is primarily as a reference for legacy search behavior and migration testing.

Historical strengths

  • Low-cost web grounding

    GPT-4o Mini token pricing made search-backed answers economical compared with the full GPT-4o Search Preview.

  • Simple Chat Completions integration

    Developers could select a dedicated search model instead of orchestrating a separate search tool.

  • Inline citations

    Search-backed responses included source-linked citation data that applications could expose to users.

  • 128K context

    The model had substantial capacity for prompts, search evidence, conversation history, and generated answers.

What to consider

  • Already shut down

    OpenAI removed the preview search models from service on July 23, 2026.

  • Search was always on

    The Chat Completions search model searched before responding instead of exposing search as an optional hosted tool.

  • No function calling

    The model could not combine its search specialization with arbitrary application-defined function calls.

  • Text-only modality

    Image input available in standard GPT-4o Mini was not supported by GPT-4o Mini Search Preview.

Move legacy search integrations forward

Re-evaluate your search workflow with current models

Use the queries and answer-quality checks from your old GPT-4o Mini Search Preview integration as a baseline, then compare a current web_search workflow for freshness, citations, controls, latency, and cost.

Start Free

The legacy preview model has been shut down. New search integrations should use current OpenAI web-search paths rather than depend on gpt-4o-mini-search-preview.

Common Questions

What was GPT-4o Mini Search Preview?

GPT-4o Mini Search Preview was a specialized OpenAI model for web search through the Chat Completions API. It was designed to understand search-oriented queries and generate answers grounded in web results.

Is GPT-4o Mini Search Preview still available?

No. OpenAI shut down the GPT-4o preview search models on July 23, 2026.

What should replace GPT-4o Mini Search Preview?

OpenAI recommends migrating to the Responses API web_search tool. If an application must stay on Chat Completions, OpenAI recommends gpt-5-search-api.

How much did GPT-4o Mini Search Preview cost?

Its text-token rates were $0.15 per 1M input tokens and $0.60 per 1M output tokens. Web search also carried a separate query or tool-call charge.

What was the GPT-4o Mini Search Preview context window?

The model supported a 128,000-token context window and up to 16,384 output tokens.

Did GPT-4o Mini Search Preview always search the web?

Yes. OpenAI's legacy Chat Completions search models searched before responding. Current Responses API search instead exposes web_search as a tool.

Did GPT-4o Mini Search Preview return citations?

Yes. Search-backed responses included inline citations and URL citation annotations that applications could render as visible, clickable sources.

Did GPT-4o Mini Search Preview support images?

No. The specialized search preview model supported text input and text output. Image, audio, and video input were not supported.

Did GPT-4o Mini Search Preview support function calling?

No. Function calling was not supported by this search-specialized model.

Did GPT-4o Mini Search Preview support Structured Outputs?

Yes. OpenAI lists Structured Outputs as supported, while fine-tuning and predicted outputs were not supported.

How was GPT-4o Mini Search Preview different from GPT-4o Mini?

GPT-4o Mini was a general-purpose low-cost multimodal model with image input, function calling, fine-tuning, and other application features. GPT-4o Mini Search Preview was a text-only specialized search model that automatically searched the web before answering.

Model information

Last updated

Specifications, historical pricing, search behavior, feature support, citations, and lifecycle information on this page are based on the official OpenAI GPT-4o Mini Search Preview model page, web search guide, March 2025 launch materials, and deprecation documentation.

GPT-4o Mini Search Preview — Web Search, Pricing & Migration | EidoStack