OpenAI legacy search model

GPT-4o Search Preview

A retired GPT-4o-based search model that combined higher-capability web-grounded answering with built-in Chat Completions search and source citations.

Context window
128K
tokens
Max output
16.4K
tokens
Input
$2.50
per 1M tokens
Output
$10.00
per 1M tokens
Shutdown
Jul 23
2026

01 / Overview

What GPT-4o Search Preview Was

GPT-4o Search Preview was OpenAI's higher-capability specialized search model for the Chat Completions API, built to understand web-search queries and answer with information retrieved from the web.

Full GPT-4o Economics Applied to Search

OpenAI launched gpt-4o-search-preview in March 2025 alongside GPT-4o Mini Search Preview and the Responses API web-search tool. Both preview models exposed search directly through Chat Completions, but the full GPT-4o version used the more expensive GPT-4o pricing tier.

The model was designed for applications where search answer quality mattered more than minimizing per-token cost. Typical workloads included current-events assistants, research features, product and market research, travel planning, shopping research, and other requests that depended on information newer than the model's October 2023 knowledge cutoff.

Unlike standard GPT-4o, this search-specialized model was text-only. It did not accept image input and did not support function calling or fine-tuning.

The model is now retired. OpenAI shut down the preview search models on July 23, 2026.

  • Specialized for web search through Chat Completions.
  • Higher-cost alternative to GPT-4o Mini Search Preview.
  • Text input and text output only.
  • Retired on July 23, 2026.
Search model profile
Provider
OpenAI
Family
GPT-4o Search
Model ID
gpt-4o-search-preview
Knowledge cutoff
Oct 1, 2023
Input modality
Text
Output modality
Text
Lifecycle
Shut down

02 / Search quality

Why GPT-4o Search Preview Existed Alongside the Mini Version

The full search preview traded much higher model cost for a stronger search-answering baseline than GPT-4o Mini Search Preview.

Search Quality Was the Main Reason to Pay More

At launch, OpenAI reported 90% SimpleQA accuracy for GPT-4o Search Preview and 88% for GPT-4o Mini Search Preview. SimpleQA measures short factual question answering, so the difference did not prove that the full model would win every production search task, but it illustrated the intended positioning.

This made the full model relevant when incorrect synthesis, missed context, or weak interpretation of search results had a higher business cost than the model bill. A research assistant for analysts, for example, may tolerate higher inference cost if it improves the reliability of web-grounded answers.

The Mini version was the better economic fit when query volume dominated the decision and its answer quality met the application's threshold. The full version was the candidate to test when the same queries exposed quality gaps.

For historical evaluations, this distinction is useful: compare accepted-answer rate and citation quality, not only token prices.

Full vs Mini search
  1. 01

    Run the same queries

    Use one representative search dataset for both full and Mini search variants.

  2. 02

    Score factual accuracy

    Measure whether the answer correctly reflects the retrieved evidence and the user's question.

  3. 03

    Inspect citations

    Check whether important claims are backed by relevant and useful sources.

  4. 04

    Calculate cost per accepted answer

    A more expensive search model can still be economical if it materially reduces failures or retries.

04 / Citations

Source Citations Were Part of the Search Experience

GPT-4o Search Preview returned web-grounded answers with source information so applications could show users where current claims came from.

Search Quality Includes Source Quality

A fluent answer can still be a poor search result if the underlying sources are weak, irrelevant, or difficult for the user to inspect.

OpenAI's web-search responses include inline citations and URL citation annotations. Those annotations provide source metadata such as page URLs and titles, allowing the client interface to render clear links next to the claims they support.

OpenAI's current web-search guidance requires citations to be visible and clickable when search-derived information is shown to end users. That requirement should survive a migration from the retired preview model.

For evaluation, score citations separately from prose quality. A useful test can ask whether each important current claim is supported, whether the cited page is relevant, and whether the user can actually open the source.

Citation workflow
  1. 01

    Generate a grounded answer

    Use retrieved web evidence rather than relying only on static model knowledge.

  2. 02

    Attach URL annotations

    Associate relevant answer spans with the web pages used to support them.

  3. 03

    Render visible citations

    Expose source links clearly in the product instead of hiding provenance in raw API data.

  4. 04

    Audit source relevance

    Evaluate whether cited pages actually substantiate the important claims in the response.

05 / Pricing

GPT-4o Search Preview Pricing

GPT-4o Search Preview used $2.50 per million input tokens and $10.00 per million output tokens, with a separate charge for web-search calls.

The Full Search Variant Was Much More Expensive Than Mini

The token rates matched the standard GPT-4o tier rather than GPT-4o Mini. That made the full search model appropriate only when its quality advantage justified a much higher inference budget.

At the March 2025 search launch, OpenAI said GPT-4o search pricing started at $30 per 1,000 queries, compared with $25 per 1,000 queries for GPT-4o Mini search. Search pricing and tool-call accounting are separate from the model's text-token rates, so historical cost analysis should include both components.

For a real product, the relevant metric is cost per useful search result. A higher-priced model may still make sense if it reduces failed answers, follow-up calls, or manual verification.

  • $2.50 per 1M input tokens.
  • $10.00 per 1M output tokens.
  • Separate web-search call fees applied.
  • Launch search pricing started at $30 per 1,000 queries.
Historical pricing

1M tokens Β· USD

Input
$2.50
Output
$10.00

Example: 10K input + 2K output

Input token cost
$0.0250
Output token cost
$0.0200
Token subtotal
$0.0450

06 / Context

GPT-4o Search Preview Had a 128K Context Window

The search preview supported 128,000 tokens of total context and up to 16,384 output tokens.

Retrieved Web Evidence Uses Context Too

In a search workflow, the context budget includes more than user messages and conversation history. Retrieved web evidence also has to be passed to the model so it can synthesize a grounded answer.

The legacy search interface exposed search_context_size, which controlled how much search-result context was made available before answer generation. Increasing the search context could expose more evidence, while decreasing it could reduce latency and token usage.

This means search quality and context efficiency were linked. A difficult query might benefit from richer evidence, while a narrow factual query could be answered with less retrieved material.

Modern Responses API web search keeps search context as a separate design concern and adds newer controls unavailable to the legacy Chat Completions preview models.

  • Context window: 128,000 tokens.
  • Maximum output: 16,384 tokens.
  • Retrieved web evidence consumes part of the effective working context.
  • Search context size influenced quality, latency, and cost.
Search context capacity

Context window

128,000

Max output

16,384

Prompt + search evidenceOutput limit

Search-grounded context can include instructions, conversation history, retrieved web content, citation metadata, and the generated answer.

07 / Capabilities

GPT-4o Search Preview API Capabilities

The model was deliberately specialized for text-based web search: streaming and Structured Outputs were supported, while image input, function calling, fine-tuning, and predicted outputs were not.

Search Specialization Removed Several Standard GPT-4o Features

The gpt-4o base model accepted image input and supported function calling and fine-tuning. GPT-4o Search Preview did not inherit that full feature set.

This distinction matters when migrating a legacy application. A product that needs both web search and arbitrary application tools should not recreate the old specialized-model architecture if a current Responses API model can combine search with other tools more naturally.

Structured Outputs were supported, which allowed search answers to be constrained into compatible machine-readable structures. Streaming was also available for progressive rendering.

Legacy capabilities
  • Text

    Accepted text input and generated text output.

    Supported
  • Web search

    Specialized to understand and execute web-search queries in the Chat Completions search path.

    Supported
  • Streaming

    Supported progressive response delivery.

    Supported
  • Structured outputs

    Supported schema-constrained output for compatible workflows.

    Supported
  • Image input

    The search preview model was text-only, unlike standard GPT-4o.

    Not listed
  • Function calling

    Application-defined function calling was not supported.

    Not listed
  • Fine-tuning

    Fine-tuning was not supported for this specialized search model.

    Not listed

08 / Migration

Migrating From GPT-4o Search Preview

OpenAI shut down GPT-4o Search Preview on July 23, 2026, so any remaining production integration must use a supported search path.

Move From a Dedicated Search Model to a Search Tool

OpenAI recommends migrating gpt-4o-search-preview to the Responses API web_search tool. If an application must stay on Chat Completions, OpenAI points to gpt-5-search-api.

The Responses approach is more flexible because search is a hosted tool rather than the identity of the model itself. It supports newer controls such as domain filters, external web-access settings, and returned-token budgets. Search can also be combined with broader tool-oriented workflows.

Migration testing should preserve the behavior users actually depend on: current-information accuracy, citation visibility, source relevance, response latency, search frequency, structured-output validity, and total cost.

Do not validate migration with a handful of generic prompts. Use the real search queries and difficult cases that justified the full GPT-4o search model in the first place.

Migration path
  1. 01

    Preserve the legacy evaluation set

    Save representative search queries, expected facts, source requirements, and historical quality metrics.

  2. 02

    Move to Responses web_search

    Prefer a current model with the hosted web_search tool when the application can use the Responses API.

  3. 03

    Use the Chat Completions alternative if necessary

    Use gpt-5-search-api when the integration must remain on Chat Completions.

  4. 04

    Compare quality before cost

    Validate factual accuracy, citations, source quality, and failure rate before optimizing the economics of the replacement.

09 / Evaluation

GPT-4o Search Preview Strengths and Limitations

Historically, GPT-4o Search Preview was the higher-quality option in OpenAI's first dedicated Chat Completions search pair. Today, its main value is as a benchmark for legacy search behavior and migration quality.

Historical strengths

  • Stronger launch search baseline

    OpenAI reported a higher SimpleQA score for GPT-4o Search Preview than for the Mini search variant at launch.

  • Web-grounded answers

    The model combined search retrieval with GPT-4o-class synthesis for current-information questions.

  • Source citations

    Responses exposed web citations that applications could render as visible source links.

  • Simple Chat Completions integration

    A dedicated search model reduced the orchestration required to add web-grounded answers in the legacy API path.

What to consider

  • Already shut down

    OpenAI removed the preview search models from service on July 23, 2026.

  • Much higher model cost than Mini

    The $2.50 input / $10 output token rates were dramatically higher than GPT-4o Mini Search Preview.

  • No function calling or image input

    Search specialization removed important capabilities available in standard GPT-4o.

  • Search was always on

    The dedicated Chat Completions search path searched before responding instead of exposing web search as an optional tool.

Preserve search quality while modernizing the stack

Compare your legacy GPT-4o search workload with current web search

Reuse the queries, citation requirements, freshness checks, and answer-quality evaluations from your old GPT-4o Search Preview integration, then measure a current web_search workflow before migrating.

Start Free

The preview model is no longer available. Treat it as a historical search baseline and migrate production integrations to supported search paths.

Common Questions

What was GPT-4o Search Preview?

GPT-4o Search Preview was a specialized OpenAI model for web search through the Chat Completions API. It used GPT-4o-class search synthesis to answer text queries with current web information and source citations.

Is GPT-4o Search Preview still available?

No. OpenAI shut down GPT-4o Search Preview and GPT-4o Mini Search Preview on July 23, 2026.

What should replace GPT-4o Search Preview?

OpenAI recommends the Responses API web_search tool. If an application must stay on Chat Completions, OpenAI recommends gpt-5-search-api.

How much did GPT-4o Search Preview cost?

The model used $2.50 per 1M input tokens and $10.00 per 1M output tokens, plus a separate web-search query or tool-call fee. OpenAI's March 2025 launch materials said GPT-4o search pricing started at $30 per 1,000 queries.

What was the GPT-4o Search Preview context window?

The model supported a 128,000-token context window and up to 16,384 output tokens.

Did GPT-4o Search Preview always search the web?

Yes. OpenAI's legacy Chat Completions search models searched before responding. Current Responses API web_search instead exposes search as a tool.

Did GPT-4o Search Preview return citations?

Yes. Web-grounded responses included inline citations and URL citation annotations so applications could display visible, clickable sources.

Did GPT-4o Search Preview support image input?

No. The specialized search model accepted text input and generated text output. Image, audio, and video input were not supported.

Did GPT-4o Search Preview support function calling?

No. Function calling was not supported by this specialized search model.

Did GPT-4o Search Preview support Structured Outputs?

Yes. OpenAI lists Structured Outputs and streaming as supported, while fine-tuning and predicted outputs were not supported.

How was GPT-4o Search Preview different from GPT-4o Mini Search Preview?

Both were text-only dedicated search models with 128K context and legacy Chat Completions search behavior. GPT-4o Search Preview used much higher GPT-4o token pricing and was positioned as the stronger search-quality tier, while Mini prioritized lower cost.

Model information

Last updated

Specifications, historical pricing, search behavior, benchmark context, citations, feature support, and lifecycle information on this page are based on official OpenAI documentation for GPT-4o Search Preview, the March 2025 web-search launch, the current web-search guide, and OpenAI deprecation notices.

GPT-4o Search Preview β€” Web Search, Citations, Pricing & Migration | EidoStack