OpenAI legacy model

GPT-4 Turbo

A deprecated GPT-4 generation that expanded context to 128K, added image input to the production model line, and became a common baseline for long-context and tool-connected applications before GPT-4o.

Context window
128K
tokens
Max output
4K
tokens
Input
$10.00
per 1M tokens
Output
$30.00
per 1M tokens
Shutdown
Oct 23
2026

01 / Overview

What GPT-4 Turbo Is

GPT-4 Turbo is the later, larger-context generation of GPT-4 that OpenAI positioned as a cheaper and improved successor to the original GPT-4 API model.

The Bridge Between Original GPT-4 and GPT-4o

GPT-4 Turbo occupies an important point in OpenAI's model history. It kept the GPT-4 family identity while expanding the context window to 128,000 tokens, supporting image input, and providing function calling for application integrations.

That made it substantially more practical than the original 8K-context GPT-4 model for long conversations, document-heavy prompts, visual analysis, and tool-connected software.

At the same time, GPT-4 Turbo now reflects an older API generation. Its output limit is only 4,096 tokens, token prices are high compared with later models, and it lacks newer features such as Structured Outputs and configurable reasoning controls.

OpenAI currently marks the model as deprecated and recommends moving to a newer model rather than adopting GPT-4 Turbo for new applications.

  • Model ID: gpt-4-turbo.
  • 128,000-token context window.
  • Text and image input with text output.
  • Function calling and streaming supported.
  • Deprecated with shutdown scheduled for October 23, 2026.
Model profile
Provider
OpenAI
Family
GPT-4 Turbo
Model ID
gpt-4-turbo
Knowledge cutoff
Dec 1, 2023
Input modalities
Text, Image
Output modality
Text
Lifecycle
Deprecated

02 / Why it mattered

Why GPT-4 Turbo Was a Major GPT-4 Upgrade

GPT-4 Turbo made GPT-4-style capability easier to use in production by dramatically increasing context capacity and reducing token cost relative to the original GPT-4 model.

Long Context Changed the Application Design Space

Original GPT-4 models were constrained by much smaller context windows. GPT-4 Turbo expanded that ceiling to 128K tokens, allowing applications to send substantially more conversation history, documentation, examples, retrieved material, or source text in a single request.

The model also became a practical vision-capable GPT-4 endpoint, allowing text and images to be combined in the same workflow. Together with function calling, this made GPT-4 Turbo useful for assistants that needed to inspect screenshots or documents and then interact with application-defined tools.

Those capabilities are now common in newer models, but they explain why many legacy systems still have prompts, evaluators, cost assumptions, and fallback logic designed around GPT-4 Turbo.

For current teams, the main value of the model is therefore compatibility and comparison: reproduce the old production baseline, then verify that a replacement preserves or improves the behavior that actually matters.

GPT-4 Turbo improvements
  1. 01

    128K context

    A major increase over the original GPT-4 context window, enabling much larger prompts and histories.

  2. 02

    Image input

    Text and images could be analyzed together in the same general GPT-4 Turbo workflow.

  3. 03

    Function calling

    Applications could connect model decisions to predefined software actions and APIs.

  4. 04

    Lower cost than original GPT-4

    GPT-4 Turbo reduced GPT-4-era token prices while increasing practical context capacity.

03 / Vision

GPT-4 Turbo for Image and Document Analysis

GPT-4 Turbo supports image input alongside text, making it relevant to legacy applications that analyze screenshots, photographed documents, charts, diagrams, and other visual evidence.

Vision Is a Migration Test, Not Just a Checkbox

An application that has depended on GPT-4 Turbo vision for years may have accumulated prompt wording and preprocessing rules tuned to the model's behavior.

That matters when migrating. A newer model may be stronger overall but interpret dense screenshots, charts, scanned forms, or unusual image crops differently. The replacement should therefore be tested against the same images and expected outputs rather than judged only by general benchmark claims.

Useful migration sets include both successful cases and failures: small text, ambiguous layouts, low-resolution scans, screenshots with multiple possible targets, and visual inputs that previously required retries or extra instructions.

The model's vision support does not extend to native audio or video input. GPT-4 Turbo accepts text and images and generates text.

Legacy vision workloads
  1. 01

    UI screenshots

    Inspect interface states, error messages, forms, dashboards, and visual support cases.

  2. 02

    Document images

    Extract or explain information from scanned pages, receipts, forms, and photographed documents.

  3. 03

    Charts and diagrams

    Interpret graphical information together with written instructions and surrounding context.

  4. 04

    Regression testing

    Compare the same production images against a newer model before retiring the GPT-4 Turbo dependency.

04 / Pricing

GPT-4 Turbo Pricing

GPT-4 Turbo costs $10.00 per million input tokens and $30.00 per million output tokens.

Legacy Economics Are Now a Strong Migration Signal

GPT-4 Turbo was introduced as a cheaper GPT-4 generation, but its pricing is expensive by current standards. Output tokens cost three times as much as input tokens, so long responses can quickly dominate request cost.

The 128K context window also creates the possibility of expensive input requests when applications send large documents or long histories. A context window describes what can fit; it does not mean filling the window is economical.

This makes cost one of the easiest migration benefits to quantify. Replay representative production traffic against the replacement model and compare cost per accepted result, not only nominal token prices. A newer model may also reduce retries, produce shorter answers, or eliminate application-side recovery logic.

OpenAI's current GPT-4 Turbo model page lists standard input and output pricing; it does not list a separate cached-input rate for this model.

  • $10.00 per 1M input tokens.
  • $30.00 per 1M output tokens.
  • No separate cached-input price is listed on the current model page.
  • Evaluate vision and long-context workloads with real production payloads.
Token pricing

1M tokens · USD

Input
$10.00
Output
$30.00

Example: 20K input + 2K output

Input cost
$0.2000
Output cost
$0.0600
Estimated total
$0.2600

05 / Context

128K Context With a 4K Output Limit

GPT-4 Turbo supports a 128,000-token context window but only up to 4,096 output tokens, creating a large gap between how much information it can read and how much it can return in one response.

The Output Ceiling Is the More Important Legacy Constraint

The 128K input capacity remains sufficient for many document, retrieval, and conversational workloads. The 4K output limit is much more restrictive compared with later GPT-4o and current-generation models.

Applications built around GPT-4 Turbo may compensate with continuation prompts, pagination, chunked report generation, or application-side stitching. Those patterns should be reviewed during migration because a newer model with a larger output budget may simplify the architecture.

Context should also be measured, not assumed. If a production request typically uses 15K tokens, moving to a million-token model does not automatically improve results. But if the application routinely truncates documents or history to stay under 128K, larger-context replacements may remove a real constraint.

  • Context window: 128,000 tokens.
  • Maximum output: 4,096 tokens.
  • Strong historical long-context capability.
  • Output ceiling is much smaller than later GPT generations.
Context capacity

Context window

128,000

Max output

4,096

Input + conversation contextOutput limit

Context can include system instructions, user messages, conversation history, image representations, retrieved material, tool results, and generated output.

06 / Capabilities

GPT-4 Turbo API Capabilities

GPT-4 Turbo supports the core integration features of its generation—text, image input, streaming, and function calling—but it predates several capabilities that are standard in newer OpenAI models.

Function Calling Without Modern Structured Outputs

Function calling allows GPT-4 Turbo to select application-defined functions and produce arguments for tool-connected workflows. That made the model useful for assistants that needed to move beyond free-form chat.

However, OpenAI's current model documentation lists Structured Outputs as unsupported. Fine-tuning and predicted outputs are also unsupported for GPT-4 Turbo.

That distinction matters for legacy systems. An application may rely on manual JSON validation, retries, schema repair, or defensive parsing because the model cannot enforce modern strict output schemas directly.

Migration can therefore improve not only model quality but also application architecture. A newer model with Structured Outputs and stronger tool support may allow developers to remove some error-handling code that existed specifically because of GPT-4 Turbo limitations.

Supported capabilities
  • Text

    Accept text input and generate text output.

    Supported
  • Image input

    Analyze screenshots, photos, scanned documents, diagrams, and other supported images.

    Supported
  • Streaming

    Receive generated output progressively.

    Supported
  • Function calling

    Generate calls to application-defined functions.

    Supported
  • Structured outputs

    Strict Structured Outputs are not supported by GPT-4 Turbo.

    Not listed
  • Fine-tuning

    Fine-tuning is not supported for this model.

    Not listed
  • Predicted outputs

    Predicted outputs are not supported.

    Not listed

07 / Migration

GPT-4 Turbo Is Deprecated — Plan the Migration

OpenAI has deprecated GPT-4 Turbo and schedules gpt-4-turbo and gpt-4-turbo-2024-04-09 for API shutdown on October 23, 2026.

Use GPT-4 Turbo as a Baseline, Not a New Dependency

The deprecation changes the model-selection question. For a new product, GPT-4 Turbo should not be the starting point. For an existing product, the priority is to understand what behavior must survive migration.

OpenAI's deprecation schedule lists GPT-5.6 Sol as the recommended substitute. The current GPT-4 Turbo model page also advises developers to use a newer model such as GPT-4o rather than GPT-4 Turbo.

A migration evaluation should include the cases most likely to regress: long prompts, visual inputs, function calls, JSON-like outputs, high-value instructions, and requests that approach the 4K output ceiling.

Compare quality and operational behavior together. A replacement can be better even if a few outputs differ stylistically, provided it improves correctness, reliability, latency, cost, or maintainability for the workload that matters.

Migration plan
  1. 01

    Capture the legacy baseline

    Save representative GPT-4 Turbo outputs, costs, latency, tool-call behavior, and known failure cases.

  2. 02

    Test GPT-5.6 Sol or another current candidate

    Run the same evaluation set under equivalent application conditions.

  3. 03

    Inspect architectural differences

    Check whether larger output limits, Structured Outputs, reasoning controls, or newer tool support can simplify the integration.

  4. 04

    Migrate before shutdown

    Remove the GPT-4 Turbo dependency before OpenAI disables API access on October 23, 2026.

08 / Evaluation

GPT-4 Turbo Strengths and Limitations

GPT-4 Turbo remains useful as a historical production baseline, but its high cost, small output limit, missing modern API features, and imminent shutdown make migration the dominant consideration in 2026.

Historical strengths

  • 128K context

    GPT-4 Turbo dramatically expanded the amount of source material and conversation history available to GPT-4-era applications.

  • Vision input

    Text and image input enabled screenshot, document-image, chart, and multimodal workflows.

  • Function calling

    The model supported tool-connected application patterns rather than only free-form text generation.

  • Useful legacy baseline

    Existing systems can compare a replacement against years of known GPT-4 Turbo behavior and metrics.

What to consider

  • Imminent API shutdown

    OpenAI schedules gpt-4-turbo for removal on October 23, 2026.

  • High token cost

    $10 input and $30 output per million tokens are expensive compared with newer model families.

  • 4K maximum output

    The 4,096-token output ceiling is restrictive for long reports, code generation, and verbose structured responses.

  • Missing modern API features

    Structured Outputs, fine-tuning, predicted outputs, and configurable reasoning are not available.

Migrate with evidence

Compare GPT-4 Turbo with its replacement on your real workload

Run the prompts, image inputs, function calls, long-context requests, and edge cases your legacy application depends on, then compare quality, latency, token usage, and cost before switching models.

Start Free

GPT-4 Turbo is deprecated. Use evaluation as a migration tool rather than starting a new production dependency on this model.

Common Questions

What is GPT-4 Turbo?

GPT-4 Turbo is a later GPT-4 generation that expanded the context window to 128,000 tokens, reduced pricing compared with original GPT-4, added image input to the production model line, and supports function calling and streaming.

How much does GPT-4 Turbo cost?

OpenAI lists GPT-4 Turbo at $10.00 per 1M input tokens and $30.00 per 1M output tokens.

What is the GPT-4 Turbo context window?

GPT-4 Turbo has a 128,000-token context window and a maximum output of 4,096 tokens.

What is the GPT-4 Turbo knowledge cutoff?

OpenAI lists a December 1, 2023 knowledge cutoff for GPT-4 Turbo.

Does GPT-4 Turbo support image input?

Yes. GPT-4 Turbo accepts text and image input and generates text output. Native audio and video input are not supported.

Does GPT-4 Turbo support function calling?

Yes. OpenAI lists function calling as a supported GPT-4 Turbo feature.

Does GPT-4 Turbo support Structured Outputs?

No. OpenAI's current GPT-4 Turbo documentation lists Structured Outputs as unsupported.

Can GPT-4 Turbo be fine-tuned?

No. Fine-tuning is not supported for GPT-4 Turbo according to the current OpenAI model documentation.

Is GPT-4 Turbo deprecated?

Yes. OpenAI has deprecated GPT-4 Turbo and schedules API shutdown for October 23, 2026.

What should replace GPT-4 Turbo?

OpenAI's current deprecation schedule lists GPT-5.6 Sol as the recommended replacement. Existing applications should compare the replacement against their real GPT-4 Turbo prompts, image inputs, function calls, output-length edge cases, latency, and cost before switching production traffic.

How is GPT-4 Turbo different from GPT-4o?

Both support 128K context and image input, but GPT-4o has a much larger output limit, lower token pricing, and newer API capabilities such as Structured Outputs and fine-tuning. GPT-4 Turbo is deprecated and should mainly be treated as a legacy compatibility and migration baseline.

Model information

Last updated

Specifications, pricing, modalities, supported features, snapshots, and lifecycle information on this page are based on the official OpenAI GPT-4 Turbo model documentation and OpenAI deprecation schedule. OpenAI lists gpt-4-turbo as deprecated with API shutdown scheduled for October 23, 2026.

GPT-4 Turbo — Pricing, 128K Context, Vision & Migration | EidoStack