OpenAI legacy model
GPT-4 Turbo
A deprecated GPT-4 generation that expanded context to 128K, added image input to the production model line, and became a common baseline for long-context and tool-connected applications before GPT-4o.
- Context window
- 128K
- tokens
- Max output
- 4K
- tokens
- Input
- $10.00
- per 1M tokens
- Output
- $30.00
- per 1M tokens
- Shutdown
- Oct 23
- 2026
01 / Overview
What GPT-4 Turbo Is
GPT-4 Turbo is the later, larger-context generation of GPT-4 that OpenAI positioned as a cheaper and improved successor to the original GPT-4 API model.
The Bridge Between Original GPT-4 and GPT-4o
GPT-4 Turbo occupies an important point in OpenAI's model history. It kept the GPT-4 family identity while expanding the context window to 128,000 tokens, supporting image input, and providing function calling for application integrations.
That made it substantially more practical than the original 8K-context GPT-4 model for long conversations, document-heavy prompts, visual analysis, and tool-connected software.
At the same time, GPT-4 Turbo now reflects an older API generation. Its output limit is only 4,096 tokens, token prices are high compared with later models, and it lacks newer features such as Structured Outputs and configurable reasoning controls.
OpenAI currently marks the model as deprecated and recommends moving to a newer model rather than adopting GPT-4 Turbo for new applications.
- Model ID:
gpt-4-turbo. - 128,000-token context window.
- Text and image input with text output.
- Function calling and streaming supported.
- Deprecated with shutdown scheduled for October 23, 2026.
- Provider
- OpenAI
- Family
- GPT-4 Turbo
- Model ID
- gpt-4-turbo
- Knowledge cutoff
- Dec 1, 2023
- Input modalities
- Text, Image
- Output modality
- Text
- Lifecycle
- Deprecated
02 / Why it mattered
Why GPT-4 Turbo Was a Major GPT-4 Upgrade
GPT-4 Turbo made GPT-4-style capability easier to use in production by dramatically increasing context capacity and reducing token cost relative to the original GPT-4 model.
Long Context Changed the Application Design Space
Original GPT-4 models were constrained by much smaller context windows. GPT-4 Turbo expanded that ceiling to 128K tokens, allowing applications to send substantially more conversation history, documentation, examples, retrieved material, or source text in a single request.
The model also became a practical vision-capable GPT-4 endpoint, allowing text and images to be combined in the same workflow. Together with function calling, this made GPT-4 Turbo useful for assistants that needed to inspect screenshots or documents and then interact with application-defined tools.
Those capabilities are now common in newer models, but they explain why many legacy systems still have prompts, evaluators, cost assumptions, and fallback logic designed around GPT-4 Turbo.
For current teams, the main value of the model is therefore compatibility and comparison: reproduce the old production baseline, then verify that a replacement preserves or improves the behavior that actually matters.
- 01
128K context
A major increase over the original GPT-4 context window, enabling much larger prompts and histories.
- 02
Image input
Text and images could be analyzed together in the same general GPT-4 Turbo workflow.
- 03
Function calling
Applications could connect model decisions to predefined software actions and APIs.
- 04
Lower cost than original GPT-4
GPT-4 Turbo reduced GPT-4-era token prices while increasing practical context capacity.
03 / Vision
GPT-4 Turbo for Image and Document Analysis
GPT-4 Turbo supports image input alongside text, making it relevant to legacy applications that analyze screenshots, photographed documents, charts, diagrams, and other visual evidence.
Vision Is a Migration Test, Not Just a Checkbox
An application that has depended on GPT-4 Turbo vision for years may have accumulated prompt wording and preprocessing rules tuned to the model's behavior.
That matters when migrating. A newer model may be stronger overall but interpret dense screenshots, charts, scanned forms, or unusual image crops differently. The replacement should therefore be tested against the same images and expected outputs rather than judged only by general benchmark claims.
Useful migration sets include both successful cases and failures: small text, ambiguous layouts, low-resolution scans, screenshots with multiple possible targets, and visual inputs that previously required retries or extra instructions.
The model's vision support does not extend to native audio or video input. GPT-4 Turbo accepts text and images and generates text.
- 01
UI screenshots
Inspect interface states, error messages, forms, dashboards, and visual support cases.
- 02
Document images
Extract or explain information from scanned pages, receipts, forms, and photographed documents.
- 03
Charts and diagrams
Interpret graphical information together with written instructions and surrounding context.
- 04
Regression testing
Compare the same production images against a newer model before retiring the GPT-4 Turbo dependency.
04 / Pricing
GPT-4 Turbo Pricing
GPT-4 Turbo costs $10.00 per million input tokens and $30.00 per million output tokens.
Legacy Economics Are Now a Strong Migration Signal
GPT-4 Turbo was introduced as a cheaper GPT-4 generation, but its pricing is expensive by current standards. Output tokens cost three times as much as input tokens, so long responses can quickly dominate request cost.
The 128K context window also creates the possibility of expensive input requests when applications send large documents or long histories. A context window describes what can fit; it does not mean filling the window is economical.
This makes cost one of the easiest migration benefits to quantify. Replay representative production traffic against the replacement model and compare cost per accepted result, not only nominal token prices. A newer model may also reduce retries, produce shorter answers, or eliminate application-side recovery logic.
OpenAI's current GPT-4 Turbo model page lists standard input and output pricing; it does not list a separate cached-input rate for this model.
- $10.00 per 1M input tokens.
- $30.00 per 1M output tokens.
- No separate cached-input price is listed on the current model page.
- Evaluate vision and long-context workloads with real production payloads.
1M tokens · USD
- Input
- $10.00
- Output
- $30.00
Example: 20K input + 2K output
- Input cost
- $0.2000
- Output cost
- $0.0600
- Estimated total
- $0.2600
05 / Context
128K Context With a 4K Output Limit
GPT-4 Turbo supports a 128,000-token context window but only up to 4,096 output tokens, creating a large gap between how much information it can read and how much it can return in one response.
The Output Ceiling Is the More Important Legacy Constraint
The 128K input capacity remains sufficient for many document, retrieval, and conversational workloads. The 4K output limit is much more restrictive compared with later GPT-4o and current-generation models.
Applications built around GPT-4 Turbo may compensate with continuation prompts, pagination, chunked report generation, or application-side stitching. Those patterns should be reviewed during migration because a newer model with a larger output budget may simplify the architecture.
Context should also be measured, not assumed. If a production request typically uses 15K tokens, moving to a million-token model does not automatically improve results. But if the application routinely truncates documents or history to stay under 128K, larger-context replacements may remove a real constraint.
- Context window: 128,000 tokens.
- Maximum output: 4,096 tokens.
- Strong historical long-context capability.
- Output ceiling is much smaller than later GPT generations.
Context window
128,000
Max output
4,096
Context can include system instructions, user messages, conversation history, image representations, retrieved material, tool results, and generated output.
06 / Capabilities
GPT-4 Turbo API Capabilities
GPT-4 Turbo supports the core integration features of its generation—text, image input, streaming, and function calling—but it predates several capabilities that are standard in newer OpenAI models.
Function Calling Without Modern Structured Outputs
Function calling allows GPT-4 Turbo to select application-defined functions and produce arguments for tool-connected workflows. That made the model useful for assistants that needed to move beyond free-form chat.
However, OpenAI's current model documentation lists Structured Outputs as unsupported. Fine-tuning and predicted outputs are also unsupported for GPT-4 Turbo.
That distinction matters for legacy systems. An application may rely on manual JSON validation, retries, schema repair, or defensive parsing because the model cannot enforce modern strict output schemas directly.
Migration can therefore improve not only model quality but also application architecture. A newer model with Structured Outputs and stronger tool support may allow developers to remove some error-handling code that existed specifically because of GPT-4 Turbo limitations.
- Supported
Text
Accept text input and generate text output.
- Supported
Image input
Analyze screenshots, photos, scanned documents, diagrams, and other supported images.
- Supported
Streaming
Receive generated output progressively.
- Supported
Function calling
Generate calls to application-defined functions.
- Not listed
Structured outputs
Strict Structured Outputs are not supported by GPT-4 Turbo.
- Not listed
Fine-tuning
Fine-tuning is not supported for this model.
- Not listed
Predicted outputs
Predicted outputs are not supported.
07 / Migration
GPT-4 Turbo Is Deprecated — Plan the Migration
OpenAI has deprecated GPT-4 Turbo and schedules gpt-4-turbo and gpt-4-turbo-2024-04-09 for API shutdown on October 23, 2026.
Use GPT-4 Turbo as a Baseline, Not a New Dependency
The deprecation changes the model-selection question. For a new product, GPT-4 Turbo should not be the starting point. For an existing product, the priority is to understand what behavior must survive migration.
OpenAI's deprecation schedule lists GPT-5.6 Sol as the recommended substitute. The current GPT-4 Turbo model page also advises developers to use a newer model such as GPT-4o rather than GPT-4 Turbo.
A migration evaluation should include the cases most likely to regress: long prompts, visual inputs, function calls, JSON-like outputs, high-value instructions, and requests that approach the 4K output ceiling.
Compare quality and operational behavior together. A replacement can be better even if a few outputs differ stylistically, provided it improves correctness, reliability, latency, cost, or maintainability for the workload that matters.
- 01
Capture the legacy baseline
Save representative GPT-4 Turbo outputs, costs, latency, tool-call behavior, and known failure cases.
- 02
Test GPT-5.6 Sol or another current candidate
Run the same evaluation set under equivalent application conditions.
- 03
Inspect architectural differences
Check whether larger output limits, Structured Outputs, reasoning controls, or newer tool support can simplify the integration.
- 04
Migrate before shutdown
Remove the GPT-4 Turbo dependency before OpenAI disables API access on October 23, 2026.
08 / Evaluation
GPT-4 Turbo Strengths and Limitations
GPT-4 Turbo remains useful as a historical production baseline, but its high cost, small output limit, missing modern API features, and imminent shutdown make migration the dominant consideration in 2026.
Historical strengths
128K context
GPT-4 Turbo dramatically expanded the amount of source material and conversation history available to GPT-4-era applications.
Vision input
Text and image input enabled screenshot, document-image, chart, and multimodal workflows.
Function calling
The model supported tool-connected application patterns rather than only free-form text generation.
Useful legacy baseline
Existing systems can compare a replacement against years of known GPT-4 Turbo behavior and metrics.
What to consider
Imminent API shutdown
OpenAI schedules gpt-4-turbo for removal on October 23, 2026.
High token cost
$10 input and $30 output per million tokens are expensive compared with newer model families.
4K maximum output
The 4,096-token output ceiling is restrictive for long reports, code generation, and verbose structured responses.
Missing modern API features
Structured Outputs, fine-tuning, predicted outputs, and configurable reasoning are not available.
Migrate with evidence
Compare GPT-4 Turbo with its replacement on your real workload
Run the prompts, image inputs, function calls, long-context requests, and edge cases your legacy application depends on, then compare quality, latency, token usage, and cost before switching models.
Start FreeGPT-4 Turbo is deprecated. Use evaluation as a migration tool rather than starting a new production dependency on this model.
Common Questions
What is GPT-4 Turbo?
GPT-4 Turbo is a later GPT-4 generation that expanded the context window to 128,000 tokens, reduced pricing compared with original GPT-4, added image input to the production model line, and supports function calling and streaming.
How much does GPT-4 Turbo cost?
OpenAI lists GPT-4 Turbo at $10.00 per 1M input tokens and $30.00 per 1M output tokens.
What is the GPT-4 Turbo context window?
GPT-4 Turbo has a 128,000-token context window and a maximum output of 4,096 tokens.
What is the GPT-4 Turbo knowledge cutoff?
OpenAI lists a December 1, 2023 knowledge cutoff for GPT-4 Turbo.
Does GPT-4 Turbo support image input?
Yes. GPT-4 Turbo accepts text and image input and generates text output. Native audio and video input are not supported.
Does GPT-4 Turbo support function calling?
Yes. OpenAI lists function calling as a supported GPT-4 Turbo feature.
Does GPT-4 Turbo support Structured Outputs?
No. OpenAI's current GPT-4 Turbo documentation lists Structured Outputs as unsupported.
Can GPT-4 Turbo be fine-tuned?
No. Fine-tuning is not supported for GPT-4 Turbo according to the current OpenAI model documentation.
Is GPT-4 Turbo deprecated?
Yes. OpenAI has deprecated GPT-4 Turbo and schedules API shutdown for October 23, 2026.
What should replace GPT-4 Turbo?
OpenAI's current deprecation schedule lists GPT-5.6 Sol as the recommended replacement. Existing applications should compare the replacement against their real GPT-4 Turbo prompts, image inputs, function calls, output-length edge cases, latency, and cost before switching production traffic.
How is GPT-4 Turbo different from GPT-4o?
Both support 128K context and image input, but GPT-4o has a much larger output limit, lower token pricing, and newer API capabilities such as Structured Outputs and fine-tuning. GPT-4 Turbo is deprecated and should mainly be treated as a legacy compatibility and migration baseline.
Model information
Last updated
Specifications, pricing, modalities, supported features, snapshots, and lifecycle information on this page are based on the official OpenAI GPT-4 Turbo model documentation and OpenAI deprecation schedule. OpenAI lists gpt-4-turbo as deprecated with API shutdown scheduled for October 23, 2026.