OpenAI legacy model
GPT-4
The original GPT-4 API model: a text-only legacy baseline with an 8K context window, high historical token cost, and years of production behavior to preserve when migrating older applications.
- Context window
- 8.2K
- tokens
- Max output
- 8.2K
- tokens
- Input
- $30.00
- per 1M tokens
- Output
- $60.00
- per 1M tokens
- Shutdown
- Oct 23
- 2026
01 / Overview
What GPT-4 Is
GPT-4 is OpenAI's original GPT-4 API model and an important historical baseline for applications built before GPT-4 Turbo, GPT-4o, and newer reasoning-capable model families.
The First GPT-4 Production Baseline
For many teams, gpt-4 was the model that moved large language models from experimentation into higher-stakes production workflows. Applications were built around its response style, prompt sensitivity, context limits, output patterns, and relatively high cost.
The current OpenAI model entry describes GPT-4 as an older high-intelligence model usable in Chat Completions. It is text-only, with an 8,192-token context window and an 8,192-token maximum output.
That profile is very different from GPT-4 Turbo and GPT-4o. GPT-4 Turbo expanded context to 128K and added image input, while GPT-4o significantly reduced cost and broadened the mature multimodal production path.
Today, the main value of GPT-4 is not that it is the best model to choose for a new application. Its value is as a reproducible legacy reference for systems that still depend on behavior established during the first GPT-4 generation.
- Original GPT-4 API family baseline.
- Text-only input and output.
- 8,192-token context window.
- Deprecated with an announced API shutdown date.
- Provider
- OpenAI
- Family
- GPT-4
- Model ID
- gpt-4
- Knowledge cutoff
- Dec 1, 2023
- Input modality
- Text
- Output modality
- Text
- Lifecycle
- Deprecated
02 / Legacy baseline
Why GPT-4 Still Matters for Legacy Applications
An older model can remain operationally important when years of prompts, tests, parsing logic, fine-tuned behavior, and user expectations were built around it.
Migration Risk Lives in Application Assumptions
Changing gpt-4 to a newer model ID is easy. Proving that the application still behaves correctly is the difficult part.
A legacy system may depend on subtle properties that are not documented in a model specification: how verbose answers are, how the model handles ambiguous instructions, which wording produces reliable classifications, how often it follows a requested format, or how a fine-tuned variant behaves on domain-specific inputs.
Those assumptions should be made explicit before migration. Collect representative prompts and known edge cases, record the current GPT-4 output, and define what counts as an acceptable replacement.
This converts an urgent deprecation project into an evaluation problem. Instead of asking whether the new model is generally better, you can ask whether it preserves or improves the exact behaviors your application needs.
- 01
Capture known behavior
Save representative GPT-4 responses, failure cases, output formats, and quality thresholds before changing the model.
- 02
Find hidden dependencies
Identify prompts, parsers, retries, post-processing, and business rules that assume GPT-4-specific behavior.
- 03
Test the replacement
Run the same workload against a current model under equivalent application conditions.
- 04
Migrate with evidence
Switch only after quality, compatibility, latency, and cost meet the acceptance criteria.
03 / Pricing
GPT-4 Pricing
GPT-4 costs $30 per million input tokens and $60 per million output tokens, making it dramatically more expensive than later GPT families.
Historical Cost Is a Strong Migration Incentive
GPT-4 pricing reflects an earlier generation of API economics. Output tokens cost twice as much as input tokens, and even moderate request volumes can become expensive compared with current alternatives.
For example, a request with 4,000 input tokens and 1,000 output tokens costs about $0.18 before accounting for retries or repeated application calls.
That can matter even more in older systems where prompts were written before aggressive cost optimization became common. Large few-shot examples, duplicated instructions, and verbose responses can multiply spend.
When evaluating a replacement, compare cost per successful task rather than token price alone. A newer model may also reduce retries, allow shorter prompts, handle structured output more reliably, or consolidate several old processing steps.
- $30.00 per 1M input tokens.
- $60.00 per 1M output tokens.
- No cached-input price is listed for the legacy GPT-4 model.
- Production cost can fall substantially when moving to a newer model that passes the same evaluation set.
1M tokens Β· USD
- Input
- $30.00
- Output
- $60.00
Example: 4K input + 1K output
- Input cost
- $0.1200
- Output cost
- $0.0600
- Estimated total
- $0.1800
04 / Context
GPT-4 Has an 8,192-Token Context Window
The original GPT-4 model supports 8,192 tokens of context, a small window by current standards and one of the clearest architectural constraints of legacy GPT-4 applications.
Old Applications Often Had to Manage Context Aggressively
With an 8K window, conversation history, system instructions, user input, examples, and generated output all compete for limited space.
That pushed early GPT-4 applications toward context-management strategies that may still exist in the codebase today: trimming old messages, summarizing history, retrieving only a few documents, splitting source material into chunks, or performing several sequential calls.
A modern replacement with a much larger window can simplify some of that logic, but migration should not automatically remove it. Context selection and retrieval can still improve quality and cost even when the replacement has far more capacity.
The key migration question is whether old context-management code remains useful or is now limiting the application unnecessarily.
- Context window: 8,192 tokens.
- Maximum output: 8,192 tokens.
- Far smaller input capacity than GPT-4 Turbo, GPT-4o, GPT-4.1, and current-generation models.
- Legacy truncation and summarization logic should be reviewed during migration.
Context window
8,192
Max output
8,192
Legacy GPT-4 applications often use explicit history trimming, summarization, retrieval, or chunking because the available context is small by modern standards.
05 / Capabilities
GPT-4 API Capabilities
The current GPT-4 model entry is intentionally limited compared with newer OpenAI models: text input and output, streaming, and fine-tuning are supported, while multimodal input and newer structured application features are not.
A Much Narrower Integration Surface Than Modern Models
GPT-4 is text-only. Image, audio, and video input are not supported by the standard model entry.
OpenAI currently lists streaming as supported. Fine-tuning is also supported, which matters for legacy deployments that invested in specialized GPT-4 behavior.
By contrast, the current model documentation lists function calling and Structured Outputs as unsupported for gpt-4. Predicted outputs are also unavailable.
This narrower capability surface is another reason migrations can become architectural upgrades rather than simple model substitutions. A replacement can potentially eliminate custom parsing, add multimodal inputs, expose modern tool workflows, or provide schema-constrained responses that were not available in the original GPT-4 path.
- Supported
Text
Text input and text output.
- Supported
Streaming
Receive generated output progressively.
- Supported
Fine-tuning
Supported for legacy GPT-4 specialization workflows.
- Not listed
Image input
The standard GPT-4 model entry is text-only.
- Not listed
Function calling
The current OpenAI GPT-4 model page lists function calling as unsupported.
- Not listed
Structured outputs
Modern schema-constrained Structured Outputs are not supported.
- Not listed
Predicted outputs
Predicted output optimization is not available.
06 / Fine-tuning
Fine-Tuned GPT-4 Workloads Need Their Own Migration Plan
Fine-tuning makes migration more complex because the production dependency is not only the GPT-4 base model but also behavior learned from a custom training dataset.
Preserve the Evaluation Dataset, Not Just the Model ID
If an application uses a fine-tuned GPT-4 model, the training examples and evaluation cases are valuable migration assets.
The first step is to establish how much the fine-tune improves over the base model on the actual workload. Then test whether a modern base model already matches that quality without customization.
If specialization is still required, reproduce the behavior using a currently supported fine-tuning path rather than assuming the old training configuration should be copied unchanged.
OpenAI's deprecation schedule also lists fine-tuned GPT-4 versions for removal on October 23, 2026, with GPT-5.6 Sol as the recommended replacement base model.
- 01
Preserve training examples
Keep the dataset and formatting that produced the legacy GPT-4 specialization.
- 02
Build a held-out evaluation
Separate representative test cases from training examples so replacement quality can be measured objectively.
- 03
Test a modern base model first
Determine whether a newer model already meets the task requirements without another fine-tune.
- 04
Specialize only when needed
Use a supported modern customization path if evaluation still shows a meaningful benefit.
07 / Migration
Migrating From GPT-4 Before October 23, 2026
OpenAI has deprecated the GPT-4 alias and schedules API access to end on October 23, 2026, with GPT-5.6 Sol listed as the recommended replacement.
Treat the Deadline as a Compatibility Project
A safe migration should reproduce the old workload rather than starting with synthetic benchmark prompts.
Collect real requests from the production distribution: short questions, long instructions, difficult edge cases, formatting-sensitive outputs, fine-tuned tasks, and prompts that historically required retries.
Run those requests against GPT-4 and the replacement while GPT-4 is still available. Compare correctness, output format, latency, token usage, cost, and any downstream validation failures.
The much larger context windows and newer capabilities of current models may allow architectural simplification, but do that after establishing parity. Separating model replacement from broader application redesign makes regressions easier to diagnose.
- 01
Freeze the GPT-4 baseline
Capture representative prompts, outputs, quality scores, failure cases, token usage, and cost before shutdown.
- 02
Test GPT-5.6 Sol
Run the same evaluation set against OpenAI's recommended replacement under equivalent application conditions.
- 03
Inspect compatibility
Compare output shape, instruction following, verbosity, latency, context behavior, and downstream parser success.
- 04
Retire legacy assumptions
After parity is established, remove obsolete context trimming, custom formatting workarounds, or other GPT-4-era constraints where appropriate.
08 / Evaluation
GPT-4 Strengths and Limitations
GPT-4's current value is primarily historical and operational: it gives legacy applications a known baseline that can be measured before migration, but its cost, context size, modality support, and API features are far behind current models.
Legacy strengths
Known production behavior
Older applications may have years of prompt tuning, acceptance tests, and operational knowledge built around GPT-4.
Useful migration baseline
The model provides a concrete reference for proving that a replacement preserves or improves existing behavior.
Fine-tuning support
Legacy customized GPT-4 deployments can be evaluated against newer base or fine-tuned alternatives.
Simple text model
For text-only legacy systems, the model's integration surface is narrow and well understood.
What to consider
Imminent API shutdown
OpenAI schedules GPT-4 removal for October 23, 2026.
Very high token cost
$30 input and $60 output per million tokens are expensive compared with newer model generations.
8K context window
The small context limit creates constraints that later GPT-4 and current-generation models largely remove.
No multimodal or modern structured features
The standard GPT-4 model is text-only and does not support Structured Outputs or the modern tool-oriented feature set.
Preserve behavior before retirement
Benchmark your GPT-4 workload before migrating
Run the prompts, long conversations, fine-tuned behaviors, output formats, and edge cases your legacy application depends on, then compare them against a current replacement model.
Start FreeGPT-4 is deprecated and scheduled for API shutdown on October 23, 2026. Use evaluation data to move legacy production traffic before the deadline.
Common Questions
What is GPT-4?
GPT-4 is OpenAI's original GPT-4 API model. It is now a deprecated text-only legacy model with an 8,192-token context window and a long history of production use.
How much does GPT-4 cost?
GPT-4 costs $30 per 1M input tokens and $60 per 1M output tokens.
What is the GPT-4 context window?
The current OpenAI model page lists an 8,192-token context window and an 8,192-token maximum output.
Does GPT-4 support images?
No. The standard gpt-4 model entry supports text input and text output. Image, audio, and video input are not supported.
Does GPT-4 support function calling?
OpenAI's current GPT-4 model documentation lists function calling as unsupported for the gpt-4 model entry.
Does GPT-4 support Structured Outputs?
No. Structured Outputs are not supported by the legacy gpt-4 model.
Can GPT-4 be fine-tuned?
Yes. OpenAI currently lists fine-tuning as supported for GPT-4, although fine-tuned GPT-4 variants are also scheduled for retirement.
Is GPT-4 deprecated?
Yes. OpenAI has deprecated gpt-4 and schedules API shutdown for October 23, 2026.
What should replace GPT-4?
OpenAI's current deprecation schedule recommends GPT-5.6 Sol as the replacement for gpt-4 and legacy GPT-4 fine-tuned workloads.
How is GPT-4 different from GPT-4 Turbo?
GPT-4 is the older text-only model with an 8K context window and $30/$60 token pricing. GPT-4 Turbo expanded context to 128K, added image input, and reduced pricing to $10/$30, making it a distinct later generation of the GPT-4 family.
How is GPT-4 different from GPT-4o?
GPT-4o is a later multimodal model with image input, a 128K context window, broader application capabilities, and substantially lower token pricing. GPT-4 is primarily relevant today as a legacy compatibility and migration baseline.
Model information
Last updated
Specifications, pricing, capabilities, fine-tuning support, and lifecycle information on this page are based on the official OpenAI GPT-4 model documentation and deprecation schedule. OpenAI currently lists gpt-4 as deprecated and recommends GPT-5.6 Sol for migration.