OpenAI legacy model

GPT-4

The original GPT-4 API model: a text-only legacy baseline with an 8K context window, high historical token cost, and years of production behavior to preserve when migrating older applications.

Context window
8.2K
tokens
Max output
8.2K
tokens
Input
$30.00
per 1M tokens
Output
$60.00
per 1M tokens
Shutdown
Oct 23
2026

01 / Overview

What GPT-4 Is

GPT-4 is OpenAI's original GPT-4 API model and an important historical baseline for applications built before GPT-4 Turbo, GPT-4o, and newer reasoning-capable model families.

The First GPT-4 Production Baseline

For many teams, gpt-4 was the model that moved large language models from experimentation into higher-stakes production workflows. Applications were built around its response style, prompt sensitivity, context limits, output patterns, and relatively high cost.

The current OpenAI model entry describes GPT-4 as an older high-intelligence model usable in Chat Completions. It is text-only, with an 8,192-token context window and an 8,192-token maximum output.

That profile is very different from GPT-4 Turbo and GPT-4o. GPT-4 Turbo expanded context to 128K and added image input, while GPT-4o significantly reduced cost and broadened the mature multimodal production path.

Today, the main value of GPT-4 is not that it is the best model to choose for a new application. Its value is as a reproducible legacy reference for systems that still depend on behavior established during the first GPT-4 generation.

  • Original GPT-4 API family baseline.
  • Text-only input and output.
  • 8,192-token context window.
  • Deprecated with an announced API shutdown date.
Legacy model profile
Provider
OpenAI
Family
GPT-4
Model ID
gpt-4
Knowledge cutoff
Dec 1, 2023
Input modality
Text
Output modality
Text
Lifecycle
Deprecated

02 / Legacy baseline

Why GPT-4 Still Matters for Legacy Applications

An older model can remain operationally important when years of prompts, tests, parsing logic, fine-tuned behavior, and user expectations were built around it.

Migration Risk Lives in Application Assumptions

Changing gpt-4 to a newer model ID is easy. Proving that the application still behaves correctly is the difficult part.

A legacy system may depend on subtle properties that are not documented in a model specification: how verbose answers are, how the model handles ambiguous instructions, which wording produces reliable classifications, how often it follows a requested format, or how a fine-tuned variant behaves on domain-specific inputs.

Those assumptions should be made explicit before migration. Collect representative prompts and known edge cases, record the current GPT-4 output, and define what counts as an acceptable replacement.

This converts an urgent deprecation project into an evaluation problem. Instead of asking whether the new model is generally better, you can ask whether it preserves or improves the exact behaviors your application needs.

Why GPT-4 still matters
  1. 01

    Capture known behavior

    Save representative GPT-4 responses, failure cases, output formats, and quality thresholds before changing the model.

  2. 02

    Find hidden dependencies

    Identify prompts, parsers, retries, post-processing, and business rules that assume GPT-4-specific behavior.

  3. 03

    Test the replacement

    Run the same workload against a current model under equivalent application conditions.

  4. 04

    Migrate with evidence

    Switch only after quality, compatibility, latency, and cost meet the acceptance criteria.

03 / Pricing

GPT-4 Pricing

GPT-4 costs $30 per million input tokens and $60 per million output tokens, making it dramatically more expensive than later GPT families.

Historical Cost Is a Strong Migration Incentive

GPT-4 pricing reflects an earlier generation of API economics. Output tokens cost twice as much as input tokens, and even moderate request volumes can become expensive compared with current alternatives.

For example, a request with 4,000 input tokens and 1,000 output tokens costs about $0.18 before accounting for retries or repeated application calls.

That can matter even more in older systems where prompts were written before aggressive cost optimization became common. Large few-shot examples, duplicated instructions, and verbose responses can multiply spend.

When evaluating a replacement, compare cost per successful task rather than token price alone. A newer model may also reduce retries, allow shorter prompts, handle structured output more reliably, or consolidate several old processing steps.

  • $30.00 per 1M input tokens.
  • $60.00 per 1M output tokens.
  • No cached-input price is listed for the legacy GPT-4 model.
  • Production cost can fall substantially when moving to a newer model that passes the same evaluation set.
Legacy token pricing

1M tokens Β· USD

Input
$30.00
Output
$60.00

Example: 4K input + 1K output

Input cost
$0.1200
Output cost
$0.0600
Estimated total
$0.1800

04 / Context

GPT-4 Has an 8,192-Token Context Window

The original GPT-4 model supports 8,192 tokens of context, a small window by current standards and one of the clearest architectural constraints of legacy GPT-4 applications.

Old Applications Often Had to Manage Context Aggressively

With an 8K window, conversation history, system instructions, user input, examples, and generated output all compete for limited space.

That pushed early GPT-4 applications toward context-management strategies that may still exist in the codebase today: trimming old messages, summarizing history, retrieving only a few documents, splitting source material into chunks, or performing several sequential calls.

A modern replacement with a much larger window can simplify some of that logic, but migration should not automatically remove it. Context selection and retrieval can still improve quality and cost even when the replacement has far more capacity.

The key migration question is whether old context-management code remains useful or is now limiting the application unnecessarily.

  • Context window: 8,192 tokens.
  • Maximum output: 8,192 tokens.
  • Far smaller input capacity than GPT-4 Turbo, GPT-4o, GPT-4.1, and current-generation models.
  • Legacy truncation and summarization logic should be reviewed during migration.
Context capacity

Context window

8,192

Max output

8,192

Shared context budgetOutput ceiling

Legacy GPT-4 applications often use explicit history trimming, summarization, retrieval, or chunking because the available context is small by modern standards.

05 / Capabilities

GPT-4 API Capabilities

The current GPT-4 model entry is intentionally limited compared with newer OpenAI models: text input and output, streaming, and fine-tuning are supported, while multimodal input and newer structured application features are not.

A Much Narrower Integration Surface Than Modern Models

GPT-4 is text-only. Image, audio, and video input are not supported by the standard model entry.

OpenAI currently lists streaming as supported. Fine-tuning is also supported, which matters for legacy deployments that invested in specialized GPT-4 behavior.

By contrast, the current model documentation lists function calling and Structured Outputs as unsupported for gpt-4. Predicted outputs are also unavailable.

This narrower capability surface is another reason migrations can become architectural upgrades rather than simple model substitutions. A replacement can potentially eliminate custom parsing, add multimodal inputs, expose modern tool workflows, or provide schema-constrained responses that were not available in the original GPT-4 path.

Supported capabilities
  • Text

    Text input and text output.

    Supported
  • Streaming

    Receive generated output progressively.

    Supported
  • Fine-tuning

    Supported for legacy GPT-4 specialization workflows.

    Supported
  • Image input

    The standard GPT-4 model entry is text-only.

    Not listed
  • Function calling

    The current OpenAI GPT-4 model page lists function calling as unsupported.

    Not listed
  • Structured outputs

    Modern schema-constrained Structured Outputs are not supported.

    Not listed
  • Predicted outputs

    Predicted output optimization is not available.

    Not listed

06 / Fine-tuning

Fine-Tuned GPT-4 Workloads Need Their Own Migration Plan

Fine-tuning makes migration more complex because the production dependency is not only the GPT-4 base model but also behavior learned from a custom training dataset.

Preserve the Evaluation Dataset, Not Just the Model ID

If an application uses a fine-tuned GPT-4 model, the training examples and evaluation cases are valuable migration assets.

The first step is to establish how much the fine-tune improves over the base model on the actual workload. Then test whether a modern base model already matches that quality without customization.

If specialization is still required, reproduce the behavior using a currently supported fine-tuning path rather than assuming the old training configuration should be copied unchanged.

OpenAI's deprecation schedule also lists fine-tuned GPT-4 versions for removal on October 23, 2026, with GPT-5.6 Sol as the recommended replacement base model.

Fine-tuned migration
  1. 01

    Preserve training examples

    Keep the dataset and formatting that produced the legacy GPT-4 specialization.

  2. 02

    Build a held-out evaluation

    Separate representative test cases from training examples so replacement quality can be measured objectively.

  3. 03

    Test a modern base model first

    Determine whether a newer model already meets the task requirements without another fine-tune.

  4. 04

    Specialize only when needed

    Use a supported modern customization path if evaluation still shows a meaningful benefit.

07 / Migration

Migrating From GPT-4 Before October 23, 2026

OpenAI has deprecated the GPT-4 alias and schedules API access to end on October 23, 2026, with GPT-5.6 Sol listed as the recommended replacement.

Treat the Deadline as a Compatibility Project

A safe migration should reproduce the old workload rather than starting with synthetic benchmark prompts.

Collect real requests from the production distribution: short questions, long instructions, difficult edge cases, formatting-sensitive outputs, fine-tuned tasks, and prompts that historically required retries.

Run those requests against GPT-4 and the replacement while GPT-4 is still available. Compare correctness, output format, latency, token usage, cost, and any downstream validation failures.

The much larger context windows and newer capabilities of current models may allow architectural simplification, but do that after establishing parity. Separating model replacement from broader application redesign makes regressions easier to diagnose.

Migration workflow
  1. 01

    Freeze the GPT-4 baseline

    Capture representative prompts, outputs, quality scores, failure cases, token usage, and cost before shutdown.

  2. 02

    Test GPT-5.6 Sol

    Run the same evaluation set against OpenAI's recommended replacement under equivalent application conditions.

  3. 03

    Inspect compatibility

    Compare output shape, instruction following, verbosity, latency, context behavior, and downstream parser success.

  4. 04

    Retire legacy assumptions

    After parity is established, remove obsolete context trimming, custom formatting workarounds, or other GPT-4-era constraints where appropriate.

08 / Evaluation

GPT-4 Strengths and Limitations

GPT-4's current value is primarily historical and operational: it gives legacy applications a known baseline that can be measured before migration, but its cost, context size, modality support, and API features are far behind current models.

Legacy strengths

  • Known production behavior

    Older applications may have years of prompt tuning, acceptance tests, and operational knowledge built around GPT-4.

  • Useful migration baseline

    The model provides a concrete reference for proving that a replacement preserves or improves existing behavior.

  • Fine-tuning support

    Legacy customized GPT-4 deployments can be evaluated against newer base or fine-tuned alternatives.

  • Simple text model

    For text-only legacy systems, the model's integration surface is narrow and well understood.

What to consider

  • Imminent API shutdown

    OpenAI schedules GPT-4 removal for October 23, 2026.

  • Very high token cost

    $30 input and $60 output per million tokens are expensive compared with newer model generations.

  • 8K context window

    The small context limit creates constraints that later GPT-4 and current-generation models largely remove.

  • No multimodal or modern structured features

    The standard GPT-4 model is text-only and does not support Structured Outputs or the modern tool-oriented feature set.

Preserve behavior before retirement

Benchmark your GPT-4 workload before migrating

Run the prompts, long conversations, fine-tuned behaviors, output formats, and edge cases your legacy application depends on, then compare them against a current replacement model.

Start Free

GPT-4 is deprecated and scheduled for API shutdown on October 23, 2026. Use evaluation data to move legacy production traffic before the deadline.

Common Questions

What is GPT-4?

GPT-4 is OpenAI's original GPT-4 API model. It is now a deprecated text-only legacy model with an 8,192-token context window and a long history of production use.

How much does GPT-4 cost?

GPT-4 costs $30 per 1M input tokens and $60 per 1M output tokens.

What is the GPT-4 context window?

The current OpenAI model page lists an 8,192-token context window and an 8,192-token maximum output.

Does GPT-4 support images?

No. The standard gpt-4 model entry supports text input and text output. Image, audio, and video input are not supported.

Does GPT-4 support function calling?

OpenAI's current GPT-4 model documentation lists function calling as unsupported for the gpt-4 model entry.

Does GPT-4 support Structured Outputs?

No. Structured Outputs are not supported by the legacy gpt-4 model.

Can GPT-4 be fine-tuned?

Yes. OpenAI currently lists fine-tuning as supported for GPT-4, although fine-tuned GPT-4 variants are also scheduled for retirement.

Is GPT-4 deprecated?

Yes. OpenAI has deprecated gpt-4 and schedules API shutdown for October 23, 2026.

What should replace GPT-4?

OpenAI's current deprecation schedule recommends GPT-5.6 Sol as the replacement for gpt-4 and legacy GPT-4 fine-tuned workloads.

How is GPT-4 different from GPT-4 Turbo?

GPT-4 is the older text-only model with an 8K context window and $30/$60 token pricing. GPT-4 Turbo expanded context to 128K, added image input, and reduced pricing to $10/$30, making it a distinct later generation of the GPT-4 family.

How is GPT-4 different from GPT-4o?

GPT-4o is a later multimodal model with image input, a 128K context window, broader application capabilities, and substantially lower token pricing. GPT-4 is primarily relevant today as a legacy compatibility and migration baseline.

Model information

Last updated

Specifications, pricing, capabilities, fine-tuning support, and lifecycle information on this page are based on the official OpenAI GPT-4 model documentation and deprecation schedule. OpenAI currently lists gpt-4 as deprecated and recommends GPT-5.6 Sol for migration.

GPT-4 β€” Pricing, 8K Context, Fine-Tuning & Migration | EidoStack