OpenAI legacy chat model

GPT-3.5 Turbo

A legacy low-cost OpenAI model that helped define the Chat Completions era, with 16K context, text-only input and output, fine-tuning support, and a large installed base of older chat and automation workflows.

Context window
16.4K
tokens
Max output
4.1K
tokens
Input
$0.50
per 1M tokens
Output
$1.50
per 1M tokens
Shutdown
Oct 23
2026

01 / Overview

What GPT-3.5 Turbo Is

GPT-3.5 Turbo is a legacy OpenAI model optimized for chat-style interaction that also became widely used for inexpensive classification, extraction, transformation, and other non-chat API tasks.

The Model That Made Chat APIs Economical at Scale

GPT-3.5 Turbo sits at an important point in the history of production LLM applications. It made conversational interfaces inexpensive enough for broad experimentation while also becoming a default model for background automation that did not require GPT-4-level capability.

The current gpt-3.5-turbo alias uses the later 16K-class model generation rather than the earliest 4K GPT-3.5 Turbo releases. OpenAI lists a 16,385-token context window, up to 4,096 output tokens, and a September 1, 2021 knowledge cutoff.

The model accepts text and produces text. It does not support image, audio, or video input.

By current standards, GPT-3.5 Turbo is no longer the natural low-cost default. OpenAI has recommended newer small models for years and now marks the GPT-3.5 Turbo line for final retirement.

  • Model ID: gpt-3.5-turbo.
  • Optimized historically for Chat Completions and useful for non-chat text tasks.
  • 16,385-token context window.
  • Text-only input and output.
  • Deprecated with final shutdown scheduled for October 23, 2026.
Legacy model profile
Provider
OpenAI
Family
GPT-3.5 Turbo
Model ID
gpt-3.5-turbo
Knowledge cutoff
Sep 1, 2021
Input modality
Text
Output modality
Text
Lifecycle
Deprecated

02 / Chat era

GPT-3.5 Turbo and the Chat Completions Era

GPT-3.5 Turbo helped establish the message-based API pattern that many production assistants, support bots, and automation systems still resemble today.

From Prompt Strings to Conversation Roles

Earlier GPT integrations often revolved around one prompt string sent to a completion model. GPT-3.5 Turbo popularized a more application-oriented conversation structure: system instructions, user messages, assistant messages, and persistent chat history.

That changed how developers designed AI products. Instead of rebuilding conversational context into one giant text prompt, applications could model the interaction as a sequence of messages.

The model was also cheap enough to use outside visible chat. Teams adopted it for intent classification, metadata generation, short summaries, extraction, rewriting, routing, and other repeated background tasks.

Many of those workloads still exist in mature products. The migration challenge is therefore broader than replacing a chatbot model: a single codebase may contain dozens of small GPT-3.5 Turbo calls that were added over several years.

Why GPT-3.5 Turbo mattered
  1. 01

    Message-based chat

    System, user, and assistant messages became a standard application pattern for conversational AI.

  2. 02

    Affordable automation

    Low pricing made repeated classification, extraction, summarization, and transformation calls practical.

  3. 03

    Fast product iteration

    Teams could add model-backed features without routing every request through the more expensive GPT-4 tier.

  4. 04

    Large legacy footprint

    Years of prompts, fine-tunes, parsers, and business logic may still depend on GPT-3.5-specific behavior.

03 / Legacy workloads

Where GPT-3.5 Turbo Appears in Existing Applications

GPT-3.5 Turbo is often found in narrow, high-volume workflow steps rather than only in user-facing chat, which can make migration inventory more important than expected.

Look for Small Calls Hidden Throughout the Product

A mature application may use one stronger model for its headline AI feature while still relying on GPT-3.5 Turbo behind the scenes.

Common examples include routing a request to the correct workflow, extracting a small set of fields, translating short text, rewriting copy, generating tags, classifying support tickets, or summarizing content before another process runs.

These tasks are good migration candidates because current small models can often deliver stronger capability at lower cost. But each replacement should still be evaluated against the actual task contract.

A classifier with 99% acceptable accuracy can be more valuable than a more articulate model that changes labels unpredictably. An extraction step may care more about parse success than prose quality. Migration criteria should match the workload.

Typical legacy workloads
  1. 01

    Classification and routing

    Intent detection, support triage, category assignment, priority labeling, and workflow selection.

  2. 02

    Extraction

    Pull names, fields, attributes, entities, or compact records from text.

  3. 03

    Short transformations

    Rewrite, summarize, translate, normalize, or reformat relatively small text inputs.

  4. 04

    Legacy chat assistants

    Customer-facing and internal bots built around GPT-3.5-era prompts and message-history logic.

04 / Pricing

GPT-3.5 Turbo Pricing

GPT-3.5 Turbo costs $0.50 per million input tokens and $1.50 per million output tokens.

Once Cheap, Now Expensive Relative to Newer Small Models

GPT-3.5 Turbo earned much of its adoption through price. Compared with GPT-4-era models, $0.50 input and $1.50 output pricing made high-volume API features dramatically easier to justify.

The market moved quickly. OpenAI's own model page now compares GPT-3.5 Turbo against GPT-4o Mini at $0.15 input and $0.60 output per million tokens, showing why GPT-3.5 Turbo is no longer the obvious economy option.

Cost migration should still be measured at the workload level. A newer model can reduce spend through lower token rates, but it may also use a different number of tokens, need fewer retries, or support more reliable structured workflows.

For narrow automation tasks, even fractions of a dollar per million tokens compound when the system executes millions of calls.

  • $0.50 per 1M input tokens.
  • $1.50 per 1M output tokens.
  • No cached-input price is listed for GPT-3.5 Turbo.
  • Newer small models can be both cheaper and more capable.
Legacy token pricing

1M tokens · USD

Input
$0.50
Output
$1.50

Example: 6K input + 1K output

Input cost
$0.0030
Output cost
$0.0015
Estimated total
$0.0045

05 / Context

GPT-3.5 Turbo Has a 16,385-Token Context Window

The current GPT-3.5 Turbo model generation provides a 16,385-token context window and up to 4,096 output tokens.

16K Was a Major Upgrade, but It Is Small Today

Earlier GPT-3.5 Turbo versions were associated with much smaller context budgets. In November 2023, OpenAI updated the default GPT-3.5 Turbo generation to a 16K-class context window.

That expansion made longer conversations, more examples, and larger source texts practical without immediate chunking. It also reduced the pressure to aggressively summarize chat history.

By 2026, 16K is a legacy-scale context window. Current models can support substantially larger prompts, which means older trimming, summarization, and chunking logic may be more restrictive than necessary after migration.

Do not remove that logic blindly, however. Selective context can still improve cost and relevance. The migration task is to determine which constraints were required by GPT-3.5 Turbo and which remain good application design.

  • Context window: 16,385 tokens.
  • Maximum output: 4,096 tokens.
  • Larger than early GPT-3.5 Turbo generations.
  • Much smaller than modern 128K and million-token models.
Context capacity

Context window

16,385

Max output

4,096

Prompt + conversationOutput limit

Legacy applications may include history trimming, summarization, retrieval, or chunking logic designed around GPT-3.5 Turbo's relatively small context budget.

06 / Fine-tuning

Fine-Tuning Was a Key GPT-3.5 Turbo Customization Path

GPT-3.5 Turbo supports fine-tuning, and many production systems used customized variants to improve stable classifications, formats, tone, or domain-specific behavior.

Fine-Tuned Models Need More Than a Model-ID Swap

A fine-tuned GPT-3.5 Turbo deployment contains value in three places: the training dataset, the evaluation set, and the application assumptions built around the resulting behavior.

Preserve all three before migration.

OpenAI's current deprecation schedule states that fine-tuned GPT-3.5 Turbo versions are also scheduled for shutdown on October 23, 2026, with GPT-5.6 Terra listed as the recommended replacement base model.

A modern base model may already outperform the old fine-tune without customization. Test that first. If the task still benefits from specialization, rebuild it on a supported current model using a held-out evaluation set rather than simply recreating the old configuration.

Fine-tuned migration
  1. 01

    Archive the training data

    Preserve the examples, labels, formatting conventions, and domain-specific patterns used by the legacy fine-tune.

  2. 02

    Separate evaluation data

    Build a held-out set that measures the production behavior you need to preserve.

  3. 03

    Test a modern base model

    Check whether current capability already eliminates the need for customization.

  4. 04

    Fine-tune only if necessary

    Use a supported current base model when specialization still produces measurable value.

07 / Capabilities

GPT-3.5 Turbo API Capabilities

The current OpenAI model page describes GPT-3.5 Turbo as a text-only legacy model with fine-tuning support and a much narrower feature surface than modern OpenAI models.

Do Not Assume Modern API Features Are Available

GPT-3.5 Turbo accepts and generates text. Image and audio modalities are not supported.

Fine-tuning remains listed as supported. However, OpenAI's current model documentation lists streaming, function calling, Structured Outputs, and predicted outputs as unsupported for the current deprecated model entry.

That matters when documenting or migrating old systems. Historical GPT-3.5 snapshots introduced capabilities such as function calling, but the current supported surface for the deprecated alias should be taken from current provider documentation rather than old integration assumptions.

A replacement model can therefore provide not only better quality but also a broader application contract: multimodal input, structured schemas, tool calling, larger context, or reasoning controls depending on the chosen model.

Current documented capabilities
  • Text

    Text input and text output.

    Supported
  • Fine-tuning

    Fine-tuning remains listed as supported for GPT-3.5 Turbo.

    Supported
  • Image input

    GPT-3.5 Turbo is text-only.

    Not listed
  • Streaming

    The current OpenAI model page lists streaming as unsupported for this deprecated model entry.

    Not listed
  • Function calling

    The current OpenAI model page lists function calling as unsupported.

    Not listed
  • Structured outputs

    Modern Structured Outputs are not supported.

    Not listed
  • Predicted outputs

    Predicted outputs are not supported.

    Not listed

08 / Migration

Migrating From GPT-3.5 Turbo

The recommended replacement path for GPT-3.5 Turbo has changed over time, which reflects how quickly the low-cost model tier has evolved.

From GPT-4o Mini to GPT-5.6 Terra

OpenAI's GPT-3.5 Turbo model page notes that, as of July 2024, developers should prefer GPT-4o Mini because it is cheaper, more capable, multimodal, and just as fast.

That was the practical migration advice during the GPT-4o generation.

The current 2026 deprecation schedule goes further. The final gpt-3.5-turbo-0125 snapshot—which backs the gpt-3.5-turbo alias—is scheduled for shutdown on October 23, 2026, and OpenAI lists GPT-5.6 Terra as the substitute. Fine-tuned GPT-3.5 Turbo versions have the same shutdown date and replacement base-model recommendation.

For existing products, migration should begin with inventory. Find every place where GPT-3.5 Turbo is called, classify each workload, and evaluate replacements task by task. The best replacement for a classifier may not be the same as the best replacement for a customer-facing assistant.

Migration path
  1. 01

    Inventory every GPT-3.5 call

    Search the application, workers, scheduled jobs, prompts, and configuration for direct or indirect dependencies.

  2. 02

    Group calls by workload

    Separate chat, classification, extraction, rewriting, summarization, routing, and fine-tuned tasks.

  3. 03

    Evaluate a current replacement

    Use GPT-5.6 Terra or another appropriate current model on representative production examples.

  4. 04

    Remove legacy assumptions

    After parity is proven, simplify context limits, parsing workarounds, retries, or other GPT-3.5-era code where appropriate.

09 / Evaluation

GPT-3.5 Turbo Strengths and Limitations

GPT-3.5 Turbo remains useful as a legacy benchmark because many applications were built around its economics and response behavior, but its context size, modalities, feature set, and lifecycle no longer make it a sensible new dependency.

Legacy strengths

  • Known production behavior

    Years of prompts, acceptance criteria, fine-tunes, and workflow logic may already be calibrated around GPT-3.5 Turbo.

  • Low historical cost

    The model helped make high-volume chat and text automation economically practical.

  • Fine-tuning support

    Many narrow production tasks were customized successfully on GPT-3.5 Turbo.

  • Useful migration baseline

    Existing quality and cost data provide a concrete reference for evaluating newer small models.

What to consider

  • Imminent shutdown

    The final GPT-3.5 Turbo alias path is scheduled for removal on October 23, 2026.

  • Text-only

    Image, audio, and video inputs are not supported.

  • 16K context

    The context window is small compared with current 128K and million-token model families.

  • Narrow current feature surface

    The current model documentation does not list streaming, function calling, Structured Outputs, or predicted outputs as supported.

Retire the legacy dependency with evidence

Benchmark GPT-3.5 Turbo before migrating

Run the chat prompts, classifiers, extraction tasks, fine-tuned behaviors, and formatting-sensitive workflows your application still depends on, then compare them against a current replacement model.

Start Free

GPT-3.5 Turbo is deprecated. Preserve the production behaviors that matter, measure a replacement, and migrate before the October 23, 2026 shutdown.

Common Questions

What is GPT-3.5 Turbo?

GPT-3.5 Turbo is a legacy OpenAI text model optimized historically for Chat Completions and also used widely for inexpensive non-chat tasks such as classification, extraction, rewriting, and summarization.

How much does GPT-3.5 Turbo cost?

GPT-3.5 Turbo costs $0.50 per 1M input tokens and $1.50 per 1M output tokens.

What is the GPT-3.5 Turbo context window?

The current GPT-3.5 Turbo model generation has a 16,385-token context window and supports up to 4,096 output tokens.

What is the GPT-3.5 Turbo knowledge cutoff?

OpenAI lists September 1, 2021 as the knowledge cutoff for GPT-3.5 Turbo.

Does GPT-3.5 Turbo support images?

No. GPT-3.5 Turbo supports text input and text output. Image and audio input are not supported.

Can GPT-3.5 Turbo be fine-tuned?

Yes. OpenAI currently lists fine-tuning as supported, although fine-tuned GPT-3.5 Turbo variants are also scheduled for shutdown on October 23, 2026.

Does GPT-3.5 Turbo support function calling?

The current OpenAI GPT-3.5 Turbo model page lists function calling as unsupported for the deprecated model entry. Historical GPT-3.5 snapshots introduced function-calling capabilities, so legacy integrations should be checked against the exact model version they used.

Does GPT-3.5 Turbo support Structured Outputs?

No. Modern Structured Outputs are not supported by GPT-3.5 Turbo.

Is GPT-3.5 Turbo deprecated?

Yes. The gpt-3.5-turbo alias is deprecated, and its final gpt-3.5-turbo-0125 snapshot is scheduled for shutdown on October 23, 2026.

What should replace GPT-3.5 Turbo?

OpenAI's current deprecation schedule lists GPT-5.6 Terra as the substitute for the final GPT-3.5 Turbo snapshot. Earlier guidance recommended GPT-4o Mini as a cheaper, more capable, multimodal replacement, illustrating how the preferred low-cost model tier has evolved.

How is GPT-3.5 Turbo different from GPT-4?

GPT-3.5 Turbo was the lower-cost, faster legacy tier with a 16K context window and $0.50/$1.50 token pricing. GPT-4 was a much more expensive higher-capability model with an 8K context window and $30/$60 pricing. Both are now legacy models approaching retirement.

Should I build a new application with GPT-3.5 Turbo?

No. GPT-3.5 Turbo is deprecated and scheduled for shutdown. New applications should start with a current supported model selected through evaluation of quality, latency, modalities, context needs, and cost.

Model information

Last updated

Specifications, pricing, modalities, fine-tuning support, snapshots, and lifecycle information on this page are based on the official OpenAI GPT-3.5 Turbo model documentation and deprecation schedule. OpenAI lists gpt-3.5-turbo as deprecated, with the final alias path scheduled for shutdown on October 23, 2026.

GPT-3.5 Turbo — Pricing, 16K Context, Fine-Tuning & Migration | EidoStack