OpenAI legacy chat model
GPT-3.5 Turbo
A legacy low-cost OpenAI model that helped define the Chat Completions era, with 16K context, text-only input and output, fine-tuning support, and a large installed base of older chat and automation workflows.
- Context window
- 16.4K
- tokens
- Max output
- 4.1K
- tokens
- Input
- $0.50
- per 1M tokens
- Output
- $1.50
- per 1M tokens
- Shutdown
- Oct 23
- 2026
01 / Overview
What GPT-3.5 Turbo Is
GPT-3.5 Turbo is a legacy OpenAI model optimized for chat-style interaction that also became widely used for inexpensive classification, extraction, transformation, and other non-chat API tasks.
The Model That Made Chat APIs Economical at Scale
GPT-3.5 Turbo sits at an important point in the history of production LLM applications. It made conversational interfaces inexpensive enough for broad experimentation while also becoming a default model for background automation that did not require GPT-4-level capability.
The current gpt-3.5-turbo alias uses the later 16K-class model generation rather than the earliest 4K GPT-3.5 Turbo releases. OpenAI lists a 16,385-token context window, up to 4,096 output tokens, and a September 1, 2021 knowledge cutoff.
The model accepts text and produces text. It does not support image, audio, or video input.
By current standards, GPT-3.5 Turbo is no longer the natural low-cost default. OpenAI has recommended newer small models for years and now marks the GPT-3.5 Turbo line for final retirement.
- Model ID:
gpt-3.5-turbo. - Optimized historically for Chat Completions and useful for non-chat text tasks.
- 16,385-token context window.
- Text-only input and output.
- Deprecated with final shutdown scheduled for October 23, 2026.
- Provider
- OpenAI
- Family
- GPT-3.5 Turbo
- Model ID
- gpt-3.5-turbo
- Knowledge cutoff
- Sep 1, 2021
- Input modality
- Text
- Output modality
- Text
- Lifecycle
- Deprecated
02 / Chat era
GPT-3.5 Turbo and the Chat Completions Era
GPT-3.5 Turbo helped establish the message-based API pattern that many production assistants, support bots, and automation systems still resemble today.
From Prompt Strings to Conversation Roles
Earlier GPT integrations often revolved around one prompt string sent to a completion model. GPT-3.5 Turbo popularized a more application-oriented conversation structure: system instructions, user messages, assistant messages, and persistent chat history.
That changed how developers designed AI products. Instead of rebuilding conversational context into one giant text prompt, applications could model the interaction as a sequence of messages.
The model was also cheap enough to use outside visible chat. Teams adopted it for intent classification, metadata generation, short summaries, extraction, rewriting, routing, and other repeated background tasks.
Many of those workloads still exist in mature products. The migration challenge is therefore broader than replacing a chatbot model: a single codebase may contain dozens of small GPT-3.5 Turbo calls that were added over several years.
- 01
Message-based chat
System, user, and assistant messages became a standard application pattern for conversational AI.
- 02
Affordable automation
Low pricing made repeated classification, extraction, summarization, and transformation calls practical.
- 03
Fast product iteration
Teams could add model-backed features without routing every request through the more expensive GPT-4 tier.
- 04
Large legacy footprint
Years of prompts, fine-tunes, parsers, and business logic may still depend on GPT-3.5-specific behavior.
03 / Legacy workloads
Where GPT-3.5 Turbo Appears in Existing Applications
GPT-3.5 Turbo is often found in narrow, high-volume workflow steps rather than only in user-facing chat, which can make migration inventory more important than expected.
Look for Small Calls Hidden Throughout the Product
A mature application may use one stronger model for its headline AI feature while still relying on GPT-3.5 Turbo behind the scenes.
Common examples include routing a request to the correct workflow, extracting a small set of fields, translating short text, rewriting copy, generating tags, classifying support tickets, or summarizing content before another process runs.
These tasks are good migration candidates because current small models can often deliver stronger capability at lower cost. But each replacement should still be evaluated against the actual task contract.
A classifier with 99% acceptable accuracy can be more valuable than a more articulate model that changes labels unpredictably. An extraction step may care more about parse success than prose quality. Migration criteria should match the workload.
- 01
Classification and routing
Intent detection, support triage, category assignment, priority labeling, and workflow selection.
- 02
Extraction
Pull names, fields, attributes, entities, or compact records from text.
- 03
Short transformations
Rewrite, summarize, translate, normalize, or reformat relatively small text inputs.
- 04
Legacy chat assistants
Customer-facing and internal bots built around GPT-3.5-era prompts and message-history logic.
04 / Pricing
GPT-3.5 Turbo Pricing
GPT-3.5 Turbo costs $0.50 per million input tokens and $1.50 per million output tokens.
Once Cheap, Now Expensive Relative to Newer Small Models
GPT-3.5 Turbo earned much of its adoption through price. Compared with GPT-4-era models, $0.50 input and $1.50 output pricing made high-volume API features dramatically easier to justify.
The market moved quickly. OpenAI's own model page now compares GPT-3.5 Turbo against GPT-4o Mini at $0.15 input and $0.60 output per million tokens, showing why GPT-3.5 Turbo is no longer the obvious economy option.
Cost migration should still be measured at the workload level. A newer model can reduce spend through lower token rates, but it may also use a different number of tokens, need fewer retries, or support more reliable structured workflows.
For narrow automation tasks, even fractions of a dollar per million tokens compound when the system executes millions of calls.
- $0.50 per 1M input tokens.
- $1.50 per 1M output tokens.
- No cached-input price is listed for GPT-3.5 Turbo.
- Newer small models can be both cheaper and more capable.
1M tokens · USD
- Input
- $0.50
- Output
- $1.50
Example: 6K input + 1K output
- Input cost
- $0.0030
- Output cost
- $0.0015
- Estimated total
- $0.0045
05 / Context
GPT-3.5 Turbo Has a 16,385-Token Context Window
The current GPT-3.5 Turbo model generation provides a 16,385-token context window and up to 4,096 output tokens.
16K Was a Major Upgrade, but It Is Small Today
Earlier GPT-3.5 Turbo versions were associated with much smaller context budgets. In November 2023, OpenAI updated the default GPT-3.5 Turbo generation to a 16K-class context window.
That expansion made longer conversations, more examples, and larger source texts practical without immediate chunking. It also reduced the pressure to aggressively summarize chat history.
By 2026, 16K is a legacy-scale context window. Current models can support substantially larger prompts, which means older trimming, summarization, and chunking logic may be more restrictive than necessary after migration.
Do not remove that logic blindly, however. Selective context can still improve cost and relevance. The migration task is to determine which constraints were required by GPT-3.5 Turbo and which remain good application design.
- Context window: 16,385 tokens.
- Maximum output: 4,096 tokens.
- Larger than early GPT-3.5 Turbo generations.
- Much smaller than modern 128K and million-token models.
Context window
16,385
Max output
4,096
Legacy applications may include history trimming, summarization, retrieval, or chunking logic designed around GPT-3.5 Turbo's relatively small context budget.
06 / Fine-tuning
Fine-Tuning Was a Key GPT-3.5 Turbo Customization Path
GPT-3.5 Turbo supports fine-tuning, and many production systems used customized variants to improve stable classifications, formats, tone, or domain-specific behavior.
Fine-Tuned Models Need More Than a Model-ID Swap
A fine-tuned GPT-3.5 Turbo deployment contains value in three places: the training dataset, the evaluation set, and the application assumptions built around the resulting behavior.
Preserve all three before migration.
OpenAI's current deprecation schedule states that fine-tuned GPT-3.5 Turbo versions are also scheduled for shutdown on October 23, 2026, with GPT-5.6 Terra listed as the recommended replacement base model.
A modern base model may already outperform the old fine-tune without customization. Test that first. If the task still benefits from specialization, rebuild it on a supported current model using a held-out evaluation set rather than simply recreating the old configuration.
- 01
Archive the training data
Preserve the examples, labels, formatting conventions, and domain-specific patterns used by the legacy fine-tune.
- 02
Separate evaluation data
Build a held-out set that measures the production behavior you need to preserve.
- 03
Test a modern base model
Check whether current capability already eliminates the need for customization.
- 04
Fine-tune only if necessary
Use a supported current base model when specialization still produces measurable value.
07 / Capabilities
GPT-3.5 Turbo API Capabilities
The current OpenAI model page describes GPT-3.5 Turbo as a text-only legacy model with fine-tuning support and a much narrower feature surface than modern OpenAI models.
Do Not Assume Modern API Features Are Available
GPT-3.5 Turbo accepts and generates text. Image and audio modalities are not supported.
Fine-tuning remains listed as supported. However, OpenAI's current model documentation lists streaming, function calling, Structured Outputs, and predicted outputs as unsupported for the current deprecated model entry.
That matters when documenting or migrating old systems. Historical GPT-3.5 snapshots introduced capabilities such as function calling, but the current supported surface for the deprecated alias should be taken from current provider documentation rather than old integration assumptions.
A replacement model can therefore provide not only better quality but also a broader application contract: multimodal input, structured schemas, tool calling, larger context, or reasoning controls depending on the chosen model.
- Supported
Text
Text input and text output.
- Supported
Fine-tuning
Fine-tuning remains listed as supported for GPT-3.5 Turbo.
- Not listed
Image input
GPT-3.5 Turbo is text-only.
- Not listed
Streaming
The current OpenAI model page lists streaming as unsupported for this deprecated model entry.
- Not listed
Function calling
The current OpenAI model page lists function calling as unsupported.
- Not listed
Structured outputs
Modern Structured Outputs are not supported.
- Not listed
Predicted outputs
Predicted outputs are not supported.
08 / Migration
Migrating From GPT-3.5 Turbo
The recommended replacement path for GPT-3.5 Turbo has changed over time, which reflects how quickly the low-cost model tier has evolved.
From GPT-4o Mini to GPT-5.6 Terra
OpenAI's GPT-3.5 Turbo model page notes that, as of July 2024, developers should prefer GPT-4o Mini because it is cheaper, more capable, multimodal, and just as fast.
That was the practical migration advice during the GPT-4o generation.
The current 2026 deprecation schedule goes further. The final gpt-3.5-turbo-0125 snapshot—which backs the gpt-3.5-turbo alias—is scheduled for shutdown on October 23, 2026, and OpenAI lists GPT-5.6 Terra as the substitute. Fine-tuned GPT-3.5 Turbo versions have the same shutdown date and replacement base-model recommendation.
For existing products, migration should begin with inventory. Find every place where GPT-3.5 Turbo is called, classify each workload, and evaluate replacements task by task. The best replacement for a classifier may not be the same as the best replacement for a customer-facing assistant.
- 01
Inventory every GPT-3.5 call
Search the application, workers, scheduled jobs, prompts, and configuration for direct or indirect dependencies.
- 02
Group calls by workload
Separate chat, classification, extraction, rewriting, summarization, routing, and fine-tuned tasks.
- 03
Evaluate a current replacement
Use GPT-5.6 Terra or another appropriate current model on representative production examples.
- 04
Remove legacy assumptions
After parity is proven, simplify context limits, parsing workarounds, retries, or other GPT-3.5-era code where appropriate.
09 / Evaluation
GPT-3.5 Turbo Strengths and Limitations
GPT-3.5 Turbo remains useful as a legacy benchmark because many applications were built around its economics and response behavior, but its context size, modalities, feature set, and lifecycle no longer make it a sensible new dependency.
Legacy strengths
Known production behavior
Years of prompts, acceptance criteria, fine-tunes, and workflow logic may already be calibrated around GPT-3.5 Turbo.
Low historical cost
The model helped make high-volume chat and text automation economically practical.
Fine-tuning support
Many narrow production tasks were customized successfully on GPT-3.5 Turbo.
Useful migration baseline
Existing quality and cost data provide a concrete reference for evaluating newer small models.
What to consider
Imminent shutdown
The final GPT-3.5 Turbo alias path is scheduled for removal on October 23, 2026.
Text-only
Image, audio, and video inputs are not supported.
16K context
The context window is small compared with current 128K and million-token model families.
Narrow current feature surface
The current model documentation does not list streaming, function calling, Structured Outputs, or predicted outputs as supported.
Retire the legacy dependency with evidence
Benchmark GPT-3.5 Turbo before migrating
Run the chat prompts, classifiers, extraction tasks, fine-tuned behaviors, and formatting-sensitive workflows your application still depends on, then compare them against a current replacement model.
Start FreeGPT-3.5 Turbo is deprecated. Preserve the production behaviors that matter, measure a replacement, and migrate before the October 23, 2026 shutdown.
Common Questions
What is GPT-3.5 Turbo?
GPT-3.5 Turbo is a legacy OpenAI text model optimized historically for Chat Completions and also used widely for inexpensive non-chat tasks such as classification, extraction, rewriting, and summarization.
How much does GPT-3.5 Turbo cost?
GPT-3.5 Turbo costs $0.50 per 1M input tokens and $1.50 per 1M output tokens.
What is the GPT-3.5 Turbo context window?
The current GPT-3.5 Turbo model generation has a 16,385-token context window and supports up to 4,096 output tokens.
What is the GPT-3.5 Turbo knowledge cutoff?
OpenAI lists September 1, 2021 as the knowledge cutoff for GPT-3.5 Turbo.
Does GPT-3.5 Turbo support images?
No. GPT-3.5 Turbo supports text input and text output. Image and audio input are not supported.
Can GPT-3.5 Turbo be fine-tuned?
Yes. OpenAI currently lists fine-tuning as supported, although fine-tuned GPT-3.5 Turbo variants are also scheduled for shutdown on October 23, 2026.
Does GPT-3.5 Turbo support function calling?
The current OpenAI GPT-3.5 Turbo model page lists function calling as unsupported for the deprecated model entry. Historical GPT-3.5 snapshots introduced function-calling capabilities, so legacy integrations should be checked against the exact model version they used.
Does GPT-3.5 Turbo support Structured Outputs?
No. Modern Structured Outputs are not supported by GPT-3.5 Turbo.
Is GPT-3.5 Turbo deprecated?
Yes. The gpt-3.5-turbo alias is deprecated, and its final gpt-3.5-turbo-0125 snapshot is scheduled for shutdown on October 23, 2026.
What should replace GPT-3.5 Turbo?
OpenAI's current deprecation schedule lists GPT-5.6 Terra as the substitute for the final GPT-3.5 Turbo snapshot. Earlier guidance recommended GPT-4o Mini as a cheaper, more capable, multimodal replacement, illustrating how the preferred low-cost model tier has evolved.
How is GPT-3.5 Turbo different from GPT-4?
GPT-3.5 Turbo was the lower-cost, faster legacy tier with a 16K context window and $0.50/$1.50 token pricing. GPT-4 was a much more expensive higher-capability model with an 8K context window and $30/$60 pricing. Both are now legacy models approaching retirement.
Should I build a new application with GPT-3.5 Turbo?
No. GPT-3.5 Turbo is deprecated and scheduled for shutdown. New applications should start with a current supported model selected through evaluation of quality, latency, modalities, context needs, and cost.
Model information
Last updated
Specifications, pricing, modalities, fine-tuning support, snapshots, and lifecycle information on this page are based on the official OpenAI GPT-3.5 Turbo model documentation and deprecation schedule. OpenAI lists gpt-3.5-turbo as deprecated, with the final alias path scheduled for shutdown on October 23, 2026.