OpenAI named Scale AI a “preferred partner” on August 24, 2023, to help businesses fine-tune OpenAI models, starting with GPT-3.5 Turbo. Scale offered data preparation and enterprise implementation support; it did not provide exclusive access to a special GPT-3.5 model. The partnership is now mainly a historical example: OpenAI’s current documentation marks GPT-3.5 Turbo as deprecated and says its fine-tuning platform is being wound down.
What OpenAI and Scale AI announced
OpenAI first announced self-serve fine-tuning for GPT-3.5 Turbo on August 22, 2023. Two days later, it said Scale AI would be a preferred partner helping enterprises customize OpenAI models with their own data. OpenAI described Scale customers as being able to fine-tune models much as they could through OpenAI itself. The designation therefore meant an endorsed enterprise-services relationship, not an exclusive route or a separately licensed model. OpenAI’s fine-tuning launch announcement and its Scale partnership announcement set out the two steps.
In the 2023 announcement, OpenAI also said GPT-4 fine-tuning was expected later that year. That was a statement of intent at the time, not a description of present-day availability.
What Scale brought beyond the API
OpenAI already offered the fine-tuning API. Scale’s proposed contribution was the work around it: helping enterprises turn proprietary information into useful examples, annotate data, rank model outputs, evaluate customized models, and move a solution toward production. OpenAI cited Scale’s enterprise AI experience and Data Engine; Scale’s own account described tools for generating prompts and ranking outputs. Scale’s account of the partnership provides its perspective.
#1 Best Overall
This kind of support could matter to a company without an experienced machine-learning or data team. But it added vendor, procurement, and data-handling considerations as well as potential implementation fees. The cited public announcements did not state Scale’s service prices or settle questions such as ownership of intermediate annotations and evaluation artifacts; those details would need to be established contractually.
What fine-tuning could—and could not—do
Fine-tuning starts with a base model and trains it further on examples of the desired task or behavior. It can be useful for consistent formatting, tone, classification, routing, or repeated task-specific response patterns. OpenAI’s 2023 examples included generating code in a particular language, summarizing text in a defined format, and producing personalized content. OpenAI’s fine-tuning guidance also describes these kinds of use cases.
Rank #2
Fine-tuning is not simply a way to load a large, changing knowledge base into a model and ensure it will recall every fact. For frequently updated policies, inventory, prices, or document collections, retrieval-augmented generation—where the system fetches relevant source material at answer time—is generally a better fit. Fine-tuning may reduce the prompt needed for a stable task and can improve consistency, latency, or cost in a suitable workload, but those gains need to be measured on that workload.
- More promising: A repeated task, reliable labeled examples, and an objective way to judge output quality.
- Less promising: Frequently changing facts, weak or inconsistent labels, no evaluation criteria, or a need to control access to individual source documents.
What the Brex example showed
OpenAI highlighted Brex, which used language models to generate employee expense memos and reports. Brex had been using GPT-4 and explored whether a fine-tuned GPT-3.5 model could deliver suitable quality at lower cost and latency. Scale’s Data Engine was used to annotate Brex data. Scale reported that the resulting model outperformed stock GPT-3.5 Turbo 66% of the time in Brex’s evaluation. Contemporary coverage by VentureBeat also reported the case study.
Rank #3
The 66% figure is a company-reported result for this particular task, not a claim that the model was “66% better” or that fine-tuning generally makes GPT-3.5 outperform its base model. The public material does not establish the evaluation-set size, task mix, scoring method, use of human or automated judges, prompt controls, or whether testing used unseen data. It also does not show how large the wins were or whether they generalize to other work.
Data, safety, and enterprise checks
OpenAI said customers owned the data sent to and returned from its fine-tuning API, and that it would not use that data to train other models. That statement addresses ownership and model-training use; it does not, by itself, answer every question about retention, access controls, data residency, contractual protections, or how a separate services provider handles information. Organizations would need to assess those points with each vendor before sharing confidential or regulated data. OpenAI’s partnership announcement and its 2023 API announcement describe its stated policy.
Rank #4
OpenAI also said fine-tuning data passed through the Moderation API and a GPT-4-powered moderation system to identify unsafe training data. That is a screening measure, not proof that every output from a tuned model will be safe. Deployment still calls for holdout and regression testing, tests for abuse and prompt injection, human review where mistakes carry serious consequences, and ongoing monitoring.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Historical launch pricing
At the August 2023 launch, OpenAI listed the following GPT-3.5 Turbo fine-tuning prices. These are historical launch figures, not current quotes or evidence of present availability.
| Charge at launch | August 2023 price |
|---|---|
| Training | $0.008 per 1,000 tokens |
| Input to a fine-tuned model | $0.012 per 1,000 tokens |
| Output from a fine-tuned model | $0.016 per 1,000 tokens |
OpenAI’s launch example estimated $2.40 to train a 100,000-token file for three epochs. That estimate covered the stated training example, not Scale’s enterprise services or the full cost of operating a production system. The original pricing and example appeared in OpenAI’s announcement.
Where the GPT-3.5 fine-tuning path stands now
OpenAI’s current GPT-3.5 Turbo model documentation labels the model deprecated and says developers should use GPT-4o mini instead for many GPT-3.5 use cases. Separately, an update dated May 8, 2026, to OpenAI’s fine-tuning API and custom-models announcement says the fine-tuning platform is being wound down: new users can no longer access it, existing users retain access for a limited period, and fine-tuned models remain available for inference until their base models are deprecated.
That means the 2023 partnership should not be read as a straightforward new GPT-3.5 implementation path. A team considering model customization now should verify current access and supported models directly, and weigh the lifespan of the base model alongside any training or integration investment.
Lessons for enterprise buyers
The durable lesson is that customization depends on more than an API endpoint. Data quality, evaluation design, governance, deployment, and a migration plan all affect whether a tuned model is worthwhile. Before committing to a similar project, an organization should have:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsQuick Recap
- A high-volume, well-defined task and examples that reflect real production inputs.
- Separate training and holdout evaluation data, with criteria that capture costly errors rather than only average preference.
- A fair comparison against a well-prompted base model, including quality, latency, and total cost.
- Clear agreements covering data handling, annotation ownership, tuned artifacts, support, and exit or export options.
- A fallback and migration plan, with versioned benchmarks and reproducible preprocessing and prompts.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




