Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Epoch AI Criticized Over Delayed Disclosure of OpenAI-Funded FrontierMath

OpenAI commissioned 300 FrontierMath questions and had access to much of the material. Epoch AI later acknowledged its contributor disclosures were inadequate, while the evidence does not prove training contamination or fabricated scores.
From TheFinanceBase Team6 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Epoch AI, the nonprofit behind the FrontierMath mathematics benchmark, faced criticism after it became public that OpenAI had commissioned and funded 300 questions and had access to much of the related test material. The relationship was disclosed around OpenAI’s December 20, 2024, announcement of its o3 model. Contributors and outside observers questioned whether they had been told enough, and Epoch later acknowledged that its communication was inadequate. The established concern is privileged access and reduced confidence in interpreting OpenAI’s results—not proof that OpenAI trained on the questions or falsified a score.

What happened with Epoch AI and OpenAI?

FrontierMath is a benchmark developed by Epoch AI to test advanced mathematical problem-solving by AI systems. OpenAI paid to commission a substantial part of it. Under the arrangement, OpenAI owned the commissioned questions and received access to much of the problem and solution material, subject to holdout exceptions. That made OpenAI both a funder of benchmark content and a company whose models were evaluated on it.

The arrangement drew attention around the launch of OpenAI’s o3 model, when FrontierMath featured in discussion of model performance. Epoch’s initial benchmark announcement, dated November 8, 2024, described the test; the OpenAI relationship became public around December 20. On January 19, 2025, TechCrunch reported criticism from contributors and researchers. Epoch published a detailed clarification on January 23.

Epoch’s initial FrontierMath benchmark page and its January 2025 chronology and reporting on contributor concerns provide context for the timeline.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What did OpenAI fund, own and access?

Epoch said OpenAI commissioned 300 core mathematics questions. OpenAI owned those commissioned questions, and the arrangement gave it access to problem statements and solutions, with designated holdouts excluded. Epoch’s January 2025 account described an original plan for a 50-problem holdout in which OpenAI would receive the statements but not the solutions. That figure describes the arrangement at that point, not every later benchmark version.

Epoch’s later benchmark documentation describes version-specific arrangements. It cites 53 solutions withheld from OpenAI in a version dated February 28, 2025. For a Tier 4 version comprising 50 problems, it says OpenAI had access to 30 while 20 were held out. These are distinct snapshots and should not be combined into one timeless count. Epoch’s clarification of the agreement and its benchmark overview and version details describe the terms.

Epoch also said its agreement required OpenAI’s permission before Epoch could publicly disclose the partnership. Epoch acknowledged that this did not prevent it from telling contributors that an AI company sponsored the work, and said it should have communicated more systematically. The distinction matters: a restriction on public disclosure is not the same as a prohibition on informing contributors.

Why did contributors and observers object?

The criticism was aimed primarily at Epoch’s handling of the arrangement. TechCrunch reported that contributors questioned whether they had been told the named funder, who would own or access their work, and whether OpenAI would have access unavailable to other AI developers. The report included an anonymous contractor’s criticism of the transparency and attributed concerns about contributors’ awareness to mathematician Carina Hong. These are reported accounts; Epoch separately acknowledged that many contributors lacked details and that its communication should have been more systematic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The issue was not simply that a company paid for research. Benchmark creators need resources, and industry funding can support expensive, technically demanding work. The concern was that contributors and readers could not fully assess the implications when a model developer had privileged access to material used to evaluate models across the field.

  • Contributor consent: People creating questions may care who will receive their work and how it might be used.
  • Unequal access: If one AI company knows test questions or answers that competitors do not, its results are harder to compare fairly.
  • Ownership and openness: Ownership can limit whether an evaluator may share questions or solutions with other labs.
  • Public confidence: Even a valid score can be difficult to interpret when the test was not equally blind for all developers.

Did OpenAI cheat or train on FrontierMath?

The cited evidence establishes privileged access and a disclosure failure; it does not establish that OpenAI trained on FrontierMath, fabricated a result or deliberately manipulated an evaluation. Epoch said there was a verbal understanding that OpenAI would not use the benchmark material for model training. A verbal assurance is not the same as a publicly documented technical control or independent audit, so it does not resolve the question of how access was governed.

Access alone is not proof of training contamination. But it creates a material risk: test questions or solutions could enter training data, or a model could be tuned against known questions. Even absent misuse, unequal knowledge can affect comparisons. TechCrunch reported that Epoch’s lead mathematician could not independently vouch for the o3 result at the time, while expressing his personal belief that it was legitimate. That reported assessment is not evidence of wrongdoing, nor does it establish the status of later evaluations.

Epoch’s later position is more qualified than a claim that the benchmark is simply independent or compromised. It says OpenAI was the only AI company with access to some material and that this has diminished confidence in results for OpenAI models. A score from a model without that access may remain informative, but comparisons require the precise benchmark version, access rules and evaluation procedure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What did Epoch acknowledge, and what did it say it would change?

Epoch identified two shortcomings: it had not adequately informed contributors about access to their work, and it had not made transparency a non-negotiable condition of its agreement with OpenAI. Its June 5, 2025, statement acknowledged the mistake and described intended changes. Epoch said it planned to disclose funders and data-access arrangements proactively, inform contributors about industry sponsorship before participation, retain benchmark ownership where possible, and offer more equitable access to future benchmarks—through public release or structured access.

Those statements describe policies Epoch said it intended to adopt; they do not by themselves verify that every policy was implemented in every later project. The broader account is on Epoch’s statement on its approach and future transparency.

How has FrontierMath changed since the controversy?

FrontierMath is versioned, and its contents and holdouts have changed. Epoch’s Tier 4 page reports that a June 12, 2026 update corrected errors in 42% of problems. For the full dataset in that cited version, Epoch reported 338 problems: 295 base problems and 43 Tier 4 problems. Those counts are Epoch-reported metadata for that version, not a fixed total for every release. The page also describes a limited public subset and the benchmark’s correction history.

Because revisions can alter questions, solutions and scoring, a result is interpretable only alongside its version and evaluation conditions. A score from an earlier release should not automatically be compared with one from a later release as though the test were unchanged. See Epoch’s Tier 4 page for version-specific counts and updates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What makes a benchmark credible when an AI company funds it?

Industry funding does not automatically make a benchmark worthless. It does raise the bar for governance. Readers assessing a benchmark should be able to find clear answers to these questions:

  • Were funders and access arrangements disclosed to contributors before they began work?
  • Who owns the questions, solutions, verification tools and derivative material?
  • Do all model developers receive equivalent access, or is a sponsor’s advantage explicitly limited?
  • Are holdouts genuinely unavailable to the sponsor and large or representative enough to support the claims being made?
  • Are restrictions on training contractual and enforceable, technically controlled, independently audited, or only verbal?
  • Can an outside evaluator reproduce the result and publish it, including an unfavorable result?
  • Are versions, scoring rules, corrections and human comparison conditions documented?

Possible safeguards include separate public and private test sets, equal-access licensing, third-party evaluation, independently audited contamination controls, pre-registered scoring rules, and public version histories. None is sufficient on its own: a hidden test, for example, still needs credible administration and a documented method for interpreting scores.

What the FrontierMath dispute means for readers of AI scores

FrontierMath may still provide useful evidence about mathematical capabilities, but the OpenAI relationship changes how its results should be read. Epoch developed and evaluated the benchmark, while OpenAI funded and owned specified content and had access to some material. That is more precise than calling FrontierMath wholly independent, and more measured than saying the benchmark was proven invalid.

When a benchmark score is used to support a claim about a model, check who funded and owned the test, which version was used, what the model developer could access, how holdouts were administered, and whether an independent party could verify the result. The credibility of an evaluation depends not just on hard questions, but on whether its rules let outsiders understand what the score does—and does not—show.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase09 OCT 267 minMortgage Escrow FAQs: Taxes, Insurance, Shortages, and Refunds
  2. The Money DeskBlogTheFinanceBase09 OCT 265 minHow Mortgage Escrow Accounts Work and What Homeowners Pay For
  3. The Money DeskBlogTheFinanceBase09 OCT 265 minHow to Read a Stock Chart, Volume and Market-Cap Data
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.