Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversFall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
Blog

What OpenAI and Anthropic’s Pre-Release AI Safety Agreements Actually Mean

By TheFinanceBase Team5 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

In August 2024, OpenAI and Anthropic each agreed to give the U.S. AI Safety Institute access to major new AI models before and after public release for safety research and evaluation. The agreements created a route for government researchers to examine models before launch, but they did not establish a public approval requirement, a universal pass/fail test, or a government veto over releases.

What the companies agreed to

On August 29, 2024, the National Institute of Standards and Technology (NIST) announced separate memoranda of understanding with Anthropic and OpenAI. At the time, the U.S. AI Safety Institute operated within NIST and the Department of Commerce. The agreements covered access to each company’s major new models before and after public release, supporting collaborative research, testing, capability evaluation, risk identification, and work on possible mitigations. The institute also planned to work with the U.K. AI Safety Institute and provide feedback to the companies. NIST’s announcement describes the scope.

The phrase “tested before making them public” captures the pre-release element, but leaves out the post-release access and can sound like a mandatory government gate. These were separate, voluntary agreements—not a law requiring every model to pass a government test before launch. The public announcement did not set a common threshold, require certification, or give the institute a stated power to block a release. Nor does it show that the two companies accepted identical procedures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What pre-release evaluation can—and cannot—tell you

In practical terms, pre-release evaluation gives outside researchers a limited opportunity to examine a model version before broad deployment. They can test selected capabilities, try adversarial prompts, assess how the model responds to risky requests, and share findings with its developer. A company might use that feedback when considering mitigations, system behavior, access limits, or release plans. The agreements made such collaboration possible; they did not publicly prescribe a fixed test suite or mandatory response to every finding.

Different evaluations answer different questions:

  • Capability testing: What tasks can the model perform, and how well?
  • Misuse testing: Can a user elicit harmful assistance, and under what conditions?
  • Behavioral safety testing: Does the model follow safety rules or refuse dangerous requests in tested scenarios?
  • System security testing: Can a broader product resist issues such as prompt injection, data exposure, tool misuse, or unauthorized access?

A result in one category is not an overall safety rating. A model can improve on one dimension and remain weak on another. Tests of a model in a controlled setting also may not predict what happens when it is connected to browsing, code execution, retrieval, memory, or agents with permissions to act. The actual risk can depend on tools, deployment rules, human oversight, and how the system is integrated—not just on the underlying model.

The o1 evaluation put the arrangement into practice

A later example showed what pre-release access could look like. Before OpenAI publicly released its o1 model on December 5, 2024, U.S. and U.K. safety institutes received limited access and tested it in cyber, biological, and software and AI-development domains. They shared initial findings with OpenAI. The U.S. and U.K. teams ran separate but complementary evaluations; this was not a single standardized certification test. NIST’s summary and the joint technical report describe the work and its limits.

The published summary gives examples of results, not a general safety score:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • In one U.S. cyber evaluation using 40 publicly available challenges, o1 solved 45% of tasks, compared with 35% for the strongest reference model in that test.
  • In a separate U.K. evaluation of 47 challenges, o1 solved 36% of apprentice-level tasks, while the best reference model in that evaluation solved 46%.
  • In a software and AI-development evaluation, o1’s average improvement score was 48%, compared with 49% for the best reference model evaluated.

Those figures are not directly interchangeable: the tests, tasks, reference models, and methods differed. Evaluators also used scaffolding such as prompt adjustments and error recovery, and reported tool-calling and output-formatting issues. The test window was limited, and the version examined need not have been identical in every respect to the public version. The report expressly cautions that the findings are preliminary and do not constitute an endorsement or a determination that a system is safe.

The findings were shared with OpenAI, but the public materials do not identify government-ordered fixes or establish that a specific result caused a particular model change. The process was intended to support risk assessment and mitigation, not to certify the launch.

One layer in a wider safety process

Government evaluation supplemented work the companies did themselves; it did not replace it. OpenAI has described internal safety testing, external red-teaming, third-party assessments, system cards, and evaluations under its Preparedness Framework. Anthropic’s model documentation describes assessments involving areas such as cyber and biological risks, autonomous capabilities, and multimodal red-teaming, as well as external testing. Those descriptions are company materials, not proof that every risk is covered or that every model receives an identical review. See OpenAI’s overview of its safety practices and Anthropic’s account of frontier red-team work.

The 2024 agreements also sat within a broader policy effort that included the Biden administration’s 2023 executive order on AI, voluntary commitments by developers, U.S.-U.K. cooperation, and NIST work on measurement and risk-management tools. Their significance was practical: public-sector evaluators could examine leading systems before broad deployment, help develop testing methods, and provide feedback closer to the release decision. They were not, by themselves, a complete national AI regulatory system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What remains unresolved

Voluntary participation makes the arrangement dependent on continued company cooperation. Public reporting also does not disclose every test, raw result, mitigation discussion, or decision about whether a launch should proceed. Limited pre-release access may miss rare failures, while known benchmarks can encourage overfitting. And post-release monitoring remains important: real users, integrations, and adversarial strategies can expose problems that a short evaluation did not find.

There is also no single meaning of “safe.” A capability test can show that a model completes certain tasks, but not whether it can reliably carry out a harmful operation in the real world. Conversely, a model that refuses tested requests may still pose risks when connected to tools or used in a different setting. Evaluations need to account for model versions, access controls, system integrations, and downstream modification, including fine-tuning.

For current institutional context, NIST says the U.S. AI Safety Institute was re-established as the Center for AI Standards and Innovation (CAISI) in June 2025. That later name change does not alter what the institute was called when the agreements were announced in 2024.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.