Meta released Llama 4 Scout and Llama 4 Maverick on April 5, 2025, and users quickly reported results that varied sharply by provider, prompt format and task. Meta said immature deployments and bugs explained much of the variation, while denying that it trained on benchmark test sets. A separate credibility problem came from Meta highlighting an experimental chat version of Maverick for an LMArena result rather than clearly presenting the ordinary public checkpoint. The evidence supports inconsistent early releases and a disclosure problem; it does not prove that Meta trained on test answers or that every poor result was a bug.
What Meta released on April 5, 2025
Meta announced three members of the Llama 4 family, but only two were available to download at launch. Scout and Maverick are mixture-of-experts (MoE) models that accept text and images. Behemoth was described as a larger teacher model still in training, not as a publicly downloadable release.
| Model | Total parameters | Active parameters | Experts | Advertised context | Launch status |
|---|---|---|---|---|---|
| Llama 4 Scout | 109 billion | 17 billion | 16 | 10 million tokens | Released |
| Llama 4 Maverick | 400 billion | 17 billion | 128 | 1 million tokens | Released |
| Llama 4 Behemoth | Not stated | Not stated | Not stated | Not stated | Still training at announcement |
These specifications come from Meta’s launch announcement and the Llama 4 model card. Active parameters are the approximate number used for each token; total parameters still matter for memory, loading and serving. A 10-million-token maximum is an advertised context limit, not proof of uniformly accurate retrieval or reasoning across ten million tokens.
Hugging Face’s launch description says the two models were trained on up to roughly 40 trillion tokens combined. The models use Meta’s custom Llama 4 Community License Agreement, so commercial users should read the license and acceptable-use terms before deployment.
Recommended Free Tools
#1 Best Overall
Why early users saw different “Llama 4s”
The weights appeared through Meta’s downloads, Hugging Face, cloud partners, LMArena and community serving stacks. In an open-weight release, the name alone does not identify a single reproducible system. A provider may change any of the following:
- the exact checkpoint or instruction-tuning state;
- chat template, system prompt or safety layer;
- sampling and decoding settings;
- quantization and inference kernels;
- context truncation and batching behavior;
- image resizing and other multimodal preprocessing;
- tool-use wrappers or even the model selected behind an API label.
Meta executive Ahmad Al-Dahle told VentureBeat that reports differed “across different services” and that implementations needed time to stabilize. That is a plausible mechanism for inconsistent output, but Meta did not publish a complete incident report identifying each bug, affected provider, reproduction case or fix. The company’s explanation should therefore be treated as its account of the rollout, not as a demonstrated root-cause analysis.
Rank #2
What users complained about
Early reports included weak coding on some tests, juvenile or odd conversational tone in certain hosted versions, uncertainty about the long-context claims and confusion between Scout and Maverick on provider platforms. One VentureBeat report cited an early Aider Polyglot result in which Maverick scored 16% on a 225-task coding evaluation. That was one independent, launch-period measurement—not a universal score for every Maverick checkpoint or task.
Some disappointing results may reflect genuine model weaknesses, unsuitable prompts or a different model variant rather than a software defect. Conversely, a successful demonstration on one service cannot establish that every public deployment behaves the same way.
The benchmark controversy had two separate parts
An unverified test-set-training allegation
An online post, whose authenticity was not established, alleged that Meta researchers had been encouraged to incorporate benchmark test sets into post-training or optimize directly for benchmark targets. Meta denied training on test sets. The available reporting provides no independent proof that the released models were trained on benchmark answers, so the allegation should not be presented as a finding.
The experimental Maverick used for LMArena promotion
Meta’s own launch post identified the high LMArena result as coming from an “experimental chat version” of Maverick. Critics argued that this variant was optimized for conversational preference ratings and was not necessarily the same configuration as the downloadable public checkpoint. The disclosure appears in the announcement, so calling it secret would overstate the evidence. The legitimate issue is comparability: a score from a special chat-tuned system can be read as a claim about standard Maverick unless the distinction is made prominent.
That distinction matters even without improper training. Benchmark optimization, a private system prompt or fine-tuning for an arena is not automatically equivalent to training on the arena’s answer key. It can still make a result less representative of the model a developer downloads.
What the evidence establishes
| Claim | Evidence status |
|---|---|
| Some early services produced inconsistent results | Supported by contemporaneous user reports and coverage |
| Bugs and immature integrations caused the variation | Meta’s explanation; independent confirmation was not published in the cited coverage |
| Meta trained on benchmark test sets | Unverified allegation, denied by Meta |
| Meta used an experimental chat variant for the highlighted LMArena result | Disclosed in Meta’s launch announcement |
| The public Maverick checkpoint matched that experimental variant | Not established |
| Llama 4 was universally poor | Not established; results varied by task and deployment |
How to compare Llama 4 results fairly
A score is meaningful only when the tested system is identified precisely. Record these details for every evaluation:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallBest Value
- Checkpoint and variant: specify Scout or Maverick, base or instruct, precision, quantization and revision.
- Serving path: name the provider, inference engine and API version; note any substitutions or system prompts.
- Prompt protocol: preserve the chat template, role messages, tools and multimodal preprocessing.
- Decoding: record temperature, top-p, token limits, seeds and pass@k methodology.
- Task and sample: publish the exact benchmark version, number of examples and date. A 16% result on one 225-task coding test is not a general capability rating.
- Context test: distinguish the advertised maximum from measured retrieval accuracy, latency and memory use at each length.
A high conversational-arena score does not prove superiority in coding, mathematics, factuality or long-context retrieval. A poor coding score does not prove broad failure. The same provenance standard applies to comparisons with DeepSeek V3, Claude, GPT-family, Gemma, Mistral or older Llama models: date the comparison and identify the exact checkpoints.
What developers should check before adopting Llama 4
Match the model to the workload
- Coding and agents: run your own repository and tool-use tests; do not substitute launch leaderboards for task-specific evidence.
- Images: test the actual image sizes, formats and preprocessing used in production.
- Long documents: measure retrieval and answer quality as prompts grow; do not assume the listed context limit guarantees reliable reasoning.
- General chat: compare tone, refusal behavior and instruction following on the exact endpoint you will operate.
Account for infrastructure and licensing
Meta positioned Scout as able to fit on one NVIDIA H100 with Int4 quantization. Maverick’s roughly 400-billion total parameters require a substantially larger serving footprint even though only about 17 billion are active per token. Total memory, networking, quantization quality, monitoring and failover determine cost—not active parameters alone.
Teams can obtain official artifacts through Meta’s Llama resources or the Hugging Face Meta collection, including the Scout Instruct and Maverick Instruct pages. Self-hosting offers control and data governance but shifts GPU operations and reliability work to the buyer. A managed endpoint trades weight-level control for simpler operations. No current prices are established by the cited sources, so compare live quotes and total cost of ownership rather than relying on launch claims.
Bottom line on the Llama 4 launch
Llama 4’s April 2025 controversy was not one proven scandal. It combined real variation among early deployments, a bug explanation that Meta did not fully document, and a benchmark disclosure that made the highlighted LMArena score hard to compare with the public checkpoint. The careful conclusion is that release clarity and evaluation provenance failed before the evidence justified either “Meta cheated” or “Llama 4 was simply bad.”
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




