All writing

AI Security Jul 2026 · 8 min read

Benchmark Theater: Kimi K3 and the 51% Number They Left Out

A spotlit golden trophy on a theatre stage while a cracked AI model core hides behind the curtain

In July 2026, Moonshot AI released Kimi K3, the largest open-weight model ever shipped, and its benchmark charts showed it beating American frontier models on coding and agentic tasks. Independent testing then found a hallucination rate around 51 percent that those charts never mentioned. After seventeen years in enterprise security, two hundred plus audits, and more ISO 27001 programmes than I care to count, I know this pattern intimately. It is compliance theater with a leaderboard, and the way you defend against it is the same way a good auditor works: never accept the assessed party’s own evidence as the whole story.

This is not really a post about one Chinese model. Kimi K3 just happens to be the cleanest specimen of a failure mode the whole AI industry is now running at scale.

What happened, and what was left out

The release itself was genuinely impressive: 2.8 trillion parameters, a million-token context window, open weights anyone can download. It topped several public leaderboards, including a frontend-coding arena where it beat the best closed US models, and it landed in the middle of a geopolitical firestorm, with the White House alleging the model was built by distilling a US competitor and Beijing drafting export controls in response.

Then the independent numbers arrived. Testing found the model hallucinating in roughly half of the measured cases, a figure absent from Moonshot’s published benchmark charts, and the UK’s Cyber Institute found its cyber and mathematics capabilities lag far behind the closed frontier despite the headline wins. Nobody has accused Moonshot of fabricating a score. They did not need to. They simply published the benchmarks the model wins and omitted the measurement it fails.

Selective disclosure is not a Chinese invention, and it would be naive to read this as one lab’s sin. Every vendor’s model card is a curated document. Kimi K3 matters because the gap between the curated story and the measured reality was quantified within days, in public, by third parties. Most of the time, nobody checks.

I have seen this movie in two hundred audit rooms

Here is what seventeen years of compliance work teaches you: organisations do not optimise for being secure. They optimise for the measurement of being secure. Given a checklist, people build to the checklist. Given an audit date, controls work beautifully on the audit date. The certificate on the wall is real; the security it implies is a separate question that the certificate was never designed to answer.

I wrote about the operational version of this in security controls that fail silently: the backup that has succeeded every night for three years and has never once been restored from, the alert rule that fires into a channel nobody reads. The control passes its check. The check has quietly stopped measuring the thing you care about. Goodhart’s law is the academic name, “when a measure becomes a target, it ceases to be a good measure,” but auditors knew it long before economists named it.

AI benchmarks are now the fastest-moving instance of this law I have ever watched. Training on the test set has a polite industry name, contamination. Cherry-picking evaluation suites is called a model card. Tuning a release for the leaderboard while the deployed variant behaves differently is called optimisation. Every one of these has an exact analogue in compliance, and every one of them would get flagged in a half-decent audit.

Benchmark theater and compliance theater, side by side

Compliance theaterBenchmark theater
The artefactCertificate, audit reportLeaderboard rank, model card
Optimised forThe auditor’s checklistThe benchmark suite
What gets hiddenCompensating-control gaps, scope exclusionsHallucination rates, contaminated training data
The tellControls that only work on audit dayScores that only hold on the published suite
Who finds outIncident responders, eventuallyYour users, in production
The fixIndependent testing beyond the checklistIndependent evals on your own workload

The last row is the entire lesson. In both worlds, the assessed party’s self-reported evidence is a starting point, never a conclusion. The UK Cyber Institute’s independent measurement did more for the truth about Kimi K3 in one report than every leaderboard it topped. That is what independent assessment is for, and it is exactly the discipline I argued for in moving from CVSS to attack-success-rate when measuring AI risk: measure the behaviour you will actually live with, not the score someone else chose to publish.

What benchmarks structurally cannot tell you

Even honest benchmarks, run cleanly, answer a narrow question: how does this model perform on this fixed task set, at this moment, under these settings. Three gaps matter for anyone deploying AI in production:

Evaluate models the way an auditor evaluates controls

The transferable method from two hundred audits fits in five moves:

The uncomfortable close

The AI industry is currently repeating, at ten times the speed, the fifteen-year journey the security industry took from “we passed the audit” to “we test our own controls because the audit is the floor, not the ceiling.” Some organisations learned that second posture after one incident. Others are still learning it. The ones deploying models on the strength of leaderboard screenshots are queueing up for the same education, and as this year’s autonomous-attack incidents keep demonstrating, the tuition fees are rising.

Kimi K3’s missing 51 percent was not hidden by lying. It was hidden by omission, which is how almost everything in my field gets hidden. The defence has not changed in seventeen years: independent measurement, on your own ground, on a schedule, with the assumption that every self-reported number is marketing until verified. The lesson generalises well beyond one model, the way most compliance lessons do. Benchmarks are useful. Benchmarks from the party being benchmarked are theatre until a third party has been backstage.


If you are choosing or deploying AI models on benchmark evidence and want an independent evaluation of how they behave on your actual workload, request a review. I run AI security and evaluation engagements anchored in 17+ years of enterprise cybersecurity.

Betting production on a leaderboard screenshot?

Independent evals on your workload, before your users run them for you.

Request an AI evaluation review