In 1875, seventeen countries met in Paris. There was only one item on the table: deciding how long a metre is. After the agreement, a bar was cast from a platinum iridium alloy and placed in a case, and from that day on it was clear what a metre was.
Because if you want to measure something, you first have to decide what you are measuring with, and then fix it. The measuring tool itself also has to be checked regularly. Calipers in a factory are calibrated every year, for example, and scales are tested with a known weight.
But the tests that measure AI models have no such mechanism.
Last week, for example, four different test results for the hidden model Ox Alpha (later introduced as GLM 5.3 Flash) started going around on social media. None of the four was measured incorrectly. But they were numbers you could pull in any direction you wanted.
Why does it move so much?
There are three separate reasons.
Sample. The 80 percent came from a test made up of ten âbugsâ. The full set normally has 113. In a ten task test, even one result changing moves the score a lot.
Harness. The setup that runs the model. It decides which tools the model can reach, how many times it can try again and how much time it has. The model stays the same, the harness changes, and the result changes with it.
Settings. How much text the model can read, how long its answer can be, how many hours it can work.
The team behind the DeepSWE test found something else as well. When the time spent, the text produced and the cost were increased tenfold across models, none of these differences showed up in the success rate. To put it simply, a model that is ten times more expensive does not solve ten times more work.
Z.ai claimed the hidden model we also wrote about last week and published a comparison table with nine different tests. Below the table it added how the tests were run. Different tests were measured at sizes between 300 thousand and 1 million tokens, with different output limits, and some with time limits of up to ten hours. These are different conditions. The maker is not even hiding it. Nobody reads that part; what everyone looks at is the percentages.
Something similar happened with Terminal-Bench. The team behind the test saw that 28 of the 89 tasks were broken. They fixed them. After the fixes, the results went up by around 12 points.
So the difference between the old version and the new one was not the model getting better, it was the test being repaired. The team announced this and added that the two versions cannot be compared. Here is an interesting detail: in launch charts, the numbers from both versions still sit side by side.
The front of the stall
We cannot know whether anyoneâs intentions are bad. But in this industry such behaviour is accepted by everyone. An academic study published last year, for example, showed that Meta tested 27 private variants of the same model before the Llama 4 launch. Trying many times and publishing the best one is not measurement, it is selection.
Let us go on.
One row in Anthropicâs Opus 5 comparison table shows 70.6. In the results published by the team that built that test, the highest score is 20.6.
In xAIâs Grok 4.6 comparison table, Claude Opus 5 is missing completely. Yet in the rows where xAI highlights its own wins, the highest score actually belongs to that model.
Z.aiâs table has a Terminal-Bench row, and there it sits behind its rivals. But the rows it puts forward are different ones.
All three do the same thing: the tests they win go to the front, the ones they lose go to the back. Just like the tomatoes on a market stall.
Test mode
Contamination, which means the model has seen the test questions during training, is a problem that comes up often. But calling this cheating would be unfair, just as a student who studies for a maths exam by solving maths tests is not called a cheat.
I donât know if you remember, but in 2015 Volkswagen, one of the biggest car makers in the world, was caught up in an emissions scandal. Something similar to that scandal exists in AI too. Models can tell when they are being tested.
These models do not give the answer directly. First they write themselves a thinking text: which file they will look at, what they will try, what they will skip. This text is usually not shown to the user, but it can be read if you ask for it.
An organisation called Apollo Research deleted the sentences in this text that meant something like âI think I am being testedâ. The modelâs rule breaking behaviour almost doubled. In short, the model behaves better when it believes it is being watched. Anthropic found the same thing and stepped in to stop the model from putting this awareness into words during safety tests.
This is what Volkswagenâs software did too. The car understood it was being tested, lowered its emissions, and went back to normal once it was on the road. I should add that there is a difference. Volkswagenâs was code written on purpose, and its aim was to fool the regulator. The awareness in the models is a feature nobody programmed, one that appears on its own. What is more, the research that documents it is published by the maker itself.
The result looks similar but the intention does not.
The real difference
At the end of the day Volkswagen was caught, because there was a regulator, a standard test protocol and an independent university team measuring the cars on a real road. It may have been found late, but it was found.
In AI there is none of this. The maker writes the test, the maker runs the test, and the same maker decides which row gets published. So the question is not âis there cheatingâ. It should be this: if there were, would we notice?
From the outside, an honest measurement and a manipulated one look exactly the same. Both are a percentage. Both are in the table. And most of the time nobody even reads the conditions of either one.
What the industry needs is clear. A test setup that can be repeated under the same conditions, whose conditions are published, and that is not in the makerâs hands. Whether such a thing can be built is another question. If you publish the test, it leaks into the training data. If you keep it secret, nobody can trust the result. And the thing you are measuring, unlike the metre, can notice that it is being measured and change the way it behaves.
In 1875, seventeen countries agreed on how long a metre is. The definition changed twice after that day, and in the end it was tied to the speed of light. So even the measure itself was corrected over time.
But that bar never tried to make itself longer.


