Each figure on this site says whose it is (Anthropic's, Artificial Analysis's, a named person's, or ours) and links to where it came from, with the time we read it. When we tested something ourselves, the figure says so and gives the number of runs, the setting each run used and the date.
Every video page has a numbers table with all of this, and the same data as JSON and CSV.
Our launch-day videos mostly report other people's measurements. When that is all we have, the video and its page say "Not measured yet". A number of ours only appears once we have run the test.
Some figures change after a video is made, for example when Artificial Analysis re-runs a model on a fixed build. Those carry a ⟳ mark. When we read one again and it has moved, the page shows the value from the video next to the new one, both dated. We don't overwrite what the video said.
Our own tests run on subscription plans, through Claude Code and Codex, at each tool's default settings unless the page says otherwise. Where we give a dollar cost for our runs, it is the notional cost at API list prices, not money we were billed.
Yellow marks the number the sentence is about. It never means better or cheaper. Grey marks the other data beside it, and it never means worse. BB, the lilac bar, is not data unless he is labelled as a bar in a chart.
We list every change after publication on the corrections page, newest first, and say what kind of change it was.