What each line is telling you.
Read the stamp first. Then read the rest, because the stamp is only as good as what sits under it.
§1The three verdicts
- Consistent
- Same tells. Nothing separates the two sides at this sample size. It is not a certificate; a bigger sample can still find a gap.
- Inconsistent
- Different tells. The two sides do not answer like the same model, and chance does not explain it.
- Inconclusive
- Too many requests failed to call it either way.
§2The lines below it
- Answers
- The three favourite answers on each side, per question, with the distance for that question. This is where you see the difference with your own eyes.
- Prompt token counts
- The same model counts the same prompt the same way. If one side is always higher by the same amount, something is being added to your prompt. If the gap jumps around, the two sides count tokens differently, which means different model families.
- Model name returned
- What each side says about itself. Shown as given. It proves nothing; it is the claim being tested.
- Upstream provider
- Reported by gateways that pass your request on to others, such as OpenRouter.
- Median latency
- For context only. Speed depends on far more than the model.
→Keep reading
- Next
- Where it stops · the cases it will let you down
- Before this
- How it decides
- All pages
- The docs