Models are bad at random, each in its own way.
That is the whole trick. The rest of Tellstone is counting.
Ask a person to say a random number and you will hear seven far too often. People are bad at random, and so are language models. One model names a pangolin most of the time. Another says giraffe. A third says otter almost every time.
Nobody designed these habits. They are residue from training, and two models trained differently carry different residue. That makes them a tell: easy to read, and hard to fake without being the real model.
It does not try to judge how clever an endpoint is. It asks idle questions on both sides and checks whether the habits line up.
It ships with no stored fingerprints. The comparison is always between two live endpoints that you pick, so there is nothing of ours to trust or to go stale.
The habits themselves are on show in the atlas.
- Next
- The eight questions · exactly what gets sent
- Before this
- Two-minute start
- All pages
- The docs