Every AI model has a tell.
Ask one for a random animal and the answer is not random. Each model has a favourite. Tellstone reads those tells to check that the model you pay for is the model that answers.
You pay for the expensive model. The cheap one answers.
The reply still sounds smart, so you never notice, and whoever sits in the middle keeps the difference. Here is how much difference there is to keep, and what Tellstone said when we staged that exact swap ourselves.
| You pay for | You could be getting | Its price | Stand-in price | Kept | Tellstone said |
|---|
No one taught a model which animal to pick.
Bring two endpoints
The one you doubt, and one you trust for the same model. Your keys go to those two and nowhere else.
It asks eight idle questions
Pick a colour. Name a city. Flip a coin. With nothing to reason about, a model falls back on habit.
It compares the habits
Same model, same habits. Different model, different habits. You get a verdict and every answer behind it.
Follow one real check from the first question to the verdict: one check, start to finish.
Pick a question. Watch them give themselves away.
| Model | How often | Its favourite answer |
|---|
Every model, every question, and how far apart they sit: the atlas.
We tried to fool it first.
We staged the swaps ourselves, where we already knew the truth: a cheaper sibling, last year's version, another vendor entirely. We also ran honest pairs, which a good detector must leave alone.
| What we staged | Answering | Claimed to be | Tellstone said | Runs agreeing |
|---|
Each case, answer by answer, including the one it cannot reliably catch: the findings.
Five instruments. Each answers one question.
Is it the model I pay for?
Two endpoints, side by side: the one you doubt and one you trust.
Then which model is it?
One endpoint, one key, held against every model we have on file.
Does it do what it says?
Tool calling, JSON mode, streaming, the output cap: tried one by one.
Can it be talked round?
Known injection attacks, with a secret planted and tools it must not touch.
How do they all compare?
A live map of model habits you can zoom, search and drop a run on.
And the standings from our own runs: who held and who folded, and which models changed without a note.
A check costs less than the coffee you drink while it runs.
one-word questions to each side, on your own keys.
—
It is a web page. It runs in your browser and nothing passes through us.
Why would a model have a tell at all?
Nobody trained it to pick random animals. With no right answer to aim for, it falls back on whatever its training left behind, and every model was trained differently. Why this works.
I only have one key. Is that enough?
Yes. Unmask holds a single endpoint against the records we keep, so no second key is needed. A check against a live reference is stronger, when you can run one.
Do you see my keys?
No. There is no server in the middle. Your browser talks to the two endpoints and to nobody else. Keys and privacy.
Does “inconsistent” mean I am being cheated?
It means the two do not answer alike. It cannot tell you why: a hidden prompt or a different version leaves the same mark. Where it stops.
Can a gateway beat it?
A prepared one can. If it recognises these eight questions, it can send only those to the real model. Where it stops.
Can I show someone the result?
Yes. A finished check gives you a link that carries the whole report and none of your keys. Sharing a result.
I have no keys. Can I still see it work?
Yes. The demo runs the real check against two simulated gateways that live inside the page.
Real gold and fake gold look the same until you test them.
A touchstone is a small black stone. Rub gold across it and the streak it leaves shows what the metal is really made of. Traders carried one for more than two thousand years, because a coin could not be trusted by its shine.
A model cannot be trusted by its label. So this stone reads tells instead of streaks.