A distance, then a test for luck.
Two batches of answers never match exactly, even from the same model. So the real question is not "are they different" but "are they more different than chance allows".
For each question, the answers from A and B are compared with the Jensen–Shannon divergence. It reads 0 when both sides give the same answers in the same proportions, and 1 when they share nothing. The score of a run is the average over the questions.
Tellstone throws all the answers to a question into one pile, deals them back out to A and B at random, and scores that. It does this 2,000 times.
The p-value is the share of random deals that look at least as far apart as the real one. A small p-value means the real split is not something chance produces.
If fewer than one random deal in a hundred looks that far apart (p under —), the verdict is inconsistent. Otherwise it is consistent.
A question only counts if each side gave at least 5 answers. With fewer than 4 such questions, the verdict is inconclusive.
One check, start to finish walks through these steps with real answers.
- Next
- Reading a result · what each line is telling you
- Before this
- The eight questions
- All pages
- The docs