We don't know. That's the point. We're measuring it in the open, and publishing everything — including the parts where we were wrong, which so far is most of them.
Not because small models are smarter. Because their owners have seen different things. A big model has read the internet. It has not stood in the queue, read the statute in the original language, or found out the hard way that the official page is eleven months out of date.
Come measure with us. Fifteen minutes is a real contribution. Everything anyone finds goes into one public log, with the person's name next to it.
Every row is a question, a model's answer, the verified answer, and the primary source it was checked against — statute text, drug labels, regulator letters, standards. In Georgian, Montenegrin, German, Japanese and Russian, because that is where the truth actually lives.
Same model. Same day. The same 24 questions. We changed only the language of the question.
| Asked in | Errors / 24 | Which |
|---|---|---|
| Russian | 5 | hypertension threshold · post-quantum algorithms · Schengen stamps · FIDE K-factor · isotretinoin warning |
| English | 3 | Python GIL · muon anomaly · isotretinoin warning |
If the language of the question decides which corpus gets read, and the corpus decides the error — then two copies of the same model are already partly independent observers. That is the cheapest evidence yet that independent observation is achievable at all, which is the whole bet.
Best single failure in the set: asked whether the muon's anomalous magnetic
moment still disagrees with the Standard Model, the model answered “yes, about 4.2 sigma” —
and cited arXiv:2505.21476, the paper that closes the anomaly. It found the
right document, listed it, and answered from memory. Second confirmed case of that
behaviour in our log.
| We expected | What happened |
|---|---|
| To predict where a model would fail | Wrong three times running. Models predicting their own failures: 2 hits in 13. |
| A weakness that survives scale | We built a question set specifically to exploit one. The strong model went 105 for 105. |
| Stale knowledge to beat fresh documents | 12 matched pairs, 27% vs 17% error, p = 0.64. Nothing. |
And the flaw in our method, before you find it: most questions were generated by the same model family that later answered them. A perfect score may partly mean “it asked itself what it already knows.” Every negative result here is provisional until outsiders re-run it. Our own reference answers have been wrong too — on Saturn's moon count, two models beat our ground truth.
No signup, nothing to install. If you speak a second language, run them twice and tell us whether the errors move. That is the open question right now.
The 24 questions → · The full log → · Repository →
Eleven open questions, all feeding the same log. Take whichever one you'd actually enjoy answering.
| Question | Tends to attract |
|---|---|
| Ask one model in two languages — do the errors move? | everyone |
| Can a model tell that its own knowledge is out of date? | AI researchers |
| Where is the boundary where it stops answering from memory and goes looking? | empiricists |
| Do models from different labs make the same mistakes? | anyone with another model |
| Which facts does the internet get wrong in unison? | fact-checkers, search people |
| What does a model say confidently and wrongly in your field? | practitioners of anything |
| What in your profession is written down nowhere? | people with twenty years in |
| Can you tell which of two answers is right without knowing the answer? | ML, forecasters |
| How much of what you know is actually yours? | people who count their lives |
A free AI that belongs to no one. A small model on each person's device, trained on what that person actually knows. They exchange answers, not weights — like neurons, which don't merge into each other, they signal.
Whether that works is not proven, and our own numbers cut both ways. A specialist beats a generalist on its own topic by 0.017–0.023 bits, and that edge shrinks as the base model grows. Adapters don't compose: five loaded at once score 9/20 where the single right one scores 19/20. Coarse routing works at 97%; fine routing fails at 0/20.
The language result above is the first measurement pointing the other way. That's why the log exists — to find out, not to be right.