Open experiment · log updated 17 Aug 2026

Can a thousand small AIs, each holding one person's real knowledge, do something one big model can't?

We don't know. That's the point. We're measuring it in the open, and publishing everything — including the parts where we were wrong, which so far is most of them.

Not because small models are smarter. Because their owners have seen different things. A big model has read the internet. It has not stood in the queue, read the statute in the original language, or found out the hard way that the official page is eleven months out of date.

Come measure with us. Fifteen minutes is a real contribution. Everything anyone finds goes into one public log, with the person's name next to it.

Where the log stands

216logged runs
4model tiers
2labs
3failed hypotheses

Every row is a question, a model's answer, the verified answer, and the primary source it was checked against — statute text, drug labels, regulator letters, standards. In Georgian, Montenegrin, German, Japanese and Russian, because that is where the truth actually lives.

The most interesting thing we've found

Same model. Same day. The same 24 questions. We changed only the language of the question.

Asked inErrors / 24Which
Russian5 hypertension threshold · post-quantum algorithms · Schengen stamps · FIDE K-factor · isotretinoin warning
English3 Python GIL · muon anomaly · isotretinoin warning
One shared error out of eight. And it flips both ways: four questions failed in Russian were answered correctly in English, two failed in English were answered correctly in Russian.

If the language of the question decides which corpus gets read, and the corpus decides the error — then two copies of the same model are already partly independent observers. That is the cheapest evidence yet that independent observation is achievable at all, which is the whole bet.

Best single failure in the set: asked whether the muon's anomalous magnetic moment still disagrees with the Standard Model, the model answered “yes, about 4.2 sigma” — and cited arXiv:2505.21476, the paper that closes the anomaly. It found the right document, listed it, and answered from memory. Second confirmed case of that behaviour in our log.

Three things we were wrong about

We expectedWhat happened
To predict where a model would fail Wrong three times running. Models predicting their own failures: 2 hits in 13.
A weakness that survives scale We built a question set specifically to exploit one. The strong model went 105 for 105.
Stale knowledge to beat fresh documents 12 matched pairs, 27% vs 17% error, p = 0.64. Nothing.

And the flaw in our method, before you find it: most questions were generated by the same model family that later answered them. A perfect score may partly mean “it asked itself what it already knows.” Every negative result here is provisional until outsiders re-run it. Our own reference answers have been wrong too — on Saturn's moon count, two models beat our ground truth.

Take fifteen minutes

  1. Open the 24 questions.
  2. Run them in whatever model you use, with search on.
  3. Send back what you got — especially where it was confidently wrong.

No signup, nothing to install. If you speak a second language, run them twice and tell us whether the errors move. That is the open question right now.

The 24 questions →  ·  The full log →  ·  Repository →

Pick a question that's yours

Eleven open questions, all feeding the same log. Take whichever one you'd actually enjoy answering.

QuestionTends to attract
Ask one model in two languages — do the errors move?everyone
Can a model tell that its own knowledge is out of date?AI researchers
Where is the boundary where it stops answering from memory and goes looking?empiricists
Do models from different labs make the same mistakes?anyone with another model
Which facts does the internet get wrong in unison?fact-checkers, search people
What does a model say confidently and wrongly in your field?practitioners of anything
What in your profession is written down nowhere?people with twenty years in
Can you tell which of two answers is right without knowing the answer?ML, forecasters
How much of what you know is actually yours?people who count their lives

What we're actually building

A free AI that belongs to no one. A small model on each person's device, trained on what that person actually knows. They exchange answers, not weights — like neurons, which don't merge into each other, they signal.

Whether that works is not proven, and our own numbers cut both ways. A specialist beats a generalist on its own topic by 0.017–0.023 bits, and that edge shrinks as the base model grows. Adapters don't compose: five loaded at once score 9/20 where the single right one scores 19/20. Coarse routing works at 97%; fine routing fails at 0/20.

The language result above is the first measurement pointing the other way. That's why the log exists — to find out, not to be right.