It’s tempting to use an LLM as a fact-checker: paste in a claim, ask “is this true?”, get a confident, well-structured verdict. The trap is that the model has no opinion of its own the way a human expert does. Change the framing or the surrounding context and it will instantly “change shoes in mid-air”, producing a different verdict on what is essentially the same question.
That isn’t a bug someone will patch out — it follows directly from how these models learn and how they process a conversation. To get genuinely useful validation out of an LLM, you need to understand two mechanisms: frequency bias, inherited from training, and context stickiness, which turns a chat into a stubborn defender of its own early mistakes. Let’s take them in turn — and end with a practical protocol.
The parrot problem: frequency instead of conviction
The most common misconception about AI is anthropocentric: people assume the model states something out of expert conviction. But look at how facts actually get into an LLM. Pretraining and supervised fine-tuning are closer to teaching a parrot — the model is trained to reproduce its datasets nearly verbatim, token by token. There is no “expertise” step. When you later ask a question, the answer is an interpolation across the cases it absorbed — a draw from the Bayesian variety of its training data.
Here’s the catch: interpolation gravitates toward frequent cases, not correct ones. If the training data repeats some claim thousands of times — “model X is best at coding”, “lab Y is ahead of everyone” — that cloud of examples becomes a powerful attractor in the model’s semantic space, a kind of semantic black hole. Answers drift toward it because of frequency, not because of logical merit. The model’s “native position” is just the frequency structure of its dataset.
The stubborn-donkey effect: why a chat digs in
The second mechanism lives inside the conversation itself. Everything already said in a session sits frozen in the model’s KV cache — the attention memory over the dialog — as fixed semantic vectors. Attention keeps being pulled back toward those early tokens (the attention-sink effect), so whatever was said first becomes the center of semantic gravity for everything that follows.
Now suppose the model made a shaky argument early in the chat. Modern RL training does teach models to reconsider — but the old, mistaken vectors don’t disappear from the context. They keep exerting pull, and the model tends to orbit its first hypothesis, defending it like a stubborn donkey even after you’ve pointed out the flaw.
The practical rule follows directly: for questions that matter, don’t keep arguing inside one long chat. Open a fresh session, and vary the initial framing — the in-context examples and angle you start from (ICL). A clean context is the only real reset button.
Where RL reasoning doesn’t save you
All frontier labs train models to reason step by step, with process-reward-style feedback on the reasoning itself. That genuinely helps — but in practice it cannot out-vote the frequency priors baked in during pretraining. A logically clean chain of reasoning can still start from a frequency-biased premise and land, very persuasively, in the wrong place.
Without good in-context framing the model defaults to playing Captain Obvious. For questions with one settled answer, that’s fine. But on genuinely contested expert questions — the kind where even scientists split into big-endians and little-endians — a model without framing is helpless: it needs you to tell it which camp of hypotheses to weigh, what evidence to prioritize, what counterarguments to take seriously. Absent that, its pick between camps can be close to random — dressed up, as always, in confident prose.
A practical protocol for validating facts with AI
Start from the default assumption: the model carries frequency bias from training and gets stubborn within a session. Neither is visible in the tone of the answer — the prose sounds equally authoritative either way.
Ask the same question several times from genuinely different angles: fresh session each time, different starting context, different framing of what’s at stake. If the verdict flips with the framing, you’ve learned something important — the model has no stable position on this question, and its “confidence” was an artifact of your prompt.
Don’t let the model collapse into a single hypothesis. Explicitly ask for alternative interpretations and the strongest case against its own conclusion. Keep several competing readings on the table — a superposition of hypotheses — and compare them, instead of accepting the first confident synthesis.
And never just dump raw context on a model “for analysis” without framing. That’s the classic way to get fooled: the output will look like science, read like expertise, and inherit every bias above. The model supplies breadth and speed; the verdict stays yours.
Ensembles: the strongest tool — and where Kumo helps
The single most effective technique on this list is an ensemble of different models. Claude, GPT, Gemini, DeepSeek and Kimi were trained on different data with different priorities — which means their frequency biases differ too. Where they independently agree, confidence is warranted; where they diverge, you’ve found exactly the spot that deserves your own judgment.
Ensembles are only practical when switching models is cheap. That’s the case for Kumo: one OpenAI-compatible base URL and one prepaid balance in front of the frontier catalog, so running the same question through three model families is a parameter change, not three accounts. You pay per token, see each model’s rate before you spend, and the model is never silently substituted — which matters when the whole point is comparing who said what.