Language models are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. Fine-tuning an LM on the recommendations of an ideal Bayesian assistant (BayesAssist in the paper, "the assistant" below) gives it near-Bayesian behavior, while standard fine-tuning on the true answers falls short. But behavior alone does not tell us why. Does the Bayes-trained LM hold Bayesian beliefs, act on them, and turn them into a choice the way Bayes' rule does?
We compare three LMs from Qiu et al. (2026) on a flight recommendation task: the untuned StartLM (Gemma-2 9B), the BayesLM, tuned on the Bayesian assistant's recommendations, and the OracleLM, tuned on the user's true choice. We follow one held-out conversation throughout this page.
The OracleLM reaches 69% and the StartLM stays near chance (33%). On correct answers, the BayesLM's policy also diverges less from the assistant's policy. We evaluate the LMs on four test sets: the original held-out conversations, a shuffled version with the evidence reordered, and two noisy versions where one or three of the user's replies are random. Neither tuned model is affected by the order of the evidence, and both lose accuracy only under high noise, as the assistant predicts.
When the BayesLM is wrong, it puts on average 51% of its probability on the assistant's second choice and 17% on its third; the OracleLM spreads its errors more (48% and 26%). The BayesLM's uncertainty is also better correlated with the assistant's uncertainty (Pearson r = 0.61, vs 0.30 for the OracleLM).
Linear probes recover the assistant's prior (its belief before the last round), likelihood update and posterior from layers 15–25 on, better for the BayesLM than for the OracleLM, while StartLM probes do no better than predicting a uniform belief. A probe for the choice policy picks the assistant's recommendation 87% of the time in the BayesLM.
We inject a prototype prior representation (the average hidden state of 500 training conversations with a strong preference, e.g. for cheap flights) into eight layers at the feedback token and the next 32 tokens. The recommendation should change after the intervention exactly when the assistant's recommendation would change, and stay the same otherwise. For the BayesLM it does at layers 18–25, with a balanced accuracy of 0.77 (the average of accuracy on cases that should change and cases that should not; 0.5 means the edit does nothing useful). A random vector of the same norm barely changes the answer, and the OracleLM stays at 0.52. No single layer or token is enough. We also explore the computational route leading to the beliefs: attention masks that allow only a round-by-round (sequential) or only an all-at-once (one-swoop) route both cost accuracy (0.84 → 0.56 and 0.68). The smaller drop for all-at-once-like computations suggests the LM gathers evidence across the conversation rather than keeping a running posterior.
Patching the BayesLM's beliefs into the OracleLM raises its balanced accuracy from 0.52 to 0.61; the reverse lowers the BayesLM from 0.77 to 0.61. The patched OracleLM still underperforms relative to the BayesLM, so the step from belief to choice differs between the two models as well.
Plugging each model's decoded beliefs into the assistant's expected-utility calculation (a neuro-symbolic read-out) adds almost nothing for the BayesLM (+0.02) but +0.09 for the OracleLM, suggesting that the OracleLM encodes more information in its beliefs than it uses when it forms recommendations internally.
The same pattern holds for the other two model families, Llama-3.1 8B and Qwen-2.5 7B; the results on this page are for Gemma-2 9B.
| R1 | R2 | R3 | R4 | |
|---|---|---|---|---|
| StartLM | × | × | × | × |
| OracleLM | ✓ | ✓ | × | × |
| BayesLM | ✓ | ✓ | (✓) | (✓) |
Hover over or tab to a mark for a one-sentence explanation. (✓) partly met: R3 does not reach ceiling, and R4 is limited by belief quality.
The advantage of Bayesian supervision. The Bayesian fine-tuning signal, although it does not always match the true user answers, installs better beliefs in the LM and improves the read-out of these beliefs for forming recommendations.
Behavioral evaluation is not enough. A wrong recommendation can come from a wrong belief or from a right belief read out badly. Behavioral evaluations alone cannot tell these failures apart, so understanding LMs requires looking at each layer of the decision-making process.
A Bayesian fine-tuned LM is still a language model. While the Bayesian assistant performs exact inference on every conversation, an LM learns one forward pass for various tasks and can only approximate the posterior, relying on distributed information. Trained to predict single tokens, it struggles more to represent high uncertainty. More Bayesian training may narrow the gap, but how the beliefs form inside the model remains open.
@article{tsvilodub2026bayesian,
title = {Bayesian Fine-tuning Yields Language Models that are
as Bayesian as their Beliefs Allow},
author = {Tsvilodub, Polina and Waldis, Andreas and Qiu, Linlu and
Linzen, Tal and Franke, Michael},
journal = {arXiv preprint arXiv:2610.00679},
year = {2026}
}