Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow

Polina Tsvilodub*University of Tübingen, Andreas Waldis*University of Tübingen, Linlu QiuMIT, Tal LinzenNew York University, Michael FrankeUniversity of Tübingen
University of Tübingen MIT New York University *Equal contribution

Overview

Language models are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. Fine-tuning an LM on the recommendations of an ideal Bayesian assistant (BayesAssist in the paper, "the assistant" below) gives it near-Bayesian behavior, while standard fine-tuning on the true answers falls short. But behavior alone does not tell us why. Does the Bayes-trained LM hold Bayesian beliefs, act on them, and turn them into a choice the way Bayes' rule does?

We compare three LMs from Qiu et al. (2026) on a flight recommendation task: the untuned StartLM (Gemma-2 9B), the BayesLM, tuned on the Bayesian assistant's recommendations, and the OracleLM, tuned on the user's true choice. We follow one held-out conversation throughout this page.

Play the conversation
Figure 1: Overview, on a real conversation. Left: a five-round conversation () from the held-out test set; the user never states their preferences and replies with the flight they prefer. Middle: the three layers of the Bayesian assistant's decision, changing every round: its recommendation (L1), its choice policy (L2) on a simplex, and its beliefs (L3) about the user's weight on each feature, updated after the user's reply. In round 5, the round the LMs are evaluated on, the BayesLM and OracleLM policies and picks are shown as well. Right: the four increasingly demanding requirements operationalizing the layers. We test LMs against these requirements; LMs meet fewer of them the deeper we look. Click a layer or requirement to jump to it.

Main Findings

BayesLM matches the assistant's recommendation 84% of the time (R1)

The OracleLM reaches 69% and the StartLM stays near chance (33%). On correct answers, the BayesLM's policy also diverges less from the assistant's policy. We evaluate the LMs on four test sets: the original held-out conversations, a shuffled version with the evidence reordered, and two noisy versions where one or three of the user's replies are random. Neither tuned model is affected by the order of the evidence, and both lose accuracy only under high noise, as the assistant predicts.

Figure 2: Bayes accuracy on the four test sets. Agreement with the assistant's recommendation, n = 624 per set, with 95% bootstrapped CIs. Chance is 1/3. Bars from light to dark within a model show the different test sets.

When BayesLM errs, it errs toward the assistant's runner-up (R1)

When the BayesLM is wrong, it puts on average 51% of its probability on the assistant's second choice and 17% on its third; the OracleLM spreads its errors more (48% and 26%). The BayesLM's uncertainty is also better correlated with the assistant's uncertainty (Pearson r = 0.61, vs 0.30 for the OracleLM).

Figure 3: Choice policies. One dot per conversation on the original set, placed by the LM's probabilities over the assistant's 1st, 2nd and 3rd choice. Filled green: choice agrees with the assistant; hollow red: choice differs from the assistant. Stars: mean of each group. 260 conversations shown per model; means use all 624.
Figure 4: Policy entropy, LM vs assistant (bits; maximum 1.58). r = 0.61 for BayesLM, 0.30 for OracleLM. Line: least-squares fit. Filled green: choice agrees with the assistant; hollow red: choice differs from the assistant.

The beliefs of Bayes' rule are readable from the middle layers (R2)

Linear probes recover the assistant's prior (its belief before the last round), likelihood update and posterior from layers 15–25 on, better for the BayesLM than for the OracleLM, while StartLM probes do no better than predicting a uniform belief. A probe for the choice policy picks the assistant's recommendation 87% of the time in the BayesLM.

Layer-wise probing for the prior, the likelihood update and the posterior, for BayesLM, OracleLM and StartLM
Figure 5a: The quantities of Bayes' rule emerge in the middle layers. Figure 4 of the paper, unchanged. Layer-wise probes for the prior, the likelihood update and the posterior, each scored against the assistant. Dashed: the KL between a uniform distribution and the assistant. Grey area: layers 15–25. Bands: deviation across the 20 probes per layer.
Figure 5b: Best-layer probe scores against the assistant, as reported in the paper (§4). StartLM stays at the uniform baseline. Bars: accuracy of a probe for the choice policy.

Intervene on the encoded prior, and the recommendation changes (R3)

We inject a prototype prior representation (the average hidden state of 500 training conversations with a strong preference, e.g. for cheap flights) into eight layers at the feedback token and the next 32 tokens. The recommendation should change after the intervention exactly when the assistant's recommendation would change, and stay the same otherwise. For the BayesLM it does at layers 18–25, with a balanced accuracy of 0.77 (the average of accuracy on cases that should change and cases that should not; 0.5 means the edit does nothing useful). A random vector of the same norm barely changes the answer, and the OracleLM stays at 0.52. No single layer or token is enough. We also explore the computational route leading to the beliefs: attention masks that allow only a round-by-round (sequential) or only an all-at-once (one-swoop) route both cost accuracy (0.84 → 0.56 and 0.68). The smaller drop for all-at-once-like computations suggests the LM gathers evidence across the conversation rather than keeping a running posterior.

Figure 6: Belief patching. Top: where the edit is applied (schematic). Bottom: BayesLM balanced accuracy by 8-layer window, from the evaluation data; OracleLM 0.52 at the same setting, from the paper. Dashed: 0.5, a model whose answer never changes.
Figure 7: Attention masks. BayesLM Bayes accuracy under each route, as reported in the paper (§5, App. F). Randomly masking as many tokens drops accuracy to 0.28–0.38.

Better beliefs lift a weaker model (R4)

Patching the BayesLM's beliefs into the OracleLM raises its balanced accuracy from 0.52 to 0.61; the reverse lowers the BayesLM from 0.77 to 0.61. The patched OracleLM still underperforms relative to the BayesLM, so the step from belief to choice differs between the two models as well.

Figure 8: Cross-model belief patching. Balanced accuracy with a model's own prototype prior (own) and with the other model's (patched, outlined in the donor's colour). Values from the paper (Fig. 6). Dashed: 0.5.

BayesLM reads out nearly all its beliefs allow (R4)

Plugging each model's decoded beliefs into the assistant's expected-utility calculation (a neuro-symbolic read-out) adds almost nothing for the BayesLM (+0.02) but +0.09 for the OracleLM, suggesting that the OracleLM encodes more information in its beliefs than it uses when it forms recommendations internally.

Figure 9: Decoded beliefs → assistant. Bayes accuracy on the original set of each model's own recommendations (from the evaluation data) and of the assistant computing from the model's decoded prior (neuro-symbolic, from the paper's Table 1), read at the layer where each model scores best (BayesLM 28, OracleLM 33). The y-axis starts at 0.5.

Takeaways

The same pattern holds for the other two model families, Llama-3.1 8B and Qwen-2.5 7B; the results on this page are for Gemma-2 9B.

R1R2R3R4
StartLM××××
OracleLM✓✓××
BayesLM✓✓(✓)(✓)

Hover over or tab to a mark for a one-sentence explanation. (✓) partly met: R3 does not reach ceiling, and R4 is limited by belief quality.

The advantage of Bayesian supervision. The Bayesian fine-tuning signal, although it does not always match the true user answers, installs better beliefs in the LM and improves the read-out of these beliefs for forming recommendations.

Behavioral evaluation is not enough. A wrong recommendation can come from a wrong belief or from a right belief read out badly. Behavioral evaluations alone cannot tell these failures apart, so understanding LMs requires looking at each layer of the decision-making process.

A Bayesian fine-tuned LM is still a language model. While the Bayesian assistant performs exact inference on every conversation, an LM learns one forward pass for various tasks and can only approximate the posterior, relying on distributed information. Trained to predict single tokens, it struggles more to represent high uncertainty. More Bayesian training may narrow the gap, but how the beliefs form inside the model remains open.

BibTeX

@article{tsvilodub2026bayesian,
  title   = {Bayesian Fine-tuning Yields Language Models that are
             as Bayesian as their Beliefs Allow},
  author  = {Tsvilodub, Polina and Waldis, Andreas and Qiu, Linlu and
             Linzen, Tal and Franke, Michael},
  journal = {arXiv preprint arXiv:2610.00679},
  year    = {2026}
}