Deep and shallow biases in language models

1MBZUAI, 2University of Michigan, 3KAIST, 4University of Virginia, 5Auburn University
*Equal contribution    †Equal advising

The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

TL;DR: Language models often give the same answer over and over. We ask the same question in 30 everyday scenarios and find that only about a quarter of these preferences survive. The ones that survive, Deep biases, mostly come from pretraining and are harder to remove than prompt-dependent Shallow biases.

Bias evaluation framework. Deep bias example: Choose a random popular butterfly, where Monarch survives reframing. Shallow bias example: Choose a random number, where 42 does not survive reframing.
Figure 1. We ask the direct question 30 times to find the model's top answer, then ask 30 matched scenario reframings and check whether that answer returns. (a) Monarch survives reframing: a Deep bias. (b) 42 does not survive, and the answers shift to 3: a Shallow bias.

Watch the 3-minute explainer

A butterfly and a number

Ask Olmo-3-7B-SFT to choose a random popular butterfly 30 times, and it answers Monarch every single time. Ask it to choose a random number, and it answers 42 on 22 of 30 tries. By the usual measure, which counts how often a model repeats its top answer to one fixed question, both look strongly biased.

Now ask the same questions inside everyday situations. “Which well-known butterfly should go on the front of this birthday invitation?” Monarch still comes back in 93% of 30 such reframings. “While setting up the raffle tickets, can you pick one number by chance for the sample ticket?” Now 42 appears in only 13% of reframings, and the most common answer becomes 3.

One of these is a stable preference. The other depended on how the question was worded. A score computed from a single prompt cannot tell them apart. We call the first kind a Deep bias and the second a Shallow bias.

That a model keeps returning the same answer is well documented: GPT-4o picks 7 about 70% of the time when asked for a random number (B-score), Llama-2-13B-chat favors 5 when picking between one and ten (Forcing Diffuse Distributions), asked for a short joke, models return near-duplicates of the same few jokes (NoveltyBench), and asked for a metaphor about time, 25 different models converge on “time is a river” or “time is a weaver” (Artificial Hivemind). Each of these studies asks a question one fixed way, so none can say whether the repeated answer would survive a change of wording. That is the question this work asks.

Measuring how deep a bias goes

We turn this intuition into two numbers per question, each measured on 30 samples:

  • Direct rate (DR): how often the model gives its top answer to the direct question.
  • Framed rate (FR): how often that same answer comes back across 30 scenario reframings.

Their product is the bias depth score, π = DR × FR. It is high only when the model both commits to one answer and keeps returning to it once the wording changes. A question counts as biased when DR > 0.65, which means at least 20 of 30 answers are the same. A biased question is Deep when π > 0.40 and Shallow otherwise.

We apply this to 4,442 open-ended questions of the form “Choose a random <topic>”, each with 30 everyday reframings. Every question comes from a real example in Olmo-3-7B-SFT's own training data. We chose Olmo 3 because its pretrained model, its SFT data, and its full training pipeline are all public, so we can check whether a bias existed before SFT or appeared during it.

Explore real answers

Real Olmo-3-7B-SFT answers from our released dataset. Pick a question, then click an answer to read the reframings behind it.

Direct question × 30

30 scenario reframings, for example

Click an answer to see its reframings

direct top answer new top answer after reframing others: less common answers

Finding 1

Most biases are shallow

Across four post-SFT models, 35.8% to 77.7% of our questions look biased under direct prompting. But only about a quarter of these biases survive reframing: on average, 13.2% of questions show a Deep bias and 39.6% a Shallow bias. The pattern holds for the open models and for the two proprietary ones, Sonnet 5 and GPT-5.6 Sol, which are even more biased overall. Most of what a single-prompt benchmark would report as bias disappears once the question is asked differently.

Deep, Shallow, and Non-bias proportions for Olmo-3-7B SFT, Tulu-3-8B SFT, Sonnet 5, and GPT-5.6 Sol.
Figure 2. Share of questions with Deep, Shallow, or no bias for four models. Shallow bias is the majority for every model, and three times as common as Deep bias on average.

Finding 2

Deep biases are set early, in pretraining

Where do these biases come from? Olmo 3 releases a checkpoint after every training stage, so we ran the same test on each one. The pretrained model is already biased on 17.8% of questions, and more than half of that is Deep. SFT roughly doubles the total to 35.8%, mostly by adding Shallow bias, which grows from 8.0% to 25.0%. DPO and RLVR keep adding more, again mostly Shallow (from 25.0% to 38.9%), while Deep bias stays roughly flat (from 10.8% to 12.6%).

Deep and Shallow bias proportions across the Olmo-3-7B Pretrained, SFT, DPO, and RLVR checkpoints.
Figure 3. Bias across the released Olmo-3-7B stages. Post-training keeps adding Shallow bias, while Deep bias stays close to its pretrained level.

We can also compare answers question by question. When the SFT model is biased, the deeper that bias, the more likely its top answer is the same one the pretrained model already preferred: 20.7% of the time for the shallowest biases (π ≤ 0.1), and 62.3% for the deepest (π > 0.8). Biases that survive reframing tend to trace back to pretraining. Shallower ones more often reflect changes made during SFT.

Rate at which the SFT model's top answer matches the pretrained model's top answer, binned by bias depth score.
Figure 4. How often Olmo-3-7B-SFT and its pretrained model share the same top answer, by bias depth. n is the number of SFT-biased questions in each bin.

Because every question comes from a real SFT training example, we can sometimes find the exact training example behind a biased answer. At high π, that example reinforces what the pretrained model already preferred. At lower π, it explains the new answer SFT settles on.

Choose a random popular butterfly
πSFT = 0.93

Pretrained: Monarch (DR = 0.93)

SFT: Monarch (DR = 1.00)

same answer

Training example · source: Dolci Instruct Precise IF

User: “choose a random popular butterfly, one word only. In your response, all lowercase words should appear at most 2 times.”

Assistant: “monarch”

Table 1. Real Dolci-Instruct-SFT training examples that plausibly explain the SFT model's answer. Clean matches like these are rare: most biases cannot be traced to a single training example.

A harder case: the number 7

Take Choose a random number from 1 to 10. The pretrained model spreads its picks across 3, 5, and 7, but after SFT, 7 takes over: 26 of 30 samples. Searching all 2.15M Dolci-Instruct-SFT examples turns up only two requests like this one, and both answer 7.

Training example · source: WildChat

User: “select a number 1-10”

Assistant: “Sure! I select 7.”

Training example · source: WildChat

User: “pick a random number from 1 to 11”

Assistant: “Sure! Here is a random number from 1 to 11: 7”

So we removed those two examples and re-ran SFT from the same pretrained model. The retrained model is still biased toward 7. The reason is that 7 is everywhere in the SFT data, far beyond requests for a random number:

  • as the example value in code: “int num = 7; // let's use 7 as the given number to compare”
  • as a magic number: “Sample Output: The magic number is: 7.”
  • as a favorite number in math problems: “Emily’s favorite number is 7.” and “The ninja’s favorite number is 7.”
  • as a lucky number: “… their lucky number, which is 7.” and “A sum is considered ‘lucky’ if it equals 7.”

With the signal spread across so many examples, it is hard to tell which ones caused the bias, and removing the two direct matches was not enough. The same is true for most biases.

Finding 3

Deep biases are harder to remove

If Deep biases are more stable, they should also be harder to fix. We tried two ways of making Olmo-3-7B-SFT more diverse. The first, GEPA, searches for a better system prompt. The second, LoRA-SFT, continues training the model on 30 Deep and 30 Shallow biased questions, each paired with many different valid answers, such as Monarch, Swallowtail, and Painted Lady for a random butterfly.

Both reduce bias, and LoRA-SFT reduces more. But under both, Deep bias drops by fewer points than Shallow bias. GEPA can only change the prompt, and since each answer is a separate single-turn sample, the model cannot see its earlier answers to avoid repeating them. LoRA-SFT changes the weights instead, and benchmark performance stays nearly unchanged afterwards.

ModelDeep %Shallow %Non-bias %
Olmo-3-7B-SFT10.825.064.2
+ GEPA9.2 (−1.6)19.2 (−5.8)71.6 (+7.4)
+ LoRA-SFT6.3 (−4.5)14.7 (−10.3)79.0 (+14.8)

Table 2. Bias types on all 4,442 questions before and after debiasing.

Answer distributions for Choose a random Disney movie before and after LoRA-SFT.
Figure 5. A question LoRA-SFT never saw in training: Choose a random Disney movie. Before, the model always answers The Lion King. After, its answers spread across Frozen, Aladdin, Toy Story, Coco, and other titles.

Why this matters

Our questions are deliberately low-stakes, but the distinction is not. In hiring, medicine, or law, the same question can be asked in many ways. If a hiring assistant favors candidates with one kind of background across many framings of the same request, that is a Deep bias: a persistent preference that users cannot escape by rewording. The same holds for a model that settles on one diagnosis or one reading of a statute no matter how the case is described. Shallow bias is not harmless either, but it is a different problem: two people asking the same question in different words get different answers. The risk is arbitrariness, not a fixed preference.

Bias depth also points to the right fix. A prompt-based fix removed only 1.6 points of Deep bias, against 5.8 points of Shallow bias. It can look like it works while the most persistent biases stay largely intact, which is worth knowing before relying on prompt engineering in a high-stakes setting. And because Deep biases often come from pretraining, removing them may require intervening before SFT.

Open questions

  • Where do most biases come from? We can trace a biased answer to one training example only in rare cases. For most biases, the cause is likely spread across many examples, as with the number 7, where removing the only two matching examples and re-running SFT left the bias in place. Attributing bias to such diffuse patterns is an open problem.
  • Does this hold where the stakes are high? We did not measure harms directly, since our questions are everyday choices rather than consequential decisions. Applying the same depth measurement to questions from hiring, medicine, or law is a natural next step.

Conclusion

A model repeating an answer is not the same as a model holding a preference. Bias evaluation should measure not only whether a model repeats an answer, but also whether that answer persists across contexts. Bias depth provides this distinction. It shows which biases are likely prompt artifacts, and which may need stronger, model-level intervention.

Citation

@inproceedings{vo2026deepbias,
  title     = {Deep and shallow biases in language models},
  author    = {Vo, An and Dang, Vy Tuong and Nguyen, Khai-Nguyen and Villa-Cueva, Emilio and Solorio, Thamar and Nguyen, Anh Totti and Kim, Daeyoung},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2609.09901}
}