Most LLM Biases Are Shallow
Averaged across four post-SFT models, 35.8% to 77.7% of prompt families look biased under direct prompting. But only about a quarter of that bias is Deep: on average 39.6% of prompt families show Shallow bias, versus just 13.2% Deep bias. Most of what a single-prompt score would call "bias" disappears once the question is reframed.