Representation steering is now a common way to mitigate LLM shortcuts. How much legitimate knowledge does this tend to remove? Turns out that these methods can be surprisingly precise! But also: no single steering operation will fix all shortcuts.
Led by @shanzzyy.bsky.social!
Aaron Mueller
Can steering remove LLM shortcuts without breaking legitimate LLM capabilities?
In our @eaclmeeting.bsky.social paper, we show that conceptual bias is separable from concept detection; this means inference-time debiasing is possible with minimal capability loss.