Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
We audit a deployed on-device language model, finding confident failures that user-visible signals cannot reliably detect while a black-box consistency wrapper substantially recovers reliability.