About
Hi, I’m Satwik. I’m an independent researcher working on the safety and reliability of reasoning and agentic systems.
My work spans uncertainty quantification, multimodal failure modes, reasoning and agentic evaluation, and interpretability, with a recurring focus on understanding when models fail and how those failures can be detected or audited. I’m drawn to the space between academic research and deployment, methods that hold up outside a benchmark and can actually be run on a system someone depends on. My research has appeared at ICML 2026 and COLM 2026 workshops.
By day, I’m an AI Research Engineer at VFS Global, where I work on production LLM/VLM systems for document intelligence, confidence-aware routing, and multimodal verification, serving 110+ countries at over 4M calls a day.
More recently, I’ve been moving deeper into mechanistic interpretability, with a pragmatic and actionable bent: I’m most interested in cases where understanding model internals helps answer a concrete safety question or enables a better audit, intervention, or decision. In particular, I’m interested in model forensics, memorization and unlearning, faithfulness, and chain-of-thought monitorability.
Always happy to talk research, interpretability, or reliability. Feel free to reach out over email or X!
News
New preprint: Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model.
Joined Kang Lab at MBZUAI as a Research Collaborator, advised by Prof. Jian Kang.
Attended ICML 2026 in Seoul, South Korea, including the CTB and FAGEN workshops where our work was accepted.
Don’t Blink was accepted at the Actionable Interpretability Workshop (AIW) and the Scientific Understanding of Foundation Models Workshop (SCI-FM) at COLM 2026.
Two papers accepted at the Failure Modes in Agentic AI (FAGEN) workshop at ICML 2026: SELFDOUBT and Proper Scoring Rules for Agentic Uncertainty Quantification.
Proper Scoring Rules for Agentic Uncertainty Quantification accepted as a poster at the Combining Theory and Benchmarks (CTB) workshop at ICML 2026.
New preprint: Proper Scoring Rules for Agentic Uncertainty Quantification. We introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper, trajectory-level scoring rules for evaluating uncertainty in LLM agents.
New preprint: SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio. We introduce HVR, a single-pass uncertainty signal for reasoning LLMs that outperforms Semantic Entropy at about 10x lower inference cost.
New preprint: Don’t Blink: Evidence Collapse during Multimodal Reasoning. We identify evidence collapse, a decay of visual grounding during multimodal reasoning that text-only uncertainty signals cannot detect.
Joined VFS Global as an AI Research Engineer.