Mechanistic Interpretability and Trustworthy LLMs
I investigate how language models internally represent information and produce decisions, particularly when their outputs are sensitive to irrelevant context, role framing, or subtle variations in input. My research combines behavioral evaluation with mechanistic analysis to identify the internal components and representations associated with model reasoning and failure.
The broader goal is to develop language models that are not only accurate, but also robust, interpretable, and appropriately calibrated for high-stakes decision-making.
