Back to papers
March 19, 2026cs.AIAdvanced
Quantitative Introspection in Language Models: Tracking Internal States Across Conversation
AI-Generated Summary
This paper explores whether large language models can accurately report on their own internal emotional states (like wellbeing and focus) during conversations, similar to how humans use self-assessment scales in psychology. The researchers found that by analyzing the model's predicted probabilities (logits) rather than just its final answers, they can track how these internal states change over time and verify these reports are genuinely connected to what the model is processing internally. The method works better with larger models and could become a practical tool for understanding and monitoring AI systems.
Difficulty
Advanced
Categories
cs.AI
AI Tags
interpretabilityintrospectionlanguage modelsinternal statesmechanistic interpretabilityprobingactivation steeringmodel safety