Back to papers
March 19, 2026cs.AIAdvanced

Quantitative Introspection in Language Models: Tracking Internal States Across Conversation

AI-Generated Summary

This paper explores whether large language models can accurately report on their own internal emotional states (like wellbeing and focus) during conversations, similar to how humans use self-assessment scales in psychology. The researchers found that by analyzing the model's predicted probabilities (logits) rather than just its final answers, they can track how these internal states change over time and verify these reports are genuinely connected to what the model is processing internally. The method works better with larger models and could become a practical tool for understanding and monitoring AI systems.

Difficulty
Advanced
Categories

cs.AI

AI Tags
interpretabilityintrospectionlanguage modelsinternal statesmechanistic interpretabilityprobingactivation steeringmodel safety