Back to papers
July 22, 2026cs.AIcs.CL

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

HF Upvotes

5

Categories

cs.AI, cs.CL