Natural Language Autoencoders: Turning Claude's Thoughts into Text
↗Anthropic presents Natural Language Autoencoders (NLAs) that translate model activations into readable text to audit and understand Claude’s internal thinking, showing measurable reconstruction-based quality, notable safety/auditing insights, but with high compute costs and potential factual hallucinations.
May 7, 20261%