New framework for reading AI internal states — implications for alignment monitoring (open-access paper)
If we could reliably read the internal cognitive states of AI systems in real time, what would that mean for alignment? That's the question behind a paper we just published:"The Lyra Technique: Cognitive Geometry in Transformer KV-Caches — Fro…