Papers & Explorations
Mechanistic interpretability and chain-of-thought faithfulness as complementary AI safety agendas: a literature review.