Building the tools to better understand large language models.
Cutting-edge AI models often function as black boxes, solving complicated problems through a reasoning process that is hidden within billions of artificial neurons. This lack of interpretability fundamentally limits the trustworthiness of AI technologies, as it is difficult to validate the reasoning that produced any given AI output.
This program aims to address this problem by supporting basic and applied research in AI interpretability, with a particular focus on understanding deceptive behaviors from AI models, including excessive deference to users (sycophancy) and knowingly giving harmful advice. These problems are appearing more frequently in frontier AI systems, which are trained with noisy human feedback, and they are an ideal target for interpretability tools. By inspecting what models are thinking, and not merely what they are saying, we will be able to better detect when humans are getting untrustworthy information from AI.
Funding will support researchers working on interpretability in academia and at nonprofits. We are currently supporting an AI interpretability competition in partnership with NDIF and David Bau at Northeastern University. The competition will pit red teams against blue teams to answer the question: how can we detect when LLMs produce deceptive reasoning or communication? Its blinded design will require blue teams to demonstrate that interpretability tools work in practice by revealing deceptive behaviors that are not known in advance.
Inquiries should be directed to [email protected].