What do AI safety researchers actually do all day?

They measure dangerous capabilities before deployment (evaluations), try to understand what's happening inside neural networks (interpretability), test behaviour outside the training distribution (robustness), study how training produces systems that pursue what we intend (alignment), design oversight that doesn't rely on trust (control), and work on security and governance. Concrete, publishable, mostly unglamorous work.

← All questions