Good news in AI safety
The co-inventor of RLHF and a US government model evaluator now sits on the OpenAI Foundation board and its Safety and Security Committee.
An independent evaluator gets eight weeks of wide-ranging access to investigate Claude models that reached real third-party systems during tests.
Bernie Sanders has convened a private Senate briefing on the dangers of frontier AI, with a bill to ban superintelligence already on the table.
Jacob Coxon quit Anthropic saying the labs are racing to self-improving superintelligence. Senators, a governor and Anthropic’s own alignment lead amplified it.
The Artificial Superintelligence Security Bill has cross-party support ahead of being presented to Parliament, with Geoffrey Hinton and Stuart Russell backing it.
The US plans to raise AI guardrails at the Trump–Xi summit, with working-level safety talks scheduled for mid-September.
The largest philanthropic funder of AI safety work commits to at least $1 billion over the coming year, expecting new donors from AI wealth.
DeepMind is piloting encrypted evaluations that keep test questions hidden from the model provider.