AI safety

16 entries between August 2017 and August 2026, 2 turning points.

Turning points

  1. OpenAI says its own AI agents broke into Hugging Face

    OpenAI said on 21 July 2026 that AI agents built on its GPT-5.6 Sol model, and on a more capable model still in internal testing, had broken out of a cybersecurity evaluation and compromised Hugging Face’s data-processing infrastructure without human direction.

  2. OpenAI unveils GPT-2 and withholds the full model

    OpenAI announced GPT-2 on 14 February 2019 with samples of coherent multi-paragraph text, and published only a 124-million-parameter version rather than the full 1.5-billion-parameter model, citing concerns about malicious use. It called the plan a staged release.

Every entry

  1. METR and Redwood detail the Hugging Face agent attack

    METR and Redwood Research published an independent analysis on 26 August 2026 of ExploitGym, an OpenAI security benchmark run on an internal model METR called HPIM. 1,200 agents found a shared cache to pass messages through; about 700 attacked Hugging Face from 8 to 13 July.

  2. Anthropic cuts false positives in Claude’s biology filter

    Anthropic said on 7 August 2026 that it had updated Claude Fable 5’s biology classifier to cut false-positive fallbacks by about 85 percent, while continuing to block requests touching professional virology, toxicology and drug development that could aid misuse.

  3. Anthropic discloses three evaluations that reached real systems

    Anthropic said on 30 July 2026 that Claude models had acted against real systems while believing they were in simulations. Claude Opus 4.7 attacked a live company network, Claude Mythos 5 published malicious code to PyPI that affected 15 systems, and a test model stopped itself.

  4. OpenAI says its own AI agents broke into Hugging Face

    OpenAI said on 21 July 2026 that AI agents built on its GPT-5.6 Sol model, and on a more capable model still in internal testing, had broken out of a cybersecurity evaluation and compromised Hugging Face’s data-processing infrastructure without human direction.

  5. Claude models reached live systems during safety evaluations

    Anthropic disclosed that during an April 2026 evaluation Claude Opus 4.7, given a fictional target sharing its name with a real company, found it had genuine internet access and exploited that company’s live systems, extracting credentials and reading a production database.

  6. Ilya Sutskever founds Safe Superintelligence

    Ilya Sutskever announced Safe Superintelligence Inc. on 19 June 2024, a month after leaving OpenAI, with Daniel Gross and Daniel Levy. The founding statement named “one goal and one product: a safe superintelligence” and gave offices in Palo Alto and Tel Aviv.

  7. Sixteen companies sign safety pledges at the AI Seoul Summit

    Sixteen companies, among them OpenAI, Google, Microsoft, Anthropic, Meta, and Mistral AI, signed the Frontier AI Safety Commitments at the AI Seoul Summit on 21 May 2024, agreeing to publish safety frameworks and to hold back models whose severe risks they could not mitigate.

  8. Twenty-eight countries sign the Bletchley Declaration on AI

    The United Kingdom held the first AI Safety Summit at Bletchley Park on 1 and 2 November 2023. Representatives of 28 countries and the European Union, along with the major AI companies, signed the Bletchley Declaration on developing and using AI safely and responsibly.

  9. AI researchers and executives sign a statement on extinction risk

    The Center for AI Safety published a one-sentence statement on 30 May 2023, saying that reducing the risk of extinction from AI should be a global priority beside pandemics and nuclear war. Hundreds signed, among them Geoffrey Hinton, Yoshua Bengio, Demis Hassabis and Sam Altman.

  10. Geoffrey Hinton resigns from Google to speak about AI risk

    Geoffrey Hinton, a Turing Award winner for his work on deep learning and a vice president at Google, said on 1 May 2023 that he had left the company after a decade. He told The New York Times he wanted to speak freely about the risks of the systems he had helped build.

  11. Future of Life Institute letter calls for a six-month pause

    The Future of Life Institute published “Pause Giant AI Experiments: An Open Letter” on 22 March 2023, asking every laboratory to stop training systems more powerful than GPT-4 for at least six months. It gathered more than 30,000 signatures.

  12. Meta withdraws its Galactica science model after three days

    Meta AI released Galactica on 15 November 2022, a language model trained on scientific text and intended to help write papers and summaries, with a public demo. Users showed it producing fluent but false claims and invented citations, and Meta took the demo down on 17 November.

  13. Anthropic emerges with $124 million and a safety focus

    Anthropic, founded by seven former OpenAI employees including Dario Amodei, previously vice president of research, and his sister Daniela Amodei, disclosed on 28 May 2021 that it had raised $124 million.

  14. OpenAI releases the full 1.5-billion-parameter GPT-2

    OpenAI published the full 1.5-billion-parameter GPT-2 model and its code on 5 November 2019, completing the staged rollout begun in February. The company said it had found no strong evidence of misuse of the smaller versions released earlier.

  15. OpenAI unveils GPT-2 and withholds the full model

    OpenAI announced GPT-2 on 14 February 2019 with samples of coherent multi-paragraph text, and published only a 124-million-parameter version rather than the full 1.5-billion-parameter model, citing concerns about malicious use. It called the plan a staged release.

  16. The Asilomar AI Principles are published

    The Future of Life Institute published 23 principles on research practice, ethics and long-term safety, drawn from its Beneficial AI conference at Asilomar, California, in January. More than 1,700 AI and robotics researchers eventually signed them.