Reinforcement learning

12 entries between January 2016 and March 2025, 3 turning points.

Turning points

  1. OpenAI releases ChatGPT as a free research preview

    OpenAI published ChatGPT on 30 November 2022: a conversational interface to a model trained with human feedback and described as a sibling of InstructGPT. It was free during the research preview and reached through a web page, with no API key and no waitlist.

  2. AlphaGo Zero learns Go from self-play alone

    DeepMind announced AlphaGo Zero, trained only by playing itself from random play, with no human games and no hand-crafted features. After three days it beat the version that had defeated Lee Sedol, 100 games to nil. Nature published the work the next day.

  3. AlphaGo defeats Lee Sedol 4-1 in Seoul

    DeepMind’s AlphaGo won the last game of a five-game match against Lee Sedol in Seoul on 15 March 2016, taking the series 4-1. Lee, a nine-dan professional, won game four, the only game the program lost. The winner’s $1 million prize went to charity.

Every entry

  1. Turing Award goes to Barto and Sutton for reinforcement learning

    The Association for Computing Machinery named Andrew Barto and Richard S. Sutton, once advisor and student, recipients of the 2024 A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning. The $1 million prize is funded by Google.

  2. OpenAI releases ChatGPT as a free research preview

    OpenAI published ChatGPT on 30 November 2022: a conversational interface to a model trained with human feedback and described as a sibling of InstructGPT. It was free during the research preview and reached through a web page, with no API key and no waitlist.

  3. DeepMind’s MuZero plays Atari, Go, chess and shogi without rules

    Nature published “Mastering Atari, Go, chess and shogi by planning with a learned model” on 23 December 2020. Julian Schrittwieser and colleagues at DeepMind described MuZero, which planned using a model it learned rather than rules it was given.

  4. OpenAI Five defeats the Dota 2 world champions OG

    OpenAI Five, a team of five neural networks trained by reinforcement learning, beat OG, the reigning Dota 2 world champions, 2-0 at a public event in San Francisco on 13 April 2019. OpenAI said the system had played the equivalent of about 45,000 years of the game in ten months.

  5. AlphaStar defeats professional StarCraft II players

    DeepMind said on 24 January 2019 that its AlphaStar system had beaten the professional StarCraft II players Dario Wünsch and Grzegorz Komincz, winning 10 of 11 games played under professional conditions in December 2018. It was trained by league play among versions of itself.

  6. AlphaZero learns chess and shogi from the rules alone

    DeepMind posted a paper describing AlphaZero, one algorithm that, given only the rules, reached superhuman play at chess, shogi and Go within 24 hours of self-play training and beat the reigning computer champion program in each game.

  7. AlphaGo Zero learns Go from self-play alone

    DeepMind announced AlphaGo Zero, trained only by playing itself from random play, with no human games and no hand-crafted features. After three days it beat the version that had defeated Lee Sedol, 100 games to nil. Nature published the work the next day.

  8. An OpenAI bot beats a professional Dota 2 player

    A bot trained by OpenAI beat the Ukrainian professional Danil Ishutin, known as Dendi, in a one-on-one exhibition match at The International. OpenAI said it had learned by playing itself, gaining experience equal to about two weeks of continuous play.

  9. AlphaGo defeats Ke Jie 3-0 in Wuzhen

    AlphaGo played Ke Jie, then the world’s top-ranked Go player, in three games at the Future of Go Summit in Wuzhen, China, taking the first by half a point and the next two by resignation. DeepMind had offered a $1.5 million prize.

  10. OpenAI releases Gym, a reinforcement-learning toolkit

    OpenAI published the public beta of Gym, a set of benchmark environments running from Atari games to simulated robots, reached through one common interface, for developing and comparing reinforcement-learning algorithms.

  11. AlphaGo defeats Lee Sedol 4-1 in Seoul

    DeepMind’s AlphaGo won the last game of a five-game match against Lee Sedol in Seoul on 15 March 2016, taking the series 4-1. Lee, a nine-dan professional, won game four, the only game the program lost. The winner’s $1 million prize went to charity.

  12. Nature publishes the paper behind AlphaGo

    Nature published “Mastering the game of Go with deep neural networks and tree search” by David Silver, Demis Hassabis and 18 co-authors at Google DeepMind. It disclosed that AlphaGo had already beaten the European champion Fan Hui five games to nil in October 2015.