Language model
11 entries between June 2018 and December 2023, 3 turning points.
Turning points
-
OpenAI releases GPT-4, its first model to accept images
OpenAI released GPT-4 on 14 March 2023 to ChatGPT Plus subscribers and to developers through a waitlist. The model accepted images as well as text. OpenAI reported it placing near the top 10 percent on a simulated bar exam, against roughly the bottom 10 percent for GPT-3.5.
-
OpenAI publishes GPT-3, a 175-billion-parameter language model
OpenAI researchers led by Tom B. Brown posted “Language Models are Few-Shot Learners” to arXiv on 28 May 2020. It described GPT-3, an autoregressive language model with 175 billion parameters, ten times larger than any prior non-sparse model.
-
OpenAI unveils GPT-2 and withholds the full model
OpenAI announced GPT-2 on 14 February 2019 with samples of coherent multi-paragraph text, and published only a 124-million-parameter version rather than the full 1.5-billion-parameter model, citing concerns about malicious use. It called the plan a staged release.
Every entry
-
Google launches Gemini, built to be multimodal from the start
Google and Google DeepMind announced Gemini 1.0 on 6 December 2023 in three sizes, Ultra, Pro and Nano. Google described it as built to be multimodal from the start rather than adapted afterward, and said Ultra was the first model past human-expert performance on MMLU, at 90.0%.
-
Meta releases Llama 2 with weights free for commercial use
Meta released Llama 2 on 18 July 2023 at 7, 13 and 70 billion parameters, with the weights and code free for research and for most commercial use. Microsoft was named Meta’s preferred partner, and the models were distributed through Azure, Amazon Web Services and Hugging Face.
-
Anthropic launches Claude, its first public assistant
Anthropic, founded by former OpenAI researchers, released Claude on 14 March 2023 through a chat interface and an API, after a closed alpha with partners including Notion, Quora and DuckDuckGo. Two versions shipped: Claude, and the faster and cheaper Claude Instant.
-
OpenAI releases GPT-4, its first model to accept images
OpenAI released GPT-4 on 14 March 2023 to ChatGPT Plus subscribers and to developers through a waitlist. The model accepted images as well as text. OpenAI reported it placing near the top 10 percent on a simulated bar exam, against roughly the bottom 10 percent for GPT-3.5.
-
Meta releases LLaMA; the weights leak online within days
Meta AI published LLaMA on 24 February 2023, a family of language models at 7, 13, 33 and 65 billion parameters, licensed for noncommercial research and granted case by case. On 3 March the full weights were posted as a BitTorrent link on 4chan and spread from there.
-
Bender, Gebru and coauthors publish “Stochastic Parrots”
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and a coauthor writing as Shmargaret Shmitchell published “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” at the ACM FAccT conference on 3 March 2021.
-
OpenAI publishes GPT-3, a 175-billion-parameter language model
OpenAI researchers led by Tom B. Brown posted “Language Models are Few-Shot Learners” to arXiv on 28 May 2020. It described GPT-3, an autoregressive language model with 175 billion parameters, ten times larger than any prior non-sparse model.
-
OpenAI releases the full 1.5-billion-parameter GPT-2
OpenAI published the full 1.5-billion-parameter GPT-2 model and its code on 5 November 2019, completing the staged rollout begun in February. The company said it had found no strong evidence of misuse of the smaller versions released earlier.
-
OpenAI unveils GPT-2 and withholds the full model
OpenAI announced GPT-2 on 14 February 2019 with samples of coherent multi-paragraph text, and published only a 124-million-parameter version rather than the full 1.5-billion-parameter model, citing concerns about malicious use. It called the plan a staged release.
-
Google publishes BERT, a bidirectional language model
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova published BERT on 11 October 2018, a transformer that pretrains representations of text by reading in both directions, then fine-tuned for particular tasks. Google released the code and weights the following month.
-
OpenAI publishes the GPT-1 paper on pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever published “Improving Language Understanding by Generative Pre-Training” on 11 June 2018, describing a transformer pretrained on unlabeled text and then fine-tuned for particular tasks.