Large language model
10 entries between June 2021 and July 2024, 1 turning point.
Turning points
-
OpenAI releases ChatGPT as a free research preview
OpenAI published ChatGPT on 30 November 2022: a conversational interface to a model trained with human feedback and described as a sibling of InstructGPT. It was free during the research preview and reached through a web page, with no API key and no waitlist.
Every entry
-
Meta releases Llama 3.1 with a 405-billion-parameter model
Meta released Llama 3.1 on 23 July 2024, including a 405-billion-parameter model it called the largest openly available foundation model to date, with a 128,000-token context window. Meta said it matched GPT-4o and Claude 3.5 Sonnet on knowledge, math, and tool-use benchmarks.
-
Meta releases Llama 3 under an open license
Meta released the first Llama 3 models, at 8 billion and 70 billion parameters, on 18 April 2024 under a license permitting most commercial use. Meta said the 70B model beat comparable models, including Claude 3 Sonnet, on human evaluations across twelve use cases.
-
Anthropic launches the Claude 3 model family
Anthropic released Claude 3 Opus, Sonnet, and Haiku on 4 March 2024, three models spanning a range of speed and capability, with a 200,000-token context window and image input. Anthropic reported that Opus matched or exceeded GPT-4 on several public benchmarks.
-
OpenAI releases ChatGPT as a free research preview
OpenAI published ChatGPT on 30 November 2022: a conversational interface to a model trained with human feedback and described as a sibling of InstructGPT. It was free during the research preview and reached through a web page, with no API key and no waitlist.
-
Meta withdraws its Galactica science model after three days
Meta AI released Galactica on 15 November 2022, a language model trained on scientific text and intended to help write papers and summaries, with a public demo. Users showed it producing fluent but false claims and invented citations, and Meta took the demo down on 17 November.
-
BigScience releases BLOOM with 176 billion parameters
The BigScience collective, hundreds of volunteer researchers coordinated around Hugging Face and using France’s Jean Zay supercomputer, released BLOOM on 12 July 2022. The 176-billion-parameter model wrote 46 natural languages and 13 programming languages.
-
Google introduces PaLM, a 540-billion-parameter model
Google Research introduced the Pathways Language Model on 4 April 2022, a 540-billion-parameter dense transformer trained across 6,144 TPU v4 chips using Google’s Pathways system. Google reported the best few-shot scores it had published on most of the benchmarks tested.
-
Chinchilla finds most large language models are undertrained
DeepMind posted “Training Compute-Optimal Large Language Models” on 29 March 2022. Its 70-billion-parameter Chinchilla, trained on four times the data used for the 280-billion-parameter Gopher at equal compute cost, outperformed Gopher and GPT-3 across a broad set of benchmarks.
-
Microsoft and Nvidia unveil a 530-billion-parameter model
Microsoft and Nvidia announced Megatron-Turing NLG on 11 October 2021, a 530-billion-parameter transformer language model. It was trained with DeepSpeed and Megatron-LM on 560 Nvidia DGX A100 servers, drawing largely on the Pile and Common Crawl web data.
-
BAAI unveils Wu Dao 2.0, a 1.75-trillion-parameter model
The Beijing Academy of Artificial Intelligence unveiled Wu Dao 2.0 on 1 June 2021, a multimodal pretrained model it said used 1.75 trillion parameters, about ten times the count in GPT-3. BAAI said it trained on 1.2 terabytes of English and Chinese text.