Multimodal

6 entries between January 2021 and August 2026, 2 turning points.

Turning points

  1. OpenAI launches GPT-4o, a single model for text, audio, and images

    OpenAI introduced GPT-4o on 13 May 2024, one model trained across text, image, and audio, demonstrated by chief technology officer Mira Murati in a livestream showing real-time voice conversation. OpenAI gave it to free ChatGPT users and said it handled around 50 languages.

  2. OpenAI releases GPT-4, its first model to accept images

    OpenAI released GPT-4 on 14 March 2023 to ChatGPT Plus subscribers and to developers through a waitlist. The model accepted images as well as text. OpenAI reported it placing near the top 10 percent on a simulated bar exam, against roughly the bottom 10 percent for GPT-3.5.

Every entry

  1. Meta releases Muse Glimmer under an Apache 2.0 license

    Meta released Muse Glimmer on 10 August 2026, a 30-billion-parameter multimodal model distilled from its closed Muse Spark model and published on Hugging Face under Apache 2.0. Compressed to about 4-bit precision, it runs on a single consumer GPU or Mac under 20GB.

  2. Alibaba announces Qwen3.8-Max, its largest model to date

    Alibaba announced Qwen3.8-Max on 3 August 2026 with 2.4 trillion parameters, about 95 billion active at a time through a sparse mixture of experts, a context window up to 1 million tokens, and native text, image and video input. Weights were scheduled to follow a week later.

  3. OpenAI launches GPT-4o, a single model for text, audio, and images

    OpenAI introduced GPT-4o on 13 May 2024, one model trained across text, image, and audio, demonstrated by chief technology officer Mira Murati in a livestream showing real-time voice conversation. OpenAI gave it to free ChatGPT users and said it handled around 50 languages.

  4. Google launches Gemini, built to be multimodal from the start

    Google and Google DeepMind announced Gemini 1.0 on 6 December 2023 in three sizes, Ultra, Pro and Nano. Google described it as built to be multimodal from the start rather than adapted afterward, and said Ultra was the first model past human-expert performance on MMLU, at 90.0%.

  5. OpenAI releases GPT-4, its first model to accept images

    OpenAI released GPT-4 on 14 March 2023 to ChatGPT Plus subscribers and to developers through a waitlist. The model accepted images as well as text. OpenAI reported it placing near the top 10 percent on a simulated bar exam, against roughly the bottom 10 percent for GPT-3.5.

  6. OpenAI unveils CLIP, a zero-shot image classifier

    OpenAI released CLIP on 5 January 2021, a model trained on 400 million image and text pairs to place both in a shared representation. It could sort images into categories it had never been trained on by matching them against written descriptions of the candidate labels.