The Rise of Generative AI Large Language Models (LLMs) like ChatGPT

Loading

A visualisation of major large-language models (LLMs), ranked by performance, using MMLU (Massive Multitasks Language Understanding) a benchmark for evaluating the capabilities of large language models.

» main source: Life Architect data
» see our data
» chart rendered with VizSweet

notes

The MMLU rating consists of 16,000 multiple-choice questions across 57 academic subject (link)

MMLU has some critiques, mainly that LLM creators may be wise to the metric and be pre-training their models to answer MMLU questions. (New York Times article). And here’s a general cautionary article.

If you’re interested in ongoing metrics, LLMArena has a leaderboard ranked subjectively by tens of thousands of users. Note: it privileges English language LLMs (there are no Chinese models on there for example). Intel also run a leaderboard with many measurements.

Note: we excluded specialised maths and coding LLMs like DeepSeek Coder, Mathstral, Granite Code etc. Let us know if we’ve missed any LLMs.

A visualisation of selected lawsuits from over 100 filed against AI companies as of June 2026.

» See the data & research

SOURCES: Wired, ChatGPT is Eating the World, news reports

Learn to Create Graphics Like This

Most are copyright claims by authors, record labels, newsrooms, academic publishers and um, Disney who say major AI models like OpenAI’s ChatGPT and Anthropic’s Claude have been trained illegally on their creative and commercial output.

These are the opening salvos in a structural conflict over who owns the raw material of the internet. And what constitutes “fair use” of that material in the age of AI. And whether there will be push back agains the inherently extractive nature of AI models.

Most cases haven’t been decided yet. But some are worthy of detailing:

Bartz vs Anthropic

Three authors – Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson – sued Anthropic, alleging their flagship AI model Claude was trained on millions of pirated e-books downloaded from online “shadow libraries” without permission.

Anthropic settled the case in April 2026 for $1.5billion – the largest copyright settlement in US history.

How much individual authors will receive in compensation is currently unknown, especially now the case is delayed while lawyers’ excessive fees ($320m!) and other legal tangles are ironed out.

You can see if your work or favourite book are included in the training set.

Pile On

Similar cases like Carreyrou vs Open AI, Kwon vs Anthropic, Cruz vs Anthropic, Hobbs vs Meta are using the same tack, accusing the tech firms of torrenting millions of copyright books from shadow websites rather than lawfully licensing them. Given the Bartz vs Anthropic precedent, these cases may have a strong chance of succeeding.

Simultaneously, Elsevier, a large scientific publisher, alongside four other major publishers — Hachette, Macmillan, McGraw Hill, and Cengage — and the best-selling novelist and lawyer Scott Turow are directly accusing Meta and Mark Zuckerberg of doing the same thing for their Llama model.

Disney

Disney is going after popular and powerful image generator MidJourney, calling it a “bottomless pit of plagiarism” for its alleged reproductions of the studios’ best-known characters.

Music Models

At the same time major record labels like Universal Music Group (UMG) and musicians rights holders like Danish Koda are suing the dominant music AI model Suno for apparently training itself on all of Western music. “We are witnessing the largest music theft in history,” Koda’s announcement of the lawsuit reads.

resources & further detail

» All cases tracked @ ChatGPT is Eating The World
» Every Copyright Court Case Visualised (Wired)
» AI Watchdog from The Atlantic
» Another (more legalese) tracker

total lawsuits

(as of Jun 2026)

Open AI – 24
Meta – 14
Microsoft – 10
Anthropic – 9
NVIDIA – 7
Perplexity – 7
Apple – 5
Google – 5
RunwayAI – 4
Adobe – 4
Suno – 4
Stability – 3
Uncharted Labs – 3
Midjourney – 2


CHANGE LOG UPDATES
: 16th Jun – added Who’s Suing Whom graphic
: 4th Feb 2026 – added Kimi K2 Thinking, GPT 5.1 & 5.2, Gemini 3, Claude Opus 4.5, Mistral Larger 3, Nova 2 Pro, Kimi K2.5, DeepSeek v3.2, Grok 4.1
: 17th Sep – added ‘What are people using ChatGPT for?’
: 17th Sep – added SWE-Bench vs LLMArena score chart for major models
: 16th Sep 2025 – major update, adding Claude 3.7 & 4.0 sonnet, Claude Opus 4, Deepseek v3.1, Exaone 4.0, Exaone Deep, Gemini 2.0 Pro, Gemini 2.5 Pro, Gemini 2.5 Flash preview, Gemma 3, GPT 4.5, GPT-5, Grok 3 & 4, Hunyuan Turbo S, Hunyuan A13B, Llama 4 Maverick, Kimi K2, Mistral Saba, Mistral Medium 3, Mistral Small 3.1, OpenAI o3 & o4-mini, Pangu Ultra, Pangu Ultra MoE, Qwen 2.5-max, Qwen 3-235B-A22b, Qwen3-Next, Qwen3-Thinking-2507
: 20th Dec 2024 – added new ranking chart, added LMArena Scores
: 12th Dec – added Amazon Nova Pro, Google Gemini 2.0, Meta’s LLama 3.3, LG’s ExaONE 3.5, and, of course, OpenAI latest interation of o1
: 11th Dec – added requested ‘search’ functionality to viz
: 3rd Nov – updated and created MMLU visualisation
: 31st Oct – added LLMs released since March, including: Anthropic Claude 3.5, Grok 2.0, AMD-Llama, Google’s Gemini 1.5, Meta’s Llama 3.2, OpenAI’s o1 and GPT-4o mini, Pixtral 12b and other offerings from Mistral, Qwen 2.5, EXAONE 3.0, and Apple’s first on device models.
: 31st Oct – new visualisation of LLMs based on MMLU (Massive Multitask Language Understanding) for those models that cite this metric
: 20th Mar 2024 – added 30 new notable LLMs including Anthropic Claude 3, Twitter’s Grok, all Mistral’s offerings, Google Gemini Pro, Apple’s MM1 (finally!) and Chinese LLMs like DeepSeek, GLM-4 and Xinghuo 3.5. The soon-to-be-released column includes OpenAI’s rumoured open source LLM G3PO, Amazon’s mighty Olympus, Meta’s Llama 3 and of course GPT-5.
: 6th Dec 2023 – added 2024 column including Amazon’s Olympus, Anthropic’s Claude-Next and Twitter’s Grok. Also noted the release of Google’s Gemini and Amazon’s Q business bot.
: 21st Nov – added Bichuan 2, Claude Instant, IDEFICS, Jais Chat, Japanese StableLM Alpha 7B, InternLM, Falcon 180B, Bolt 2.5B, DeciLM, Mistral-7B, Persimmon-8B, MoLM, Qwen, AceGPT, Retro48B, Ernie 4.0
: 2nd Nov – updated Amazon story with $1.25bn Anthropic investment
: 27th Jul – added Meta’s LLama2
: 12th Jun – added Claude 2.0, and ErnieBot 3.5
: 21st Jun – added Vicuna 13-B, Falcon LLM, Sail-7B, Web-LLM, OpenLLM
: 20th Jun – visualized all open-source LLMs as a diamond
: 11th Jun – added a ‘more info’ link for each LLM (click to spawn)
: 11th May – added Google’s latest LLM PaLM2 (source)
: 10th May 2023 – Uploaded first version

» All graphics made with our visualisation tool VizSweet

Learn to Create Impactful Infographics


further notes & essential reading

» Why LLMs Can’t Play Wordle and Why That Means They Won’t Lead to AGI
» How ChatGPT and other LLMs Work – And Where They Could Go Next (good Wired article)

» More detailed (understandable) explanation of how LLMs work

» While we’ve plotted these LLMs by the size of each model in billion parameters, there is a growing sense of diminishing returns for simply increasing the model size (Wired article)

» Will A.I. become the new McKinsey? Author Ted Chiang argues that AI is likely to function like larger corporate consulting firms, acting as a “willing executioner”, accelerating job loss (New Yorker)

» The Mounting Environmental & Human Costs of Generative AI (Arstechnica article) &TLDR: larger models = more consumption of planetary resources (minerals, energy, water for cooling) + AI training needs large-scale human supervision so very real possibility of ‘AI sweatshops‘ + the serious issues of copyright infringement for artists and creators.

summarised here:Some of the human & environmental costs of Generative AI

Image from Ars Technica article

A Quick Data Story

Some simple filtering reveals an interesting story in our early LLM data.


Google drove a burst of innovation in the LLM space

Sharing their knowledge and research with the AI world. (Transformers, for example, the ‘T’ in GPT, originated from research at Google). But then the broader company was pipped to the practical-application post by OpenAI. Will their latest release PaLM2 overtake ChatGPT?


OpenAI, creators of ChatGPT, stole the LLM show

They made steady, solid progress over the last three years, driving the curve. Slow and steady won the race?
The Rise of AI Large Language Models - OpenAI's Contribution


And here’s why Microsoft invested in OpenAI

You can see they weren’t directly active in the space with their own research. Instead they invested early and hard in OpenAI ($1bn in both 2019 and 2021) and that paid serious dividends.
The Rise of AI Large Language Models - Microsoft's Contribution


Meta / Facebook also drove significant early innovation in the field…

The Rise of AI Large Language Models - Meta/Facebook's Contribution
What muted their breakthroughs? Maybe the models weren’t large enough (see how many are below the ‘magic’ 175 billion parameter line). Maybe, like Google, there’s was too much emphasis on internal applications & processes versus public tools? Maybe, also, their research was chastened by the poor reception of its science-specialised LLM Galactica.


Meanwhile, in the background, China is also making steady progress

In the advent of ChatGPT, release of Chinese-language LLM’s and chatbots have significantly accelerated.
The Rise of Chinese AI Large Language Models


What about Amazon?

Well, they have steamed in at the end – too late to the party? Time will tell… Though they have recently invested heavily in Anthropic, creators of impressive LLM Claude
The Rise of AI Large Language Models - Amazon's Contribution