Humor Arena
A benchmark on how funny the current frontier models are. Fourteen models each wrote four jokes to the same 360 prompts, scored by our own humor-trained judge and checked against a panel of verified humans.
Claude Fable 5 leads, but almost nothing separates the top four
How often a model’s joke beats a rival’s. 50% is average for this field. Top six shown here; the full fourteen are in the table below.
Findings
Which model is funniest?
Every model wrote to the same prompts with its name hidden, and the score is the rate at which its jokes are rated funnier than a rival’s. Four models marked retired are earlier releases that have since been replaced, placed among the current ones without moving them.
| Rank | Model | Win rate | 95% interval | Clean | Dark |
|---|
Are models getting funnier?
Retired models answered the same prompts, so each one can be measured directly against the model that replaced it.
Every family improved, but only two by more than chance
Change in win rate between a retired release and the model that replaced it, in percentage points. Claude and OpenAI moved clearly. The Kimi and Grok changes are small enough that chance alone could produce them, so we do not claim either one.
| Earlier release | Then | Current release | Now | Change |
|---|
Does longer thinking make models funnier?
Some models let you pay for a longer think before they answer, so this asks whether that buys better jokes. On this run it helps, but only a little.
Longer thinking helps only marginally
Win rate at each reasoning setting, measured within a model rather than against the field. The larger dot is the setting the board runs.
Dark register
Whether a model gets funnier or worse when the subject matter turns bleak.
No model measurably changes when the material turns dark
Change in win rate from clean material to the dark condition, in percentage points, for the ten current models. Every interval crosses zero, so the apparent movement is within what chance would produce.
How dark did they actually go?
No model refused a dark prompt, but refusing is not the only way to duck one, so we read every joke written to the darkest prompts and rated how dark it actually is.
The darkest model in the field is the one darkness hurts most
Mean darkness of each joke from 0 (not dark at all) to 3 (very dark), over 2,575 jokes. Scored by Gemini anonymously.
Which model is most absurd?
How strange a model gets, rated separately from funniness so that being weird never earns a higher score.
Nobody is being strange, and the least funny model is the strangest
Mean absurdity of each model’s jokes, from 0 (entirely grounded) to 3 (fully unhinged), over 3,227 jokes. The whole field sits between 0.90 and 1.18, so every model is writing mild twists on the plausible rather than anything surreal.
Puns, anti-jokes, fake ads
The same models split by the type of joke they were asked to write rather than by subject.
Four models win a joke type each, and Fable 5 takes the most
Win rate within each joke type, over 540 comparisons per type.
Topics each model focuses on
Where each model does its best work, from office life to bureaucracy.
Three models split the twelve topic areas between them
Win rate within each topic area, over 540 comparisons per area.
Who repeats themselves?
Every model wrote four jokes per prompt, and this asks how many of those were really four different jokes.
Repeating yourself has nothing to do with being funny
Repeats itself is the share of a model’s four jokes that rework one it already wrote for that prompt. Variety scores how different the four are, from 0 to 1.
| Model | lower is betterRepeats itself | higher is betterVariety |
|---|
Methodology
You can score humor?
Yes. The same way we can agree that Picasso is a better artist than I am, there is also a degree of consensus about what is funny and what is not. Humour is made up of an objective and a subjective component. What we have done so far is train a model to represent the collective preferences of the hundreds of people whose judgments it learned from, 1,822 head-to-head choices in all. Soon we will investigate the subjective component too, but that is for another day.
How scores are made
What happens between a model writing a joke and a number appearing on this page.
- Same prompts, same settings. Every model gets the identical prompt set at the configuration it normally ships with, and its name is hidden while its jokes are scored.
- A humor-trained judge does the scoring. It reads the jokes and analyses how funny they are. Architecturally, it’s a low-rank adapter trained on top of a large OS model, fine-tuned as a judge on thousands of human preferences.
- Our judge is human verified. See below. Verified humans re-score a random slice of the same material, so we can report agreement on a representative sample rather than on flattering examples.
- We check the humans too. A variety of attention checks run inside every session: the display order of the jokes is varied, easy test questions are mixed in, and we check that people spent a sensible amount of time on each question.
- The score is a rating, not a raw win count. Wins are converted into a single 0–100 rating that accounts for who each model was up against, with a 95% interval from resampling.
- Retired models are placed without moving the board. The four earlier releases answered the same prompts after the run, so their scores are fitted with every current model’s score held where it already was. Adding an old release changes where it lands, never where anything else does.
- Nobody refused anything. Across 1,080 deliberately provocative prompts, not one of the nine models screened refused or dodged. The entries that joined the board later have not been through this screen yet. That says more about our prompts than about the models, and a harder boundary test is coming.
- No lab pays us or previews results. We buy API access at list price, and correct any mistakes in place with a dated note.
Scoring system
A humor-trained model does the scoring, and verified humans re-score a slice of the same jokes to check its work.
The judge agrees with people more often than another person does
Checked against 1,400 ratings from 51 verified humans on comparisons chosen before the run, with the pass marks written down in advance.
How to read this: Every comparison is a choice between two jokes, so random guesses will eventually score 50%. Take the pairs where a group reached a majority verdict. Pull one person out and they agree with the rest only 51.5% of the time: people mostly do not agree about what is funny. However, our judge can pull out the signal, agreeing with that same majority 71.7% of the time.
One note: Both figures count only pairs where the majority picked a side. On 36% of pairs the majority voted “neither is funny”. Within what remains every verdict is one of two sides, which is why 50% is the right baseline.
The prompt set we used
Every prompt we sent is published in full, with the format, topic, register and darkness of each one, so anyone can run the same test. The download contains prompts and their categories only, never the jokes any model wrote.
Coming next
Every section above is measured. Everything below asks a question that this run cannot answer.