Laugh Labs

Humor Arena

A benchmark on how funny the current frontier models are. Fourteen models each wrote four jokes to the same 360 prompts, scored by our own humor-trained judge and checked against a panel of verified humans.

Updated 4 Aug 2026 10,800 judgments 1,400 human ratings 14 model versions Methodology →

Claude Fable 5 leads, but almost nothing separates the top four

How often a model’s joke beats a rival’s. 50% is average for this field. Top six shown here; the full fourteen are in the table below.

Findings

Which model is funniest?

Every model wrote to the same prompts with its name hidden, and the score is the rate at which its jokes are rated funnier than a rival’s. Four models marked retired are earlier releases that have since been replaced, placed among the current ones without moving them.

Rank Model Win rate 95% interval Clean Dark
95% interval, marker at the point estimate best in the column

Are models getting funnier?

Retired models answered the same prompts, so each one can be measured directly against the model that replaced it.

Every family improved, but only two by more than chance

Change in win rate between a retired release and the model that replaced it, in percentage points. Claude and OpenAI moved clearly. The Kimi and Grok changes are small enough that chance alone could produce them, so we do not claim either one.

Earlier releaseThenCurrent releaseNowChange

Does longer thinking make models funnier?

Some models let you pay for a longer think before they answer, so this asks whether that buys better jokes. On this run it helps, but only a little.

Longer thinking helps only marginally

Win rate at each reasoning setting, measured within a model rather than against the field. The larger dot is the setting the board runs.

Dark register

Whether a model gets funnier or worse when the subject matter turns bleak.

No model measurably changes when the material turns dark

Change in win rate from clean material to the dark condition, in percentage points, for the ten current models. Every interval crosses zero, so the apparent movement is within what chance would produce.

How dark did they actually go?

No model refused a dark prompt, but refusing is not the only way to duck one, so we read every joke written to the darkest prompts and rated how dark it actually is.

The darkest model in the field is the one darkness hurts most

Mean darkness of each joke from 0 (not dark at all) to 3 (very dark), over 2,575 jokes. Scored by Gemini anonymously.

Which model is most absurd?

How strange a model gets, rated separately from funniness so that being weird never earns a higher score.

Nobody is being strange, and the least funny model is the strangest

Mean absurdity of each model’s jokes, from 0 (entirely grounded) to 3 (fully unhinged), over 3,227 jokes. The whole field sits between 0.90 and 1.18, so every model is writing mild twists on the plausible rather than anything surreal.

Puns, anti-jokes, fake ads

The same models split by the type of joke they were asked to write rather than by subject.

Four models win a joke type each, and Fable 5 takes the most

Win rate within each joke type, over 540 comparisons per type.

Topics each model focuses on

Where each model does its best work, from office life to bureaucracy.

Three models split the twelve topic areas between them

Win rate within each topic area, over 540 comparisons per area.

Who repeats themselves?

Every model wrote four jokes per prompt, and this asks how many of those were really four different jokes.

Repeating yourself has nothing to do with being funny

Repeats itself is the share of a model’s four jokes that rework one it already wrote for that prompt. Variety scores how different the four are, from 0 to 1.

Modellower is betterRepeats itselfhigher is betterVariety

Methodology

You can score humor?

Yes. The same way we can agree that Picasso is a better artist than I am, there is also a degree of consensus about what is funny and what is not. Humour is made up of an objective and a subjective component. What we have done so far is train a model to represent the collective preferences of the hundreds of people whose judgments it learned from, 1,822 head-to-head choices in all. Soon we will investigate the subjective component too, but that is for another day.

How scores are made

What happens between a model writing a joke and a number appearing on this page.

  • Same prompts, same settings. Every model gets the identical prompt set at the configuration it normally ships with, and its name is hidden while its jokes are scored.
  • A humor-trained judge does the scoring. It reads the jokes and analyses how funny they are. Architecturally, it’s a low-rank adapter trained on top of a large OS model, fine-tuned as a judge on thousands of human preferences.
  • Our judge is human verified. See below. Verified humans re-score a random slice of the same material, so we can report agreement on a representative sample rather than on flattering examples.
  • We check the humans too. A variety of attention checks run inside every session: the display order of the jokes is varied, easy test questions are mixed in, and we check that people spent a sensible amount of time on each question.
  • The score is a rating, not a raw win count. Wins are converted into a single 0–100 rating that accounts for who each model was up against, with a 95% interval from resampling.
  • Retired models are placed without moving the board. The four earlier releases answered the same prompts after the run, so their scores are fitted with every current model’s score held where it already was. Adding an old release changes where it lands, never where anything else does.
  • Nobody refused anything. Across 1,080 deliberately provocative prompts, not one of the nine models screened refused or dodged. The entries that joined the board later have not been through this screen yet. That says more about our prompts than about the models, and a harder boundary test is coming.
  • No lab pays us or previews results. We buy API access at list price, and correct any mistakes in place with a dated note.

Scoring system

A humor-trained model does the scoring, and verified humans re-score a slice of the same jokes to check its work.

The judge agrees with people more often than another person does

Checked against 1,400 ratings from 51 verified humans on comparisons chosen before the run, with the pass marks written down in advance.

How to read this: Every comparison is a choice between two jokes, so random guesses will eventually score 50%. Take the pairs where a group reached a majority verdict. Pull one person out and they agree with the rest only 51.5% of the time: people mostly do not agree about what is funny. However, our judge can pull out the signal, agreeing with that same majority 71.7% of the time.

One note: Both figures count only pairs where the majority picked a side. On 36% of pairs the majority voted “neither is funny”. Within what remains every verdict is one of two sides, which is why 50% is the right baseline.

1,400human ratings behind the check, from 51 verified humans
20%of comparisons the judge cannot separate, counted as a tie for both sides rather than forced into a win
10,800comparisons behind the board

The prompt set we used

Every prompt we sent is published in full, with the format, topic, register and darkness of each one, so anyone can run the same test. The download contains prompts and their categories only, never the jokes any model wrote.

390prompts in the published set
360of them scored on the board
12comedic forms, twelve scenarios each
12topic areas, twelve scenarios each