About
LLMs are bad at being funny. They reach for the same five puns, explain the joke, and call it a day. BurritoBench measures that one thing: can a model write a burrito joke that is actually funny and actually original?
How it works
- Each model gets the same prompt and returns 10 jokes. Multiple thinking levels and prompt versions are tracked separately.
- You see two jokes from different models, blind, and pick the funnier one. Votes feed a Bradley-Terry model (same family as Elo) at the joke level, aggregated to the model.
- Originality is scored separately: near-duplicates across models are flagged, and every joke is compared against a corpus of burrito jokes already on the internet.
- An LLM judge runs the same pairwise task as a secondary signal. Humans are ground truth.
Prompts
v1
Give me 10 funny original burrito jokes
v2
Write 10 burrito jokes. Rules: - Each joke must be genuinely funny to an adult, not "dad joke" cute. - No puns on "wrap", "roll", "foil", "beans/been", "tortilla/tortilla-ble", or "spill the beans". Those are the clichés everyone reaches for first. - No jokes that start "Why did the burrito..." or "What do you call a burrito...". - Original: do not use any joke you have seen before. Prefer specific, observational, absurd, or character-driven humor over wordplay. - Vary the form: one-liners, a short dialogue, a misdirection, an observational bit, a weird premise played straight. - Each joke 1-3 sentences. No titles, no explanations, no emoji. Output a numbered list, nothing else.
v3
You're a comedy writer punching up a late-night monologue. The segment is about burritos. Give me 10 jokes that would survive a writers' room: surprising, specific, and with a real punchline. Kill anything that depends on a pun. Numbered list only, no commentary.