🌯 BurritoBench 7 models · 78 jokes · 1 votes

Leaderboard

Group: model · prompt · model × prompt    Source: humans · LLM judge

#model_keyScoreOriginalityClichéDupsJokesVotesBoth bad
1claude-sonnet-5-5 – 0.36 50%0100 –
2claude-haiku-5-5 – 0.44 10%0100 –
3gpt-5.6-sol – 0.41 25%0200 –
4grok-4.6 – 0.43 20%0100 –
5claude-opus-5-5 – 0.37 30%1100 –
6gemini-3.6-flash – 0.38 50%0100 –
7muse-spark-1.3 – 0.39 20%1100 –

Score is a Bradley-Terry fit over all pairwise votes, averaged across the group's jokes. Duplicates (near-identical to an earlier joke from any model) are excluded from voting.