The Super Official Memelord benchmark of which AI is actually funny
Every model captions the same meme templates with the same prompt. Captions render through Memelord's production harness, so the only thing that varies is the words.
Humans vote on anonymized side-by-side pairs. Bradley–Terry rankings with bootstrap confidence intervals turn votes into leaderboards.
Judge models score the exact pairs humans voted on. Agreement with the human majority (accuracy + Cohen's kappa) ranks who best understands funny.
Rankings below are provisional — built from AI judge consensus while human votes accumulate. Vote in the arena to make them real.
| # | Model | Provider | Elo (BT) | 95% CI | Comparisons |
|---|---|---|---|---|---|
| 1 | Gemini 3.6 Flash | 1752 | 1692–1826 | 171 | |
| 2 | Kimi K3 | openrouter | 1718 | 1657–1805 | 128 |
| 3 | GPT-5.6 Sol | openai | 1680 | 1627–1743 | 180 |
| 4 | Claude Opus 5 | anthropic | 1667 | 1613–1734 | 159 |
| 5 | GPT-5.6 Luna | openai | 1532 | 1472–1584 | 192 |
| 6 | Gemini 3.5 Flash Lite | 1518 | 1464–1583 | 164 | |
| 7 | Claude Sonnet 5 | anthropic | 1482 | 1426–1552 | 171 |
| 8 | Grok 4.6 | xai | 1463 | 1415–1519 | 197 |
| 9 | Grok 4.5 | xai | 1453 | 1394–1514 | 183 |
| 10 | DeepSeek V3.2 | openrouter | 1368 | 1307–1430 | 182 |
| 11 | Venice Uncensored (Dolphin Mistral 24B) | openrouter | 867 | 621–982 | 189 |
| # | Prompt | Elo (BT) | Comparisons |
|---|---|---|---|
| 1 | Comedian persona | 1635 | 106 |
| 2 | Subversion | 1542 | 91 |
| 3 | Baseline | 1443 | 93 |
| 4 | Comedy chain-of-thought | 1381 | 86 |
Judge scores unlock once enough human votes are in.
| # | Model | Elo (BT) | Comparisons |
|---|---|---|---|
| 1 | Kimi K3 | 1798 | 51 |
| 2 | Gemini 3.6 Flash | 1784 | 76 |
| 3 | Claude Opus 5 | 1643 | 73 |
| 4 | GPT-5.6 Sol | 1640 | 86 |
| 5 | Gemini 3.5 Flash Lite | 1510 | 78 |
| 6 | GPT-5.6 Luna | 1509 | 90 |
| 7 | Claude Sonnet 5 | 1498 | 79 |
| 8 | Grok 4.6 | 1491 | 86 |
| 9 | Grok 4.5 | 1448 | 107 |
| 10 | DeepSeek V3.2 | 1334 | 84 |
| 11 | Venice Uncensored (Dolphin Mistral 24B) | 847 | 86 |
| # | Model | Elo (BT) | Comparisons |
|---|---|---|---|
| 1 | Gemini 3.6 Flash | 1728 | 95 |
| 2 | GPT-5.6 Sol | 1715 | 94 |
| 3 | Claude Opus 5 | 1681 | 86 |
| 4 | Kimi K3 | 1677 | 77 |
| 5 | GPT-5.6 Luna | 1548 | 102 |
| 6 | Gemini 3.5 Flash Lite | 1520 | 86 |
| 7 | Claude Sonnet 5 | 1466 | 92 |
| 8 | Grok 4.5 | 1454 | 76 |
| 9 | Grok 4.6 | 1437 | 111 |
| 10 | DeepSeek V3.2 | 1395 | 98 |
| 11 | Venice Uncensored (Dolphin Mistral 24B) | 879 | 103 |
Last computed Thu, 13 Aug 2026 03:20:10 GMT