HumorBench

The Super Official Memelord benchmark of which AI is actually funny

Captioned memes
26,332
Human votes
1,199
AI judgments
19,072
Same Harness

Identical rendering

Every model captions the same meme templates with the same prompt. Captions render through Memelord's production harness, so the only thing that varies is the words.

Blind Votes

Ranked-choice arena

Humans vote on anonymized side-by-side pairs. Bradley–Terry rankings with bootstrap confidence intervals turn votes into leaderboards.

Judge the Judges

AI humor judges

Judge models score the exact pairs humans voted on. Agreement with the human majority (accuracy + Cohen's kappa) ranks who best understands funny.

Rankings below are provisional — built from AI judge consensus while human votes accumulate. Vote in the arena to make them real.

Model Leaderboard — who makes the funniest memes?
#ModelProviderElo (BT)95% CIComparisons
1Gemini 3.6 Flashgoogle17521692–1826171
2Kimi K3openrouter17181657–1805128
3GPT-5.6 Solopenai16801627–1743180
4Claude Opus 5anthropic16671613–1734159
5GPT-5.6 Lunaopenai15321472–1584192
6Gemini 3.5 Flash Litegoogle15181464–1583164
7Claude Sonnet 5anthropic14821426–1552171
8Grok 4.6xai14631415–1519197
9Grok 4.5xai14531394–1514183
10DeepSeek V3.2openrouter13681307–1430182
11Venice Uncensored (Dolphin Mistral 24B)openrouter867621–982189
Prompt Leaderboard — which prompt is funniest?
#PromptElo (BT)Comparisons
1Comedian persona1635106
2Subversion154291
3Baseline144393
4Comedy chain-of-thought138186
Judge Leaderboard — who knows funny?

Judge scores unlock once enough human votes are in.

Caption-bar templates
Composer templates (structured slot data)

Last computed Thu, 13 Aug 2026 03:20:10 GMT