HumorBench

The Super Official Memelord benchmark of which AI is actually funny

Captioned memes
26,332
Human votes
852
AI judgments
19,072
Same Harness

Identical rendering

Every model captions the same meme templates with the same prompt. Captions render through Memelord's production harness, so the only thing that varies is the words.

Blind Votes

Ranked-choice arena

Humans vote on anonymized side-by-side pairs. Bradley–Terry rankings with bootstrap confidence intervals turn votes into leaderboards.

Judge the Judges

AI humor judges

Judge models score the exact pairs humans voted on. Agreement with the human majority (accuracy + Cohen's kappa) ranks who best understands funny.

Rankings below are provisional — built from AI judge consensus while human votes accumulate. Vote in the arena to make them real.

Model Leaderboard — who makes the funniest memes?
#ModelProviderElo (BT)95% CIComparisons
1Gemini 3.6 Flashgoogle175216921826171
2Kimi K3openrouter171816571805128
3GPT-5.6 Solopenai168016271743180
4Claude Opus 5anthropic166716131734159
5GPT-5.6 Lunaopenai153214721584192
6Gemini 3.5 Flash Litegoogle151814641583164
7Claude Sonnet 5anthropic148214261552171
8Grok 4.6xai146314151519197
9Grok 4.5xai145313941514183
10DeepSeek V3.2openrouter136813071430182
11Venice Uncensored (Dolphin Mistral 24B)openrouter867621982189
Prompt Leaderboard — which prompt is funniest?
#PromptElo (BT)Comparisons
1Comedian persona1635106
2Subversion154291
3Baseline144393
4Comedy chain-of-thought138186
Judge Leaderboard — who knows funny?

Judge scores unlock once enough human votes are in.

Caption-bar templates
Composer templates (structured slot data)

Last computed Thu, 13 Aug 2026 03:20:10 GMT