i built a benchmark to test whether LLMs can understand and create jokes
i built a benchmark to test whether LLMs can understand and create jokes

i built a benchmark to test whether LLMs can understand and create jokes

lolbench - LLMs take three tests:

- explain why jokes work (or don't)
- write jokes under shared premises
- predict which jokes humans prefer

The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%).

the part I need humans for: the vote booth. two clanker jokes, you pick blind, then it reveals the models + whether you agreed with thousands of other voters. ~30 seconds a ballot + a laugh (hopefully)

https://lolbench.lol/

happy to answer anything about the eval system and open to any sort of feedback!

submitted by /u/AffectionateGas9544
[link] [comments]