lolbench - LLMs take three tests:
- explain why jokes work (or don't)
- write jokes under shared premises
- predict which jokes humans prefer
The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%).
the part I need humans for: the vote booth. two clanker jokes, you pick blind, then it reveals the models + whether you agreed with thousands of other voters. ~30 seconds a ballot + a laugh (hopefully)
happy to answer anything about the eval system and open to any sort of feedback!
[link] [comments]