i built a benchmark to test whether LLMs can understand and create jokes
lolbench – LLMs take three tests: – explain why jokes work (or don't) – write jokes under shared premises – predict which jokes humans prefer The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on ex…