Static LLM benchmarks can miss performance changes over time – observations from 31,352 repeated measurements
Public LLM benchmarks answer a useful question: how well did this model perform on this benchmark at the time it was tested? What they usually don't tell us is whether the same model continues behaving the same way days or weeks later. This becomes…