| The headline numbers make Gemini 3.6 Flash look like a straightforward upgrade. Google says it uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index. It also reports gains on DeepSWE (49 vs. 37), MLE-Bench (63.9 vs. 49.7), OSWorld-Verified (83.0 vs. 78.4), and GDPval-AA v2 (1421 vs. 1349). The output price is also lower at $7.50 per million tokens. That is a solid aggregate story. At the same time, early screenshots in the source material claim regressions in frontend generation and spatial reasoning. The examples available there do not include original links, complete prompts, model settings, or a reproducible configuration. One example even leaves open whether the appropriate thinking setting was enabled. So I would not treat those screenshots as independent evidence that the model is broadly worse, and definitely not as evidence that it is the "worst" model overall. My read is simpler: they are enough to propose a regression case, but not enough to settle it. Aggregate gains and narrow regressions can easily coexist. A benchmark averages across its own task distribution. Your application may put most of its weight on a category that barely affects the aggregate result. A model can improve on coding agents, knowledge work, and computer use while becoming less reliable on one specific UI pattern. The overall score rises. Your product still breaks. For an actual upgrade decision, I would use a paired workload regression:
That last part seems especially important. If a team decides after the test that a preferred model's regression is "small enough," the evaluation becomes model advocacy. A predeclared gate forces the decision to follow the workload. I would also avoid forcing a single global winner. If 3.6 Flash wins on document analysis but loses on a frontend workflow, that is a routing result. Keep the incumbent for the failing category and use the new model where it clears the gate. The production unit is not just "Gemini 3.6 Flash." It is the model, settings, prompts, tools, and workload together. Official source: Google's Gemini 3.6 Flash launch post If you are evaluating 3.6 Flash in production, what specific failure gate would make you keep 3.5 Flash, or route only a subset of tasks, even if the aggregate benchmarks improve? [link] [comments] |