All posts
AI

DeepSeek's V4.1 Flash Beat Kimi K3 in Days, Not Months

ZKO blog tile: AI

DeepSeek's V4.1 Flash edged past Moonshot AI's Kimi K3 on independent Terminal-Bench 2.1 runs this week, 90.6 to 88.3. Kimi K3 held the open weight coding benchmark lead for barely five weeks after its own release. That is the real story, not the half a point gap between two model names most people outside AI research will forget by Christmas.

Benchmark leadership among open weight Chinese labs is now turning over faster than most teams can update their internal recommendations. We saw the same pattern through August: Moonshot's Kimi K3 launched claiming the largest open weight model yet, DeepSeek answered with V4-Pro within three weeks, and now V4.1 Flash has leapfrogged again. Any blog post or internal wiki page that names a specific model as the best open weight option is stale within a month.

The practical response is to stop treating which model as a decision you make once. We pick models per task based on current pricing and benchmark results, re check quarterly at minimum, and keep our integration layer thin enough that swapping a backend model is a config change, not a rebuild. The lesson from this particular week is not that DeepSeek is now the best open weight lab. It is that whichever lab holds that title will not hold it for long, and building around any single model as a permanent choice is already the wrong bet.