When to use DeepSeek-V4.1-Flash
Field-test, Max effort, private repo to 512k. Reading 60/64 each. Kimi 8/8 vs DeepSeek 7/8 repairs. Bill $1.70 vs $31.75
- DeepSeek-V4.1-Flash (
deepseek-flash) tied Kimi K3 at 60/64 on reading from 8k to 512k, ahead of GLM-5.3-Flash at 59/64 and GLM-5.3 at 55/64. - At 128k, Kimi's first patch passed 8/8 and DeepSeek's passed 5/8, and a second try that returned the public-test failure (grading tests still hidden) brought DeepSeek to 7/8 while Kimi stayed at 8/8.
- The scored grid cost $1.70 on DeepSeek and $31.75 on Kimi (about 19×), and a successful 128k repair cost $0.026 on DeepSeek against $0.405 on Kimi (about 16×).
- A 128k lookup took about 4.2 seconds on DeepSeek, 18.9 on Kimi, and 15.3 on GLM-5.3-Flash, and a cached follow-up on DeepSeek was about $0.0007 at roughly 99.8% prefix reuse.
- GLM-5.3-Flash reached 7/8 within two at 128k on a $1.77 bill, after a 0/8 on the earlier high-effort run, so the default is DeepSeek unless the loop needs Kimi's first-try 8/8.
DeepSeek-V4.1-Flash is a new cheap and fast model, but what's impressive is that it reads the prompt with only the first half of its layers, and uses the full model to write the reply. In a previous article, we unpacked the architecture. The question here is when to use it instead of Kimi K3 or either GLM.
I sent a synthetic codebase built for this test, not a client repo, to DeepSeek-V4.1-Flash, Kimi K3, GLM-5.3, and GLM-5.3-Flash. Each call went to that company's own API at max effort, the deepest thinking setting it offers. I ran two types of tests:
- Lookup: I asked the model questions about the code. I scaled the prompt from 8k tokens to about 512k.
- Repair: I asked the model to change the code.
One bug shows the split. The retry check lets one extra attempt through, and asking what that limit is counts as lookup. Changing the check so a hidden test passes counts as repair. The model was not shown that test.
On lookup, DeepSeek and Kimi each got 60 of 64 right. On repair, at 128k tokens, Kimi's first patch passed 8 of 8 and DeepSeek's passed 5 of 8. When the first patch failed the public tests, I sent that output back and asked for a new patch while keeping grading tests hidden. DeepSeek then reached 7 of 8. Kimi had already passed 8 of 8 on the first try.
Across the whole scored grid, DeepSeek cost $1.70 and Kimi cost $31.75. Each bill is that model's scored work: the lookup questions, the follow-ups, and every repair attempt.
My read is to default to DeepSeek, and to pay Kimi only when you need that 8 of 8 on the first try.
What follows is how a pass was scored, how both GLMs did on the same repairs, the reading table out to 512k, and when Kimi's bill is worth paying.
How DeepSeek-V4.1-Flash is built (a recap)
DeepSeek-V4.1-Flash is a 552B parameter mixture-of-experts model that splits its 40 layers in two, with weights open source under an MIT license. The first 20 layers read the prompt (about 8 billion active parameters per input token) and the full stack writes the answer (about 16 billion per generated token).
Sebastian Raschka dubbed the September 10 release a "big overhaul" and argued DeepSeek could have named it V5
Shared storage plus 4-bit compression is the memory trick, which squeezes the global key-value cache, the attention state the model keeps per token, to 890 bytes per token from 3,514 on V4-Flash. Part 1 covers the sparse-attention machinery behind that number.
How I ran the test (so you can understand the results)
We ran two test runs for this piece, so let's get a few definitions out of the way:
- a run is a whole study/test (out of 2)
- an attempt is one repair try (up to 128 in the main run)
- and max effort is each vendor's deepest thinking setting
The main run sent the same synthetic repo to four first-party APIs, DeepSeek's deepseek-flash (v4.1-flash), Moonshot's kimi-k3, and Z.ai's glm-5.3 and glm-5.3-flash on pay-as-you-go billing, out to 512k tokens. An earlier run used high effort with Kimi silently at max and stopped at 128k.
- The main run finished 424 calls across all four models: 256 reading tests, 64 cached follow-ups, and 104 repair calls out of 128 scheduled. Max effort was requested and confirmed in the logs, and the four models together cost $55.65, with no charge left unresolved. DeepSeek's share is the $1.70 above, and Kimi's is the $31.75.
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves