By Thorsten Meyer

On the Frontend Code Arena board of 1 August 2026, an MIT-licensed model sits nine points behind the second-best model on the board — at roughly one fifteenth of its price.

The gap to the very top is larger: 128 points, or 7.5 percent. The price gap to the very top is larger still: roughly eighty-two to one.

The model is DeepSeek-V4-Flash-High. The rating is one day old, rests on 1,319 votes, and is marked preliminary by Arena itself. That caveat governs everything that follows — and the argument survives it anyway, because the argument is not about the ninth point. It is about the shape of the curve.

AI DISPATCH · REALITY CHECK Arena board of 1 Aug 2026
DeepSeek-V4-Flash-High on the Frontend Code Arena
The Ninth Point

An MIT-licensed mixture-of-experts sits nine points behind the second-best model on the board at roughly one fifteenth of its price — and 128 points behind the leader at roughly one eighty-second. The rating is one day old and marked preliminary. The shape of the curve is the story anyway.

▲ Preliminary rating · ±18 · 1,319 of 510,194 votes
1577
Arena score, preliminary
$0.25
Blended per million tokens
284B / 13B
Total / active parameters (MoE)
MIT
Licence — commercial use, no strings
01
The frontier, drawn to scale

Six models nothing else beats on both score and price at once. The horizontal axis is logarithmic — every gridline is roughly a tenfold price increase.

$0.01 $0.10 $1.00 $10 / M blended 1200 1400 1600 1800 granite-4.1-8b 1194 laguna-xs.2 1304 deepseek-v4-flash-high 1577 · $0.25 glm-5.2-max 1586 kimi-k3-max 1676 claude-opus-5-max 1705 +9 pts · ~15× price
SOURCE: ARENA.AI FRONTEND CODE ARENA, OVERALL BOARD, 108 MODELS, 1 AUG 2026 · LOG PRICE AXIS · DEEPSEEK ROW PRELIMINARY · POSITIONS APPROXIMATE
laguna-xs.2 → deepseek-v4-flash-high
+ ~$0.07 / MMARGINAL PRICE
+273 ptsSCORE GAINED
deepseek-v4-flash-high → glm-5.2-max
~15× the rateMARGINAL PRICE
+9 pts · 0.57%SCORE GAINED
deepseek-v4-flash-high → claude-opus-5-max
~82× the rateMARGINAL PRICE
+128 pts · 7.5%SCORE GAINED
02
What moved on 31 July: post-training, nothing else

Both checkpoints sit on the board simultaneously — a rare clean record of what re-post-training alone is worth on frozen weights at a frozen price.

deepseek-v4-flash-high-preview
CHECKPOINT 0420 · 24 APR 2026
1432
  • Original public release
  • Chat Completions API
+145
on the live board
deepseek-v4-flash-high
CHECKPOINT 0731 · 31 JUL 2026
1577
  • Re-post-trained for agentic work
  • Native Responses API, Codex-adapted
  • MIT weights on Hugging Face, DSpark module attached
Unchanged between the two rows: 284B/13B MoE architecture · 1M context · 384K max output · $0.14 in / $0.28 out / $0.0028 cache-hit · the licence
03
The caveat that governs everything

Arena reports a conservative rating — mu minus three sigma — and the row is one day old. The bias cuts both ways.

Preliminary flag
1,319 votes. 0.26% of the board. ±18 stated uncertainty.

Nothing here should be read as a settled ranking. The durable claim is narrower: at the price actually published, a model of this class being on the frontier at all is the fact worth recording.

Why 1577 may rise
Three standard deviations are subtracted before reporting. A thin row is deliberately printed below its central estimate — a floor, if the model keeps winning.
Why 1577 may fall
A thin sample is a noisy one. A run of favourable early pairings inflates the central estimate itself, and no conservative offset corrects a mu that is wrong.
04
Bull and bear, for a local-first operator

A 284B MoE with 13B active, expert weights in FP4, is approximately the shape of model that already runs on high-memory Apple silicon.

Bull
  • MIT means MIT. Commercial use, modification, redistribution — no bespoke licence to interpret, no acceptable-use policy to monitor.
  • Runnable in principle. FP4 experts and 13B-active sparsity put per-token compute near a mid-size dense model, within reach of a 512GB unified-memory machine.
  • Post-training is the cheap lever. +145 points on frozen weights signals more gains of this kind, from every open-weight lab.
Bear
  • Vendor benchmarks are vendor benchmarks. Terminal-Bench, Cybergym and DeepSWE numbers come from DeepSeek’s own harness; agent scores are harness-sensitive.
  • One task family. Frontend code voting is not a general capability measure, and sub-boards disagree with the Overall board.
  • Self-hosting buys sovereignty, not savings. At $0.25 per million blended, the hosted API undercuts your own electricity and depreciation for most workloads.
The ninth point costs fifteen times the price. The last 128 cost eighty-two times.
For the first time, the model asking the question carries an MIT licence.

What the model is

V4-Flash is a sparse mixture-of-experts model: 284 billion parameters in total, roughly 13 billion active per token — about one twenty-second of the network firing on any given step. Context runs to one million tokens; a single output can run to 384,000.

The published API price is $0.14 per million input tokens, $0.28 per million output, and $0.0028 per million on a cache hit. Blended for the Arena workload, that lands around $0.25 per million.

"High" is not a separate model. It is a reasoning-effort setting on the same checkpoint — which is why the listed price is identical to the lower-effort tiers, and why the real cost per task is not. More effort means more reasoning tokens, and reasoning tokens bill as output. The price card stays flat; the invoice does not.

The weights are MIT. That is worth pausing on, because it is a materially stronger grant than the "open weights" label carried by several neighbours on the board. MIT permits commercial use, modification, and redistribution with no bespoke licence to interpret and no acceptable-use policy to monitor for changes. For anyone building sovereign or local-first infrastructure, the licence is a spec sheet line, not a footnote.

Laplink PCmover - Easy Migration of your Applications, Files and Settings from an Old PC to a New PC - Data Transfer Software with Optional Super Speed USB 3.0 Cable - Business Standard, 10 Licenses
  • Licensing Options: Flexible tiers for 1, 5, 10, or 25 transfers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The arithmetic of the frontier

A Pareto frontier is the set of options nothing else beats on both axes at once. Plot Arena score against blended price across the 108 models on the Overall board, and the frontier runs, cheapest to dearest: granite-4.1-8b at 1194, laguna-xs.2 at 1304, deepseek-v4-flash-high at 1577, glm-5.2-max at 1586, kimi-k3-max at 1676, claude-opus-5-max at 1705.

The interesting feature is the size of the step into DeepSeek and the size of the step out of it.

Stepping up from laguna-xs.2 costs about seven cents per million blended and buys 273 points. Stepping up from DeepSeek to glm-5.2-max costs roughly fifteen times as much per million and buys nine points — 0.57 percent. Going all the way to the top buys 128 points for roughly eighty-two times the blended rate.

To be clear about what this is not: it is not a claim that the cheap model is as good as the expensive one. A 128-point TrueSkill gap on half a million votes is a real gap, the sub-boards do not all agree with the Overall board, and frontend code voting is one task family, not a general capability measure.

The claim is narrower. In one specific band, the price axis and the quality axis have come badly out of proportion. Above the DeepSeek point, you are no longer buying capability at anything like the exchange rate that applied below it. Whether the last 7.5 percent is worth eighty-two times the rate depends entirely on the task — and for a meaningful class of tasks, it plainly is. But the pricing of that band is now set against an MIT-licensed floor, and floors like that do not politely go away.

Large Language Model-Based Solutions: How to Deliver Value with Cost-Effective Generative AI Applications

Large Language Model-Based Solutions: How to Deliver Value with Cost-Effective Generative AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What moved on 31 July — and why it is the real story

V4-Flash first shipped on 24 April 2026, the checkpoint usually called 0420. On 31 July, DeepSeek shipped 0731: by its own account a re-post-training of the same architecture, at unchanged pricing, adding native support for the OpenAI Responses API and compatibility with Codex-style coding clients. The official weights landed on Hugging Face the same day, shipping with the DSpark speculative-decoding module attached — which is why the repository reports 304B parameters against the 284B base.

No new parameters. No new price. No new context window. What moved was post-training — and the board happens to record the move unusually cleanly, because both checkpoints sit on it simultaneously. The unversioned deepseek-v4-flash-high row reads 1577. The -preview row, which is the April checkpoint, reads 1432.

Same parameter count, same architecture, same price card. +145 points, visible on one leaderboard on the same day.

Arena's own announcement put the jump at +154 against its launch-day figure for the older checkpoint; the live board shows +145 between the rows. The likeliest reconciliation is that the older row's rating has itself drifted as votes accumulated — a nine-point difference that is, usefully, a live demonstration of how soft these numbers are at the margins.

The strategic reading matters more than the number. For most of this cycle, the assumption has been that capability jumps require new models — more parameters, new architectures, another training run priced in the hundreds of millions. A 145-point Arena move from post-training alone, on frozen weights, at frozen prices, says the binding constraint in this band is no longer the network. It is what you do to the network after pre-training ends. That lever is dramatically cheaper than the alternative, and every open-weight lab now knows it.

End-to-End AI Evaluation: Building Effective Metrics, Pipelines, and Monitoring for LLM Systems

End-to-End AI Evaluation: Building Effective Metrics, Pipelines, and Monitoring for LLM Systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The caveat, taken seriously

Arena marks the row Preliminary: ±18 stated uncertainty, 1,319 votes, about 0.26 percent of the 510,194 votes on the board. Every figure above inherits that flag, and the direction of the bias is genuinely unclear — it cuts both ways.

Arena reports a conservative TrueSkill rating, mu minus three sigma: three standard deviations are subtracted, so a thin row is deliberately reported below its central estimate. On that reading, 1577 is a floor that should rise as votes arrive, if the model keeps winning.

But a thin sample is also a noisy one. A run of favourable pairings early on inflates the central estimate itself, and mu minus three sigma cannot correct for a mu that is wrong. The honest statement is that 1577 is a one-day-old number on a young row and may move in either direction. The claim that survives the caveat is the durable one: at the price actually published, a model of this class being on the frontier at all is the fact worth recording.

Pick Your Brain: How Claude, ChatGPT, Gemini and the rest compare — and how to choose (Fluent in AI Book 2)

Pick Your Brain: How Claude, ChatGPT, Gemini and the rest compare — and how to choose (Fluent in AI Book 2)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What it means if you run your own hardware

Here the abstract arithmetic becomes concrete, because a 284B mixture-of-experts with 13B active per token is not a hypothetical for a local-first operation — it is approximately the shape of model that already runs on high-memory Apple silicon.

The V4 family stores expert weights in FP4 with most other parameters in FP8, which compresses the memory footprint well below what the headline parameter count suggests. A 512GB unified-memory machine is in the conversation for a model of this class; the 13B-active sparsity means per-token compute is closer to a mid-size dense model than to anything with 284B in the name. The MIT licence means nothing stops you from doing it, commercially, tomorrow.

The bull case for the open-weight, local-first position writes itself: the cheapest point on the useful part of the frontier is now a model you can download, modify, fine-tune, and serve yourself, with a licence that will not change underneath you.

The bear case deserves equal weight. Vendor-run benchmarks — Terminal-Bench, Cybergym, DeepSWE numbers published alongside the release — are DeepSeek's own harness, and agent scores are notoriously harness-sensitive. At least one early report claimed the 0731 weights had not yet superseded the preview repository, before the official repo was confirmed; version confusion of exactly this kind is why the Arena board carries an unversioned row and a -preview row for what a casual reader would take to be one model. And "you can run it" is not "you should": a hosted API at $0.25 per million blended is, for most workloads, cheaper than the electricity and depreciation of serving 284B parameters yourself. Self-hosting buys sovereignty and data control. It does not, at these prices, buy savings.

The dispatch in one line

The ninth point between DeepSeek and the model above it costs fifteen times the price. The 128 points to the top cost eighty-two times. Somewhere in that spread, every builder now has to decide what capability is actually worth — and for the first time, the model asking the question carries an MIT licence.

We will restate the figures when the row loses its preliminary flag.


Sources: Arena.ai Frontend Code Arena, Overall board as of 1 August 2026 (deepseek-v4-flash-high, preliminary, ±18, 1,319 of 510,194 votes); DeepSeek API changelog and Hugging Face model card for DeepSeek-V4-Flash-0731, 31 July 2026; MarkTechPost and Fireworks AI release coverage confirming the 284B/13B architecture, DSpark module, and MIT-licensed weights; contemporaneous reporting on Responses API and Codex compatibility. Ratings marked preliminary by Arena; vendor benchmark figures are DeepSeek's own harness. Point-in-time as of 3 August 2026. Not purchasing advice.

You May Also Like

AI Is the Alibi. The Reorg Is the Signal.

In May, Coinbase cut about 700 people — 14% of its staff…

One Transformer, Sound Included: What MiniMax H3 Actually Ships — and What “Open” Means This Time

By Thorsten Meyer The interesting thing about MiniMax H3 is not that…

Cognition Acquires Windsurf: Background and Analysis

Cognition’s recent acquisition of Windsurf in July 2025 marked the climax of…

AI’s Role in Health Insurance Coverage Choices

Explore how artificial intelligence is reshaping the way health insurance coverage is determined and what it means for you.