As of August 2026, DeepSeek V4 Pro is the cheapest of the three at roughly $0.44 per million input tokens and $0.87 output. Gemini 3.7 Flash sits in the middle at $0.75 and $3.75. GPT-5.6 Sol is the premium option at $5 and $30. The three target different tiers, so match the model to the task rather than the headline price.
Three model updates landed inside a three-week window, and every headline framed it as a price war. That framing is half right. Gemini 3.7 Flash and DeepSeek V4 Pro are competing directly on cost-per-token for high-volume work. GPT-5.6 Sol is not in that fight at all — it is a maximum-effort reasoning flagship priced accordingly. Reading all three as interchangeable is how teams end up paying Sol rates for work a Flash-class model does at a fraction of the cost. This is a comparison of what each one actually charges, what you get for it, and how to choose without overpaying.
The three prices at a glance
Here is the verified pricing as published by each provider, checked on 17 August 2026. Every figure below is per one million tokens.
| Model | Input | Output | Context | Max output | Source |
| DeepSeek V4 Pro | $0.435 | $0.87 | 1M tokens | 384K | DeepSeek (2026) |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1M tokens | 65K | DataCamp (2026) |
| GPT-5.6 Sol | $5.00 | $30.00 | 1M tokens | 128K | Artificial Analysis (2026) |
The spread on output tokens is the number that matters most, because agentic and coding workloads generate far more output than input. DeepSeek V4 Pro's $0.87 output rate is roughly 34 times cheaper than GPT-5.6 Sol's $30. Gemini 3.7 Flash sits between them, closer to DeepSeek on input and far below Sol on output — $3.75 is eight times cheaper than Sol's rate.
[Inline SVG bar chart in HTML] Output price per million tokens, drawn to scale:
| Model | Output $/1M |
| DeepSeek V4 Pro | $0.87 |
| Gemini 3.7 Flash | $3.75 |
| GPT-5.6 Sol | $30.00 |
Caption: Output price per million tokens, drawn to scale. Prices per DeepSeek, DataCamp and Artificial Analysis (2026).
Why these three are not the same product
The cleanest way to think about the lineup is by tier, not by brand. DeepSeek V4 Pro and Gemini 3.7 Flash are both cost-optimised models built to run at volume — the kind you point at a pipeline that processes thousands of requests an hour. GPT-5.6 Sol is OpenAI's maximum-effort reasoning model, built for the hardest single problems where being right matters more than being cheap.
OpenAI itself makes this split explicit. Alongside Sol at $5/$30, it ships GPT-5.6 Terra at $2.50 input and $15 output, and GPT-5.6 Luna at $1 and $6, according to OpenAI's own model announcement (2026). Luna, not Sol, is the tier that competes with Gemini 3.7 Flash on price. Comparing Flash to Sol is comparing a delivery van to a race car and concluding the van is better value.
Worth knowing: Google measures Gemini 3.7 Flash as beating Claude Sonnet 5 and GPT-5.6 Terra on its coding benchmarks — while charging less than either. That is the specific claim driving the price-war coverage, and it is Google's own measurement, not an independent one.
What each model costs, and the fine print that changes it
The sticker prices above are current, but two of the three carry conditions that will move your bill.
Gemini 3.7 Flash's $0.75/$3.75 is introductory pricing that Google has committed to only through 31 December 2026. From 1 January 2027 the standard rate rises to $1.50 per million input and $7.50 per million output, per DataCamp's breakdown of Google's pricing (2026). Anything you budget on Flash today doubles at the turn of the year unless Google extends the promotion.
DeepSeek V4 Pro moved the other way. Its current $0.435/$0.87 reflects a permanent cut of roughly 75% made on 31 May 2026, down from the $1.74/$3.48 the model launched at in April, per DeepSeek's pricing page (2026). That page also lists a cache-hit input token at a fraction of the standard input rate, which makes repeated-context workloads — retrieval pipelines, long system prompts — much cheaper than the headline figure suggests.
Watch out: A widely repeated claim this month said DeepSeek raised V4 Pro prices by up to 1,100%. DeepSeek's own pricing page (2026) shows the opposite — a cut, not a hike. Check the provider's page before you budget around a secondhand number.
GPT-5.6 Sol's $5/$30 has no promotional asterisk, which is its own kind of signal: OpenAI is not trying to win the volume tier with Sol. It is selling reasoning quality, and the price is the filter.
Coding benchmarks: what the money buys
Price is only half of value; the other half is capability, and the three separate clearly on coding — the workload most teams are actually buying for.
Gemini 3.7 Flash posts the headline gains. It scores 43.6% on FrontierCode, up from 34.4% for the previous Flash release, and 65.3% on DeepSWE against 49.0% before, according to The Decoder's report on the launch (2026). Those are large jumps for a model that also halved its price. On Terminal-bench 2.1 it reaches 85.8%.
GPT-5.6 Sol competes on general reasoning rather than cost. It scores 61 on the Artificial Analysis Intelligence Index, ranking fifth of more than 1,800 evaluated models, per Artificial Analysis (2026). OpenAI reports Sol setting a new record on Terminal-Bench 2.1 in its announcement (2026). Where a task genuinely needs the strongest available reasoning — a gnarly refactor, a multi-step agent that cannot afford to go wrong — Sol earns its premium. Where it does not, you are paying about 34 times the output rate for headroom you will not use.
DeepSeek V4 Pro publishes less benchmark detail than the other two, which is itself worth noting: its case is almost entirely price and context. At $0.435/$0.87 with a one-million-token window and a 384K-token maximum output, it is built for high-volume throughput rather than benchmark-topping single answers.
Stat: 43.6% — Gemini 3.7 Flash's FrontierCode score, up from 34.4% for the prior release, per The Decoder (2026).
Speed: read the GPT-5.6 Sol number carefully
Speed is where a headline figure can mislead. OpenAI cites GPT-5.6 Sol running at up to 750 output tokens per second on a Cerebras deployment, per OpenAI (2026), a fast number that drove much of the launch coverage. But that is a specialised deployment figure. Measured on the standard API, Artificial Analysis records Sol at 73.7 output tokens per second, per its performance page (2026).
Both numbers are real; they describe different things. If you are evaluating throughput for a latency-sensitive product, benchmark the endpoint you will actually call rather than the peak figure in the announcement. The gap between the two is the difference between a hardware showcase and a production default.
Comparing Flash to Sol on price is comparing a delivery van to a race car and concluding the van is better value.
Which one should you use?
Match the model to the shape of the work, not to the lowest number on the page. The table below maps common situations to a default choice.
| If your workload is... | Default choice | Why |
| High-volume, cost-sensitive, repeated context (RAG, batch processing) | DeepSeek V4 Pro | Cheapest output at $0.87/M plus a discounted cache-hit input rate |
| Coding agents and dev tooling at scale | Gemini 3.7 Flash | Strong coding scores at $3.75 output; watch the January 2027 price rise |
| Hardest reasoning, correctness over cost | GPT-5.6 Sol | Top-five intelligence index; use where a wrong answer is expensive |
| General production traffic on a budget | GPT-5.6 Luna or Gemini 3.7 Flash | Luna at $1/$6 is OpenAI's real volume tier, not Sol |
The most common mistake we see teams make is standardising on a single flagship for everything. A routing layer that sends easy requests to a Flash- or V4-class model and escalates only the hard ones to a Sol-class model will usually cut a bill by more than half without a quality loss anyone notices. For the broader task-by-task framework across a wider model set, see our August 2026 LLM price-war breakdown; this post is the narrower head-to-head on the three models that moved this week.
Is this actually a price war?
On the volume tier, yes. Gemini 3.7 Flash undercutting its own three-week-old predecessor by 50%, and DeepSeek V4 Pro's 75% cut, are both real downward moves aimed at the same cost-sensitive buyer. On the frontier tier, no — GPT-5.6 Sol held its price because its buyer is choosing capability, not cost. The useful read is that the market has split into two races: one on price for high-volume work, one on reasoning quality for the hard problems. Deciding which race you are running is the whole decision. When teams get their model bill wrong, it is almost always because they entered the wrong race.
The bottom line
Pick the tier before you pick the brand. If the work is high-volume and cost is the constraint, DeepSeek V4 Pro is the cheapest and Gemini 3.7 Flash is the strongest coder — with a price increase pencilled in for January 2027. If the work is a hard reasoning problem where a wrong answer is costly, GPT-5.6 Sol is worth its premium and the cheaper tiers are a false economy. The teams that get their AI bill right in 2026 are not the ones chasing the lowest number; they are the ones routing each request to the cheapest model that can still do the job. If you want help designing that routing layer, our AI development team builds it into production systems, and you can hire AI developers to run the integration end to end.
The most common mistake we see teams make is standardising on a single flagship for everything. A routing layer that sends easy requests to a Flash- or V4-class model and escalates only the hard ones to a Sol-class model will usually cut a bill by more than half without a quality loss anyone notices. For the broader task-by-task framework across a wider model set, see our August 2026 LLM price-war breakdown; this post is the narrower head-to-head on the three models that moved this week.
FREQUENTLY ASKED QUESTIONS (FAQs)
