Three models landed within three weeks of each other: GLM-5.3 on 14 August, Gemini 3.8 Flash on 2 September, and Claude Opus 5 alongside them. They span a 6.7x price range and target genuinely different jobs. What they don't share is a scoreboard โ and that turns out to be the most important thing to understand before you pick one.
The comparison, by the numbers
Vendor and third-party data, September 2026
$0.75
Cheapest input per 1M
Gemini 3.8 Flash, until Jan 2027
6.7x
Opus 5 vs Gemini Flash cost
Same multiple on input and output
1M
Context on all three
Output caps differ: 66K to 128K
0
Benchmarks covering all three
Under one common harness
The short answer
If you want the verdict without the methodology: Claude Opus 5 is the frontier pick for hard agentic and long-horizon coding work and is priced like it. Gemini 3.8 Flash is the throughput pick โ cheapest per token, fastest measured output, and strong enough for most production workloads. GLM-5.3 is the value pick, landing close to frontier coding scores at roughly a quarter of Opus 5's token cost, with an open-weight variant and a flat-rate coding subscription.
- Pick Opus 5 when correctness on a long, multi-step task is worth more than the token bill โ refactors across a large codebase, autonomous agents, research-grade reasoning.
- Pick Gemini 3.8 Flash for high-volume, latency-sensitive work: classification, extraction, summarisation, chat, anything where you're paying per request at scale.
- Pick GLM-5.3 when you want frontier-adjacent coding performance on a predictable budget, or when open weights and self-hosting matter.
Why this comparison is harder than it looks
Here is the finding that should shape how you read every number below: no single benchmark has been run on all three models under one harness. Each vendor published against its own evaluation set, with its own scaffolding, and no independent lab has re-run them side by side.
That isn't a technicality. Benchmark scores swing enormously with scaffold, prompt, retry policy, and test harness. The clearest illustration is Opus 5 itself: third-party sources report its SWE-bench Verified score anywhere from 72.5% to 96%. That is a 23-point spread on the same benchmark, for the same model, because different people ran it differently.
A number without a harness attached isn't a measurement. It's a marketing claim with decimals.
Read the tables as three separate scoreboards
The benchmark tables below are not like-for-like. Each column comes from a different source using a different harness, and dashes mean not published, not scored zero. Use them to see what each vendor chose to measure โ not to rank the models to a decimal place.
Price and specifications
Pricing is the one dimension where the three models are directly and unambiguously comparable, and the spread is wide. Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output; GLM-5.3 is $1.40 and $4.40; Claude Opus 5 is $5.00 and $25.00. All three carry a 1M-token context window.
Price and specifications
The one genuinely like-for-like comparison
| Benchmark | Gemini 3.8 Flash Google | GLM-5.3 Z.ai | Claude Opus 5 Anthropic |
|---|---|---|---|
Input per 1M tokens | $0.75 | $1.40 | $5.00 |
Output per 1M tokens | $3.75 | $4.40 | $25.00 |
Cached input per 1M All three support prompt caching | Supported | $0.26 | Supported |
Context window Identical across the board | 1M | 1M | 1M |
Max output tokens | 66K | 128K | 128K |
Open weights | No | Flash variant, MIT | No |
Released | 2 Sep 2026 | 14 Aug 2026 | 2026 |
Gemini 3.8 Flash pricing doubles in January 2027
The $0.75 / $3.75 rate is promotional through the end of 2026. In January 2027 it rises to $1.50 / $7.50 โ which puts it above GLM-5.3 on both input and output. If you are modelling annual costs, model the 2027 rate, not today's.
Coding and agentic benchmarks
Coding is where all three vendors compete hardest, and where the reporting gaps are most obvious. Z.ai reports GLM-5.3 at 77.8% on SWE-bench Verified; Google reports Gemini 3.8 Flash at 61.6% on SWE-bench Pro; third parties put Opus 5 at 96% Verified and 79.2% Pro. Only the Pro row lets you compare two models on the same test.
Coding and agentic benchmarks
Dashes mean not published โ not a zero score
| Benchmark | Gemini 3.8 Flash | GLM-5.3 | Claude Opus 5 |
|---|---|---|---|
SWE-bench Verified Opus figure is third-party, disputed | โ | 77.8% | 96% |
SWE-bench Pro The only shared coding benchmark | 61.6% | โ | 79.2% |
Terminal-Bench 3.0 Different harnesses | โ | 28.3% | 42.7% |
CyberGym | โ | 84.5% | โ |
Primary source Three different methodologies | Z.ai (vendor) | Third-party |
Two things are worth pulling out. GLM-5.3's Terminal-Bench 3.0 result of 28.3% is a jump from 4.6% on the previous generation โ a sixfold improvement that came entirely from post-training, since GLM-5.3 shares GLM-5.2's base model. And Opus 5's advantage on SWE-bench Pro over Gemini Flash is real but comes at 6.7x the token cost, which is the trade the whole comparison turns on.
GLM has a flat-rate route that skips token maths entirely
Alongside per-token API access, Z.ai sells a GLM Coding Plan from about $18/month (Lite, Pro and Max tiers, cheaper billed yearly) for use inside supported coding tools. For steady daily coding it is often far cheaper than metered tokens. See GLM plans on Z.ai.
Reasoning and science benchmarks
The reasoning picture is sparser still. Opus 5 posts the strongest published figures โ GPQA Diamond around 93% and ARC-AGI-2 at 90.4% โ while GLM-5.3 reports GPQA 68.2% and AIME 2025 at 84%, and Gemini 3.8 Flash reports HLE-Verified at 54.9%. Almost none of these overlap.
Reasoning and science
Almost no overlap between what each vendor measured
| Benchmark | Gemini 3.8 Flash | GLM-5.3 | Claude Opus 5 |
|---|---|---|---|
GPQA Diamond Google published no absolute figure | Leads Google set | 68.2% | 93.2% |
HLE Opus without tools; 64.7% with tools | 54.9% | โ | 56.3% |
AIME 2025 | โ | 84% | โ |
GSM8K | โ | 97% | โ |
ARC-AGI-2 | โ | โ | 90.4% |
What Anthropic actually published for Opus 5
This deserves its own section, because it explains the 23-point spread. Anthropic did not publish an SWE-bench Verified number for Opus 5 at all. Its launch materials reported results as relative comparisons on a different set of benchmarks entirely:
- Frontier-Bench v0.1 โ surpasses all other models, more than doubling Opus 4.8's performance.
- CursorBench 3.2 โ within 0.5% of Fable 5's peak score, at half the cost.
- ARC-AGI 3 โ roughly three times the next-best model's score.
- Zapier AutomationBench โ about 1.5x the next-best model's pass rate.
- OSWorld 2.0 โ outperforms every other model at any given cost, beating Fable 5's best result at just over a third of the cost.
Every SWE-bench figure you see quoted for Opus 5 is therefore someone else's re-run. That's why they disagree so violently, and why the 96% in the table above carries an asterisk in spirit if not in print. Anthropic also published candid weaknesses โ on OSS-Fuzz exploit development, Opus 5 sits substantially behind Mythos 5.
What Z.ai actually published for GLM-5.3
Z.ai's figures are vendor-reported and have not been independently re-run under a neutral harness, which is the standard caveat for any first-party benchmark claim. Treat them as a best case rather than a settled result.
What is independently checkable is the release itself. GLM-5.3 arrived on 14 August 2026 built on GLM-5.2's base model, with all gains from post-training. Twelve days later Z.ai shipped GLM-5.3-Flash โ the series' first natively multimodal model, handling text, image and video at 320B parameters with 18B active, released under the MIT licence. That combination of permissive licensing and a leaner multimodal model is the genuinely differentiated part of the offering, and no benchmark caveat touches it.
Gemini 3.8 Flash: speed as the product
Gemini 3.8 Flash's headline claim isn't a benchmark โ it's throughput. At roughly 305 tokens per second, it posts the fastest measured output speed of the three, and it is Google's third Flash release in about six weeks.
The benchmark gains over 3.7 Flash are modest: SWE-Bench Pro moved from 60.4% to 61.6%, barely more than a point. This is an iteration on speed and price, not a leap in capability, and Google positions it accordingly as a workhorse rather than a frontier model. Its 66K output cap is also the tightest of the three, which matters if you generate long documents or large diffs in a single call.
What you will actually pay
Benchmarks are abstract; token bills are not. Take a realistic long-context task โ 100,000 input tokens and 10,000 output tokens, roughly a large-codebase question or a long document analysis:
- Gemini 3.8 Flash โ $0.075 in, $0.038 out = $0.11 per task
- GLM-5.3 โ $0.14 in, $0.044 out = $0.18 per task (about $0.07 if the input is cached)
- Claude Opus 5 โ $0.50 in, $0.25 out = $0.75 per task
Run a thousand of those a month and you're looking at roughly $112, $184, and $750 respectively. That framing usually settles the argument faster than any leaderboard: Opus 5 has to be about seven times more useful than Gemini Flash on your specific workload to justify itself, and on genuinely hard agentic tasks it often is. On classification and extraction, it very obviously isn't.
Prompt caching changes the maths substantially for all three, and GLM-5.3's published cached-input rate of $0.26 per million is the most aggressive of the group. If your workload replays a large stable prefix, measure with caching on before comparing anything.
Latency and throughput
Cost per token is only half the operating picture; tokens per second is the other half, and it is where Gemini 3.8 Flash makes its strongest case. At roughly 305 tokens per second it posts the fastest measured output of the three, which compounds into a materially different user experience on anything streaming to a person in real time.
Opus 5 answers this with an optional Fast Mode โ a research preview that runs the same model at up to 2.5x higher output speed, priced at $10 per million input and $50 per million output. That is double standard Opus 5 pricing, or roughly 13x Gemini Flash's current rate, so it buys latency rather than economy. It is also restricted to Anthropic's own API, not the cloud partner platforms.
GLM-5.3 sits between them on published throughput and competes on a different axis entirely: because the Flash variant ships under an MIT licence, latency becomes an infrastructure decision you control rather than a vendor characteristic you accept.
Run your own eval in an afternoon
Given that no shared harness exists, the highest-value thing you can do is stop reading leaderboards and measure the three on your own work. Fifty representative tasks will tell you more than every table in this article. The process is genuinely a single afternoon:
- Pull fifty real requests from your logs โ not hand-written test cases. Real inputs carry the messiness that separates models.
- Write the grading rule before you run anything. Exact match, a rubric, or a stronger model as judge. Deciding what counts as correct after seeing outputs is how you talk yourself into a conclusion.
- Run all three at the same effort and temperature settings, with the same prompt and the same tools. This is the harness discipline the public benchmarks lack.
- Record cost and latency per task alongside accuracy. A model that wins on quality but costs seven times more has not necessarily won.
- Judge cost per completed task, not per request. A cheaper model that needs three attempts to get there is not cheaper.
Prompt caching is the one setting that will most distort the result if you leave it off. All three support it, and GLM-5.3's published cached-input rate of $0.26 per million is aggressive enough to change which model wins on a workload that replays a large stable prefix. Measure with caching configured the way you will actually run it.
Which should you pick?
Stop at the first line that matches your workload:
- High-volume, latency-sensitive, cost-dominated โ Gemini 3.8 Flash. Cheapest and fastest, until the January 2027 price change.
- Daily coding on a predictable budget โ GLM-5.3 via the coding plan. Frontier-adjacent scores without metered billing anxiety.
- Open weights, self-hosting, or permissive licensing โ GLM-5.3, specifically the MIT-licensed Flash variant.
- Hard agentic work where a wrong answer is expensive โ Claude Opus 5. The price premium buys reliability over long horizons, which is exactly what Anthropic's launch benchmarks measure.
- You genuinely can't decide โ run your own eval. Given that no shared harness exists, fifty representative tasks from your actual workload will tell you more than every table above combined.
Methodology and disclosure
Pricing, context and output limits come from each vendor's published documentation as of September 2026. Anthropic's Opus 5 figures are its own launch materials; third-party Opus 5 benchmark numbers are re-runs and are flagged as such. GLM-5.3 benchmark figures are vendor-reported by Z.ai and have not been independently verified. Gemini 3.8 Flash figures are Google-published. Cost-per-task and cost-multiple calculations are ours. Disclosure: links to Z.ai are affiliate links; Google and Anthropic links are not, and no vendor reviewed this article.
The verdict
These three models aren't really competing for the same job, which is why the absence of a shared benchmark matters less than it first appears. Opus 5 is built for long-horizon autonomy, Gemini 3.8 Flash for volume, GLM-5.3 for value and openness. The pricing table separates them far more honestly than any leaderboard currently can.
If you take one thing away, make it this: the 6.7x cost gap is a hard, verified number, and the benchmark gaps are not. Build your decision on the number you can trust, then validate it against your own tasks.
API from $1.40 per 1M input tokens ยท Coding plans from about $18/month
Frequently Asked Questions
8 questions answered
Enjoyed this article?
Share it with someone who'd love it.
Written by
AI Magic Editorial Team
We write about AI image generation, creative workflows, and how creators use AI Magic to ship faster โ built on the latest from Google Gemini.