Gemini 3.8 Flash Review: Fast, Cheap — Until It Doubles in 2027
Gemini 3.8 Flash launched September 2 at $0.75/$3.75 per million tokens with strong agent benchmarks (DeepSWE 73.7%, Terminal-Bench 2.1 89.4%). We benchmark it, price it — including the thinking-token trap that makes real bills ~1.9x naive estimates — and flag the January 1, 2027 doubling.
Six weeks, three Flash releases. Google shipped Gemini 3.8 Flash on September 2, 2026 — and at first glance it looks like the best price-performance deal in the current model lineup. The fine print is where this review lives.
The Benchmarks: An Agent Workhorse
The agentic numbers are the story. On DeepSWE v1.1, 3.8 Flash scores 73.7% — beating its predecessor 3.7 Flash (65.3%) and GPT-5.6 Sol (72.7%), and landing a hair under Claude Opus 5 (74.0%). Terminal-Bench 2.1: 89.4%. Vals Finance v2: 61.4%, LVBench 87.8%, CharXiv 86.2%, HLE-Verified 54.9%. The trade: Harvey's Legal sits at 10.0%, and on OSWorld-2.0 it manages 59.0% against Opus 5's 75.4%; Terminal-Bench 4.0 comes in at 19.1% versus 51.8%. The picture is consistent — this is a coding-and-terminal agent, not a computer-use or legal specialist. Google's own framing is "works harder": more reasoning steps and iterative tool calls rather than one-shot brilliance.
The Pricing: Cheap, With a Countdown
List price is $0.75 per million input tokens and $3.75 per million output — with a 1M-token context window and 64K max output. That is introductory pricing: on January 1, 2027 it doubles to $1.50/$7.50. Context caching runs $0.075 per million (also promotional), and the Batch API halves everything to $0.375/$1.875. Knowledge cutoff is March 2026.
The Real Bill: Thinking Tokens Are Billed as Output
Here's the trap. 3.8 Flash's "works harder" behavior means it generates a lot of thinking tokens — and they're billed at the output rate, which is 5x the input rate. Artificial Analysis flagged the model as "very verbose" after its eval suites generated 120 million output tokens; blended cache-discounted pricing lands around $0.58 per million. For planning-heavy agent workloads, our rule of thumb: budget ~1.9x your naive estimate. The promotional price makes this easy to ignore now; the doubling makes it expensive to ignore in January.
Verdict
If you're building coding agents or terminal automation today, 3.8 Flash is the default pick — near-Opus DeepSWE performance at roughly a tenth of the output price. Just model your real token mix before Q4 ends, and re-run the math on January 1. 8.5/10, held back from 9 only because the intro pricing is a countdown timer, not a commitment.