model · model/google-gemini-3-8-flash
Google: Gemini 3.8 Flash
Also called Gemini 3.8 Flash, google/gemini-3.8-flash
Right now, this row costs exactly what its predecessor does: it lists at $0.75per million tokensopenrouter-models, last checked 2026-09-14 input and $3.75per million tokensopenrouter-models, last checked 2026-09-14 output, the same figures 3.6 and 3.7 carry, because Google carried the introductory window across the release. The launch post dates the window plainly: 2026-12-31; US$1.50/US$7.50 standard rate applies from 2027-01-01source, accessed 2026-09-03.
What changed is the invoice per task. Google's release note credits the gains to "a core design choice: 3.8 Flash works harder" — extra reasoning steps, iterative tool calls, and at higher effort levels, more tokens spent. An independent reading of the release, llm-releases, turns that sentence into arithmetic: average output tokens per task rose ~30%, to ~48ksource, accessed 2026-09-03, lifting cost per task to ~$0.58source, accessed 2026-09-03 at high reasoning despite the unchanged per-token price. Buyers of this row are paying the same rate and a larger bill.
Google's own launch numbers for the row are vendor-reported and dated to the release: HLE-Verified at 54.9%source, accessed 2026-09-03, Vals Finance Agent v2 at 61.4%, against 59.0% for Gemini 3.7 Flashsource, accessed 2026-09-03, Harvey's Legal Agent Benchmark at 10.0%, against 8.8% for Gemini 3.7 Flashsource, accessed 2026-09-03, and on DeepSWE v1.1, 3.8 Flash outperforms most larger frontier models at a fraction of the costsource, accessed 2026-09-03. The row keeps the family-standard 1048576tokensopenrouter-models, last checked 2026-09-14 input window with 65536tokensopenrouter-models, last checked 2026-09-14 output, and Google dates its knowledge cutoff to March 2026 for some domains; January 2025 for otherssource, accessed 2026-09-03.
The by-effort split recorded below is a pre-rebase reading. It was read from llm-releases on 3 September 2026, on the Artificial Analysis intelligence index as it then stood; the index has since been rebased to v4.2, which moved this row's score down along with every other row that kept a score through it. Those effort numbers and the 41.2openrouter-models, last checked 2026-09-14 the catalog carries today are not on the same scale.
The same launch carried a second model. 3.8 Flash Cyber is the cybersecurity-tuned sibling, available only to trusted defenders — named by llm-releases as government authorities, critical-infrastructure operators and software maintainers — through Google's new Fairwind Program, and it is not publicly token-billed. Google's claims for it, all vendor-reported, are in the launch post: frontier-level performance on the CyberGym vulnerability-discovery benchmark, surpassing 3.5 Flash Cyber and much larger frontier models; a success rate exceeding 70% on an internal real-world benchmark spanning 20 programming languages; 47.2% pass@1 on CWE-Bench, on the Pareto frontier against a leading frontier model's 47.8% at far lower cost. The deployment numbers are Google's too: Chrome Security reports 2.6 times more correct patches than the best commercial models that are much larger; Wiz measured +7.5-9.7% higher recall on its internal penetration-testing benchmark at 2.3-5.2x lower cost than other leading frontier models; and Cloud Vulnerability Research used the model to find a critical foundational vulnerability in under two hours, where the same discovery usually takes months.
Facts
- price input
- $0.75per million tokensopenrouter-models, last checked 2026-09-14
- price output
- $3.75per million tokensopenrouter-models, last checked 2026-09-14
- context window
- 1048576tokensopenrouter-models, last checked 2026-09-14
- max output tokens
- 65536tokensopenrouter-models, last checked 2026-09-14
- intelligence index
- 41.2openrouter-models, last checked 2026-09-14
- status
- activeopenrouter-models, last checked 2026-09-14
- release date
- 2026-09-02source, accessed 2026-09-03
- introductory pricing ends
- 2026-12-31; US$1.50/US$7.50 standard rate applies from 2027-01-01source, accessed 2026-09-03
- intelligence index by effort
- 59 at high reasoning, 57 at medium, 52 at low; up 3 from Gemini 3.7 Flash at high, level with sub-maximum efforts of GPT-5.6 Sol and Grok 4.6source, accessed 2026-09-03
- hle verified
- 54.9%source, accessed 2026-09-03
- vals finance agent v2
- 61.4%, against 59.0% for Gemini 3.7 Flashsource, accessed 2026-09-03
- harveys legal agent benchmark
- 10.0%, against 8.8% for Gemini 3.7 Flashsource, accessed 2026-09-03
- deepswe v1 1
- outperforms most larger frontier models at a fraction of the costsource, accessed 2026-09-03
- output tokens per task
- ~48ksource, accessed 2026-09-03
- cost per task
- ~$0.58source, accessed 2026-09-03
- knowledge cutoff
- March 2026 for some domains; January 2025 for otherssource, accessed 2026-09-03
Timeline
- released with introductory pricing declared to run through 2026-12-31source