Project MonetRequest demo
Home/Blog/Gemini 3.8 Flash: API, Pricing, Benchmarks & Features

AI · Project Monet Briefing

Gemini 3.8 Flash: What Changed, Pricing, API and Benchmarks

Google released Gemini 3.8 Flash as a GA model for long-horizon coding, agents and complex workflows, with 1M context and introductory pricing through 2026.

Published 2026-09-03 · Updated 2026-09-04 · By Project Monet Editorial Team

Gemini 3.8 Flash editorial card showing GA status, 1M context, coding and agent workflow cues

01

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google's generally available Flash model released on September 2, 2026. Google positions it for long-horizon software engineering, autonomous agents and complex enterprise workflows.

The stable Gemini API model ID is gemini-3.8-flash. Google's current model documentation lists a 1,048,576-token input limit and a 65,536-token maximum output.

02

Release and availability

Google's Gemini API release notes mark Gemini 3.8 Flash as generally available, not preview-only. Its deprecation page currently shows no announced shutdown date.

Google also says 3.8 Flash is available in Gemini Enterprise and to Google AI Pro and Ultra subscribers across selected consumer surfaces. Account-level quotas and product-surface availability can still differ.

03

Gemini 3.8 Flash pricing

Google's introductory price through December 31, 2026 is $0.75 per million input tokens and $3.75 per million output tokens.

Google has announced standard pricing from January 1, 2027 at $1.50/M input and $7.50/M output. Teams budgeting beyond 2026 should model the higher standard rate.

04

Context window, thinking and built-in tools

Gemini 3.8 Flash accepts text, image, video, audio and PDF inputs and returns text. It supports low, medium and high thinking levels; Google lists medium as the default and says minimal is unsupported.

The model page lists support for function calling, structured outputs, code execution, file search, Google Search grounding, Maps grounding, URL context, caching and computer use in Preview. The Live API, image generation and audio generation are not supported by this model.

For implementation details, model ID, pricing modes and production checks, read the Gemini 3.8 Flash API & Pricing Guide.

05

Coding and agent capabilities

Google describes 3.8 Flash as its most intelligent Flash model and emphasizes software engineering, autonomous agents and multi-step reasoning. It also says the model is now the default for its managed Antigravity agent and Antigravity SDK.

Those claims are vendor positioning, not a guarantee for every workload. Teams should compare success rate, latency, total token use and tool-call reliability on representative tasks before migrating.

06

Benchmarks: what Google claims

Google's launch materials highlight gains on software-engineering and agent benchmarks and cite results including DeepSWE v1.1, Vals Finance Agent V2, Harvey's Legal Agent Benchmark and HLE-Verified.

Treat benchmark results with attribution. Some underlying benchmarks are externally operated, but the launch framing and model-selection claims come from Google. Production decisions should use your own evaluation set.

07

Gemini 3.8 Flash vs Gemini 3.7 Flash

Google's main argument for 3.8 Flash is stronger reasoning, coding and autonomous-agent reliability while preserving the Flash cost/speed positioning. Google also says 3.7 Flash remains supported for efficiency-first workloads.

Do not migrate solely because the version number is newer. Benchmark the same prompts, tools and failure cases on both models and compare cost per successful result.

08

Limitations and operational caveats

  • Computer use remains Preview even though the base model is GA
  • The model does not support the Gemini Live API
  • Image generation and audio generation are not supported
  • Rate limits vary by account tier and active capacity
  • Introductory pricing ends December 31, 2026 unless Google changes the schedule
  • Vendor benchmarks do not guarantee application-specific performance

09

Gemini 3.8 Flash FAQ

Is Gemini 3.8 Flash generally available? Yes. Google lists gemini-3.8-flash as GA from September 2, 2026.

What is the context window? Google's model page lists 1,048,576 input tokens and a 65,536-token maximum output.

How much does Gemini 3.8 Flash cost? Google lists $0.75/M input and $3.75/M output through December 31, 2026, then $1.50/M input and $7.50/M output from January 1, 2027.

Does it support computer use? Yes, but Google marks computer use as Preview.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept