G
GOOGLE / gemini-3.8-flash

Gemini 3.8 Flash

Flash model for long-running software engineering, autonomous agents, and complex business flows.

01 / SUMMARY

The essential

ORIENTATIONFast multimodal

Primary use stated or inferred cautiously from official documentation.

CONTEXT1,048,576 entry tokens

Maximum input capacity when published by the source.

EXIT65,536 output tokens

Documented maximum generation limit.

ACCESSGemini API · Google AI Studio · stable version

Verified availability channels.

02 / TECHNICAL SHEET

Limits and integration

API IDgemini-3.8-flash
ReleaseSeptember 2, 2026
Knowledge cutoffNot published in the source consulted
EntranceText · Image · Video · Audio · PDF
ExitText
Context window1,048,576 entry tokens
maximum output65,536 output tokens
Reasoninglow, medium, and high levels
Published toolsFeatures · Web search · Google Maps · Code execution · Computer use (preview) · File search · URL context
Structured OutingsCompatible
Batch processingCompatible · 50% discount
Prompt cacheCompatible
Fine-tuningNot published in the source consulted
Verified platformsGemini API · Google AI Studio
BEST SUITED FOR

Low-cost agents, long-running software, and broad multimodal analysis.

WORTH MONITORING

The price is introductory, and using grounding may incur per-query charges.

03 / CAPABILITIES

What it can do

01

Freelance agents

Planning and orchestrating tools into multi-step objectives.

02

Code

Aimed at extensive software engineering and extensive refactorings.

03

Multimodality

Supports text, image, video, audio and PDF as input.

04

Adjustable reasoning

It offers low, medium and high levels to adjust the effort.

Input modalities
TextImageVideoAudioPDF
04 / PRICES

Documented cost

ConceptWorthUnit/condition
Standard input0.75 USD / 1M tokensStandard API rate
Cached input0.075 USD / 1 M tokensReading reused prefixes
Cache write or storage0.50 USD / 1 M tokens per hour of storageThe condition varies by provider
Standard output3.75 USD / 1M tokensMay include reasoning tokens
Batch input0.375 USD / 1 M tokensAsynchronous processing
Batch output1.875 USD / 1 M tokensAsynchronous processing

Introductory pricing until December 31, 2026; from 2027, standard rates double.

i Prices change and may depend on level, region or context length. Check the source before making a decision.

05 / EVALUATIONS

How to read the results

Google describes it as its smartest Flash model for code and agents. Prices shown are introductory until December 31, 2026.

Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.

06 / VERSIONS

Chronology

Gemini 3.8 Flash

General availability and stable version.

Gemini 3.7 Flash

Previous version for code and agents.

Gemini 3.6 Flash

First stable version of the series 3.6.

07 / SOURCES

Traceability