
Google DeepMind unveiled Gemini 3.8 Flash, an artificial-intelligence model that seeks to blend strong performance with lower latency and reduced operating costs.
The system can handle text, images, audio, video and PDF documents, and it is capable of processing up to 1 million tokens in a single context while generating as many as 64,000 tokens in response.
Agentic capabilities and tool integration
It supports function calls, built-in search tools and direct interaction with a computer environment, features that enable autonomous agents and multi-step workflows.
Developers may choose among effort levels, low, medium (default) or high, to balance speed, cost and answer quality, a flexibility highlighted in benchmark tests.
Compared with the earlier Gemini 3.7 Flash, the new version follows an industry pattern of offering more context for a similar price point. Other providers have also introduced tiered models that trade off latency against reasoning depth, making such options increasingly common.
Safety controls and availability
Safety assessments follow DeepMind’s Frontier Safety Framework, adding stronger defenses against prompt injection and jailbreak attempts, especially for cyber and biosecurity uses.
A special Gemini 3.8 Flash Cyber variant offers looser safeguards under a trust-based program for institutions that need higher capability.
Read Also: Melbourne AI expo highlights robots, drones and everyday applications
The model is reachable through the Gemini app, Google Cloud, AI Studio, the Gemini API, Search’s AI mode and the Antigravity platform, and it is now generally available.
Introductory pricing runs at $0.75 per million input tokens and $3.75 per million output tokens until the end of 2026; rates double after that date.
Cached inputs receive a discount of about $0.075 per million tokens, and actual costs may vary by region and usage volume.
Target users include developers building autonomous pipelines, enterprises handling legal or financial analysis, product teams needing fast prototyping, and organizations weighing cost against output quality.
The system can still produce inaccurate statements, experience occasional timeouts, and shows a modest drop in safety performance for non-English inputs, so oversight remains essential.
Because of these limitations, teams are encouraged to implement human-in-the-loop checks, especially when the output informs high-stakes decisions such as loan approvals or medical triage.