Skip to content

Google Cuts Agent Operating Costs with Two New Gemini Flash Models

Google has released two lower-priced Gemini Flash models for agents and high-volume tasks, creating a practical test of task costs in Arabic.

Share
Programming code on a computer screen

Listen to this article

Read by Anchor

What happened: In a late signal from the 72-hour window, Google made Gemini 3.6 Flash and Gemini 3.5 Flash-Lite available to developers, enterprises and users on July 21. The company says 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis benchmark, and prices it at $1.50 per million input tokens and $7.50 per million output tokens. Flash-Lite costs $0.30 for input and $2.50 for output per million tokens. Google says it generates 350 output tokens per second on the same benchmark.

The lens: economic transformation Competition is shifting from model rankings on benchmarks to the cost of a completed task and the speed of running it at scale. This puts pressure on model providers and gives organisations more scope to divide work between a primary model and lighter agents.

Who is affected: Product and software teams, contact centres, ecommerce businesses and organisations running automated research and high-volume document processing.

What it means for the region: Arab companies can test high-demand services in Arabic at a lower cost, but supplier figures are not enough to demonstrate local quality. Measurement must include Arabic accuracy, the number of tool calls, task duration and total task cost.

The practical takeaway: Run a sample of 100 real tasks on your current model and Flash-Lite, then compare the cost of a correctly completed task, not just the token price, before changing the production workflow.

Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

Don't miss the next story

Subscribe for updates