Google Introduces Gemini 3.6 Flash and Gemini 3.5 Flash-Lite

Google's title card for Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

My bench-supply folder is a buck-converter datasheet, a phone photo of the silkscreen, and a parts list pasted from three carts. On July 21, 2026, Google introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a narrower security model, Gemini 3.5 Flash Cyber. Tulsee Doshi, senior director of product management, posted it for the Gemini team.

Gemini 3.6 Flash is the new workhorse. Google says it is better at coding, knowledge work, and multimodal tasks than Gemini 3.5 Flash, at a lower price: $1.50 per million input tokens and $7.50 per million output tokens. A token is a chunk of text the model counts. Multimodal means the model can take more than text, such as a PDF and a photograph, in one job. Gemini 3.5 Flash-Lite is the fast, cheap 3.5-class model, at $0.30 per million input tokens and $2.50 per million output tokens. Google cites Artificial Analysis for 350 output tokens per second. Tokens per second is how many chunks the service writes each second.

Both general models are available the day of the post. Gemini 3.5 Pro is not. Google says it is in partner testing until it is ready. The two prices matter because a datasheet-plus-photo job and a long parts list are different workloads.

What actually changed?

Google's frame is the cost of an agent. An agent is software that chases a goal across many steps, often by calling tools, and then stops when the goal is done. Doshi writes that production agents need fewer tokens, lower latency, and steadier behavior. Latency is the wait before work comes back.

For 3.6 Flash, Google cites the Artificial Analysis Index for 17 percent fewer output tokens than 3.5 Flash. On some benchmarks, DeepSWE from Datacurve among them, Google says it sees cuts up to 65 percent. DeepSWE is a coding benchmark. Google also says 3.6 Flash uses fewer reasoning steps and fewer tool calls. The 17 percent figure is an index comparison. The 65 percent figure is an upper end on selected tests.

3.5 Flash-Lite is the other general release. Google calls it the fastest 3.5-series model and significantly above Gemini 3.1 Flash-Lite. Gemini 3.5 Flash Cyber is a specialized model paired with CodeMender, Google's code-security agent, and it is not a general shop tool.

How does the new piece work?

Google says 3.6 Flash came from 3.5 Flash feedback: better coding and knowledge work, a shorter answer, and a lower price. The old price is not printed.

Against 3.5 Flash, Google prints these comparisons. DeepSWE: 49 percent versus 37 percent, tied by Google to higher precision, fewer unwanted code edits, and fewer execution loops. MLE Bench, a machine-learning research benchmark: 63.9 percent versus 49.7 percent. OSWorld-Verified, a computer-use benchmark: 83.0 percent versus 78.4 percent. Computer use means the model operates a software screen, clicking and typing, as a tool. Google says that tool is now built into the client side of the Gemini API and Gemini Enterprise. GDPval-AA v2, a knowledge-work score: 1,421 versus 1,349. Google says Hebbia and Harvey found it capable at document parsing, charts, and drafts.

Flash-Lite is the high-volume model, including for search and document processing. Thinking level controls how much work it does before answering. Minimal or low thinking is the cheap, fast setting. A higher level is for multi-step subagent work. A subagent is a helper the main agent calls for one slice of the job. Google says Flash-Lite also has built-in computer use.

Against Gemini 3.1 Flash-Lite, Google reports Terminal-Bench 2.1 at 54 percent versus 31 percent, long-context GDM-MRCR v2 at 72.2 percent versus 60.1 percent, and GDPval-AA v2 at 1,140 versus 642. Against Gemini 3 Flash, Google says Flash-Lite leads on several agent and coding evals, including SWE-Bench Pro at 54.2 percent versus 49.6 percent and OSWorld-Verified at 74.0 percent versus 65.1 percent. SWE-Bench Pro is a software-repair benchmark. Those last two baselines are Gemini 3 Flash, not 3.5 Flash and not 3.6 Flash.

Google says 3.6 Flash has stronger safeguards against chemical, biological, radiological, and nuclear misuse, and against cyber-offense misuse, with more resistance to jailbreaks and fewer refusals on beneficial work. Details are on the model card. This article does not turn them into steps.

Google's Gemini 3.6 Flash benchmark chart
Google's bar chart, "Outperforms prior generations across agentic benchmarks." The panels are DeepSWE v1.1, MLE-Bench, GDPval-AA v2, and OSWorld-Verified.

What does this look like on a real project?

I would split the power-supply folder across the two models. I did not run that split on July 21. The photo and the datasheet are the careful pass. The vendor export is the bulk pass.

In Google AI Studio I would choose Gemini 3.6 Flash, attach the PDF and the bench photo, and ask for the datasheet's enable threshold plus whether the resistor code in the photo matches that divider. I would ask which section it used. A fluent paragraph with no pointer back to the PDF is how this job goes wrong.

I would send the parts export to Gemini 3.5 Flash-Lite for manufacturer part numbers, quantities, and the lines it could not parse. I would still spot-check the expensive parts. Firmware diffs stay on 3.6 Flash. A vulnerability hunt is the Cyber model's job, and this post does not hand a shop a key.

Google's chart of average output tokens per task
Google's bar chart, "Avg output tokens per task." DeepSWE v1.1 shows 276K tokens for Gemini 3.5 Flash and 97K for Gemini 3.6 Flash. The Intelligence Index panel shows 28K and 23K.

How does it compare with the previous version?

3.6 Flash takes the 3.5 Flash role: higher quality and shorter output, at a lower posted price, on Google's account. The DeepSWE move is large. The OSWorld-Verified move is smaller. The 17 percent output-token cut is the Artificial Analysis comparison Google cites.

Flash-Lite's previous small model in these charts is Gemini 3.1 Flash-Lite, with a side comparison to Gemini 3 Flash. The jumps versus 3.1 Flash-Lite are large. The SWE-Bench Pro gap versus Gemini 3 Flash is smaller. Google still calls 3.6 the workhorse. Flash-Lite does not replace it.

Gemini 3.5 Pro stays in partner testing. It is not broadly available, and "soon" is not a ship date. Google also says it has started pretraining for Gemini 4, its most ambitious pretraining run so far. Pretraining is the early pass over a huge data pile, before a model is offered. Gemini 4 is not something you can call.

Where does it sit next to other tools a maker already uses?

Developers get both general models in the Gemini API through Google AI Studio and Android Studio. Google also lists 3.6 Flash in Google Antigravity, which this post does not define. Enterprises get the models in the Gemini Enterprise Agent Platform, and 3.6 Flash is also in the Gemini Enterprise app. Everyone gets them in the Gemini app, and Flash-Lite is rolling out in Google Search. I would not treat the Search box as 3.6 Flash.

3.5 Flash Cyber sits beside CodeMender. Google says several agents work toward one combined report and that the setup is competitive on CyberGym. The model is built on 3.5 Flash, fine-tuned for finding and fixing vulnerabilities, at a lower price per token than larger models. No token price is printed. Access is for governments and trusted partners, through a limited CodeMender pilot Google says is coming soon. This article does not describe the finding or the fixing.

What does it cost, and who can use it today?

3.6 Flash is $1.50 per million input tokens and $7.50 per million output tokens. 3.5 Flash-Lite is $0.30 and $2.50 on the same scale. Google says the 3.6 rate and its output-token use are both lower than 3.5 Flash.

The two general models are on the list above. 3.5 Pro stays partner-only. Gemini 4 is pretraining news. Cyber is not on the public start list. Thinking level has no separate price in the post. A high setting still spends more tokens.

What is still unproven?

I did not run these models or reproduce the benchmarks. The 17 percent token cut and the 350 tokens per second are Google citing Artificial Analysis. "Up to 65 percent" is a ceiling on some tests, not a typical job. The quality bars are Google's. I am not quoting the testimonial images.

Disclosure: The author is a paying subscriber to ChatGPT Plus, Claude Pro, and SuperGrok and uses all three services on a daily basis. The Makers Workbench is not affiliated with OpenAI, Anthropic, xAI, Google, or any of the other major AI companies covered in our reporting. No company receives favorable editorial treatment based on the author's personal subscriptions.

Sources and image credits

Sub-Category

Add new comment

Restricted HTML

  • You can align images (data-align="center"), but also videos, blockquotes, and so on.
  • You can caption images (data-caption="Text"), but also videos, blockquotes, and so on.