Which Model Is the Best Non-Thinking Fast Model? Gemini 2.5 Flash Lite vs Gemini 2.0 Flash

By Joe @ SimpleMetrics
Published 18 June, 2025
Updated 10 September, 2026
Which Model Is the Best Non-Thinking Fast Model? Gemini 2.5 Flash Lite vs Gemini 2.0 Flash

Historical comparison, corrected September 9, 2026: Google shut down gemini-2.0-flash in the Gemini API on June 1, 2026. Do not choose it for a new Gemini API integration. The June 2025 figures below are retained for reference, not as a current ranking or a newly verified benchmark.

For a later comparison, see our April 2026 guide to Gemini 3.1 Pro, Gemini 3.1 Flash-Lite, and Gemini 3 Flash, with a September model-ID correction. That guide is also a dated comparison, not a complete ranking of the latest models.

Gemini 2.5 Flash-Lite supports thinking; it is not a model without reasoning capability. Google's June 17, 2025 announcement says thinking is off by default for Flash-Lite and can be controlled with a thinking budget. That default setting is different from whether a model supports thinking at all. Neither the label nor this old table establishes which model is fastest or best for your task.

1. Performance Benchmarks (Non-Thinking Model)

How to read this historical table: these are numbers published in the original article, not our own controlled test results. We have not independently revalidated every score, model version, thinking setting or evaluation method. The “non-thinking” designation in the original table is not a verified configuration for all these figures.

The math row compares AIME 2025 with HiddenMath, which are different benchmarks. The code-editing and SWE rows contain estimates, not verified measured competitor results. Those rows cannot support winner claims. We have withdrawn the overall win tally and all row-level rankings; the remaining figures also require source and methodology verification before comparison.

Capability Benchmark Gemini 2.5 Flash-Lite (historical figures) Gemini 2.0 Flash Evidence limitation
General Reasoning MMLU-Pro 71.6% 77.6% Configuration and methodology not revalidated
Scientific QA GPQA Diamond 64.6% 60.1% Configuration and methodology not revalidated
Math AIME 2025 49.8% 63.5% (HiddenMath) Not comparable: AIME 2025 vs HiddenMath
Code (Python) LiveCodeBench 33.7% 34.5% Configuration and methodology not revalidated
Code Editing Aider Polyglot 26.7% ~25% (est.) Estimate, not a verified measured comparison
SWE-bench (Agentic Coding) SWE Verified 42.6% ~34.5% (est.) Estimate, not a verified measured comparison
Factual QA (Simple) SimpleQA 10.7% 29.9% Configuration and methodology not revalidated
Factual QA (Grounded) FACTS Grounding 84.1% 84.6% Configuration and methodology not revalidated
Multilingual QA Global MMLU (Lite) 81.1% 83.4% Configuration and methodology not revalidated
Image Reasoning MMMU 72.9% 71.7% Configuration and methodology not revalidated
Long-Context Memory MRCR (1M) 4.1% 70.5% Configuration and methodology not revalidated

The original article linked the Google announcements listed in Sources. Those links alone do not validate each transcribed score or make different benchmarks comparable. No overall winner can be established from this table.


2. Pricing

This is the original June 2025 price table, which cited Vertex AI pricing. It has not been revalidated as a current quote, and it does not imply that the retired model remains available. Provider, model version, input type and billing conditions matter; verify the current rate for an available model before budgeting.

Model Input (1M tokens) Output (1M tokens)
Gemini 2.0 Flash $0.15 $0.60
Gemini 2.5 Flash-Lite $0.10 $0.40

Original pricing reference: https://cloud.google.com/vertex-ai/generative-ai/pricing

The listed input and output rates are each about 33% lower for Flash-Lite. That is arithmetic on historical rates, not a measured workload saving or a reason to infer worse long-context capability from unverified benchmark figures.

3. Conclusion

  • 2.5 Flash-Lite supports a configurable thinking budget, with thinking off by default in the cited launch announcement. Evaluate an available model on representative inputs rather than relying on an unverified historical winner table.
  • 2.0 Flash is retired in the Gemini API. We withdraw the original recommendation to choose it as the most balanced non-thinking model. No new model performance test was run for this correction.

4. Sources

The thinking-default explanation was checked against Google's June 17, 2025 announcement, and Gemini API shutdown status against its current schedule, on September 9, 2026. Other links below are the original references, not a fresh verification of the table.

Found this useful? Share it!

If this helped you, I'd appreciate you sharing it with colleagues.

Was this page helpful?

Your feedback helps improve this content.

Related Posts