Historical comparison, corrected September 9, 2026: Google shut down gemini-2.0-flash in the Gemini API on June 1, 2026. Do not choose it for a new Gemini API integration. The June 2025 figures below are retained for reference, not as a current ranking or a newly verified benchmark.
For a later comparison, see our April 2026 guide to Gemini 3.1 Pro, Gemini 3.1 Flash-Lite, and Gemini 3 Flash, with a September model-ID correction. That guide is also a dated comparison, not a complete ranking of the latest models.
Gemini 2.5 Flash-Lite supports thinking; it is not a model without reasoning capability. Google's June 17, 2025 announcement says thinking is off by default for Flash-Lite and can be controlled with a thinking budget. That default setting is different from whether a model supports thinking at all. Neither the label nor this old table establishes which model is fastest or best for your task.
1. Performance Benchmarks (Non-Thinking Model)
How to read this historical table: these are numbers published in the original article, not our own controlled test results. We have not independently revalidated every score, model version, thinking setting or evaluation method. The “non-thinking” designation in the original table is not a verified configuration for all these figures.
The math row compares AIME 2025 with HiddenMath, which are different benchmarks. The code-editing and SWE rows contain estimates, not verified measured competitor results. Those rows cannot support winner claims. We have withdrawn the overall win tally and all row-level rankings; the remaining figures also require source and methodology verification before comparison.
| Capability | Benchmark | Gemini 2.5 Flash-Lite (historical figures) | Gemini 2.0 Flash | Evidence limitation |
|---|---|---|---|---|
| General Reasoning | MMLU-Pro | 71.6% | 77.6% | Configuration and methodology not revalidated |
| Scientific QA | GPQA Diamond | 64.6% | 60.1% | Configuration and methodology not revalidated |
| Math | AIME 2025 | 49.8% | 63.5% (HiddenMath) | Not comparable: AIME 2025 vs HiddenMath |
| Code (Python) | LiveCodeBench | 33.7% | 34.5% | Configuration and methodology not revalidated |
| Code Editing | Aider Polyglot | 26.7% | ~25% (est.) | Estimate, not a verified measured comparison |
| SWE-bench (Agentic Coding) | SWE Verified | 42.6% | ~34.5% (est.) | Estimate, not a verified measured comparison |
| Factual QA (Simple) | SimpleQA | 10.7% | 29.9% | Configuration and methodology not revalidated |
| Factual QA (Grounded) | FACTS Grounding | 84.1% | 84.6% | Configuration and methodology not revalidated |
| Multilingual QA | Global MMLU (Lite) | 81.1% | 83.4% | Configuration and methodology not revalidated |
| Image Reasoning | MMMU | 72.9% | 71.7% | Configuration and methodology not revalidated |
| Long-Context Memory | MRCR (1M) | 4.1% | 70.5% | Configuration and methodology not revalidated |
The original article linked the Google announcements listed in Sources. Those links alone do not validate each transcribed score or make different benchmarks comparable. No overall winner can be established from this table.
2. Pricing
This is the original June 2025 price table, which cited Vertex AI pricing. It has not been revalidated as a current quote, and it does not imply that the retired model remains available. Provider, model version, input type and billing conditions matter; verify the current rate for an available model before budgeting.
| Model | Input (1M tokens) | Output (1M tokens) |
|---|---|---|
| Gemini 2.0 Flash | $0.15 | $0.60 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 |
Original pricing reference: https://cloud.google.com/vertex-ai/generative-ai/pricing
The listed input and output rates are each about 33% lower for Flash-Lite. That is arithmetic on historical rates, not a measured workload saving or a reason to infer worse long-context capability from unverified benchmark figures.
3. Conclusion
- 2.5 Flash-Lite supports a configurable thinking budget, with thinking off by default in the cited launch announcement. Evaluate an available model on representative inputs rather than relying on an unverified historical winner table.
- 2.0 Flash is retired in the Gemini API. We withdraw the original recommendation to choose it as the most balanced non-thinking model. No new model performance test was run for this correction.
4. Sources
The thinking-default explanation was checked against Google's June 17, 2025 announcement, and Gemini API shutdown status against its current schedule, on September 9, 2026. Other links below are the original references, not a fresh verification of the table.
- Gemini API shutdown schedule: https://ai.google.dev/gemini-api/docs/deprecations
- https://developers.googleblog.com/en/gemini-2-5-thinking-model-updates
- https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/#gemini-2-0
- https://blog.google/products/gemini/gemini-2-5-model-family-expands
- https://cloud.google.com/vertex-ai/generative-ai/pricing