Ranked #4 Everyday Ecosystem — The Leading AI Assistants
Google DeepMind

Gemini 3.1 Pro

Gemini 3.1 Pro is Google's deep-thinking specialist: an aging flagship that still excels at science, unfamiliar reasoning problems, long documents, and mixed-media research. The newer Flash family is faster and cheaper, but 3.1 Pro remains useful when the quality of the analysis matters more than finishing first.

Updated August 14, 2026 Multi-modalReasoningLong context
9.7out of 10
Official Website
Best for

Gemini 3.1 Pro is Google's deep-thinking specialist: an aging flagship that still excels at science, unfamiliar reasoning problems, long documents, and mixed-media research. The newer Flash family is faster and cheaper, but 3.1 Pro remains useful when the quality of the analysis matters more than finishing first.

Why It Wins

Its launch results included 77.1% on ARC-AGI-2, 94.3% on GPQA Diamond, 44.4% on Humanity's Last Exam without tools, and 80.6% on SWE-Bench Verified. It accepts text, images, audio, video, and PDFs within a one-million-token context window.

Watch out

Six months after release it is still a Preview model at $2 input / $12 output per million tokens for prompts up to 200k. Its promised 3.5 Pro successor missed Google's June target and may have been overtaken by the Gemini 4 roadmap, while 3.7 Flash now handles much of the workhorse role.

01

What It Actually Is

Imagine giving the same difficult research folder to three colleagues. The first reads every page, compares the footnotes, studies the graphs, and returns a careful argument. The second moves quickly, calls whatever tools are needed, and turns the folder into a working dashboard before lunch. The third handles a thousand similar folders at once, extracting the same twelve fields from each one without exhausting the budget.

Those three colleagues are a useful way to understand Google’s current Gemini family. Gemini 3.1 Pro is the first: the deliberate analyst. Gemini 3.7 Flash is the fast builder and agent. Gemini 3.5 Flash-Lite is the inexpensive production line. The word Pro does not automatically make the first model best for every job; it tells you which kind of compromise Google originally designed it to make.

The reasoning flagship that aged in public

Google released Gemini 3.1 Pro in Preview on February 19, 2026. Its headline improvement was not a new chat interface or another bundle of Google services. It was a large increase in core reasoning. On ARC-AGI-2—puzzles that require discovering unfamiliar visual rules rather than recalling facts—Google reported 77.1%, compared with 31.1% for Gemini 3 Pro. It also recorded 94.3% on GPQA Diamond and 44.4% on Humanity’s Last Exam without tools.

Artificial Analysis reached a similar broad conclusion from its own evaluation suite. At launch, 3.1 Pro scored 57 on its Intelligence Index, took first place, and led six of the ten component evaluations. Artificial Analysis measured roughly 114 output tokens per second: not Flash-fast, but unusually quick for a top reasoning model at that moment. It also praised the economics relative to contemporary flagships, because Google kept the familiar price of $2 per million input tokens and $12 per million output tokens for requests at or below 200,000 tokens.

That February context matters. A benchmark score is a photograph stamped with a date, not a permanent championship belt. Six months of model releases have shifted the field. Gemini 3.1 Pro has not become less intelligent, but competitors and Google’s own Flash models have become better at long-horizon coding, computer use, and enterprise workflows.

Where 3.1 Pro still earns its place

The model is at its best when a problem is difficult before it becomes repetitive. It can receive text, images, audio, video, and PDF files in one prompt, with up to one million input tokens and 64,000 output tokens. A scientist can place a paper beside microscopy images. An investigator can compare a recorded interview with a timeline and financial report. A product team can inspect a screen recording alongside the relevant source code.

This is analysis, not native media generation. Gemini 3.1 Pro produces text. Veo creates video, Nano Banana creates images, and Lyria creates music. Google may connect those systems inside a product, but describing all of them as abilities of the 3.1 Pro model is like crediting a conductor with personally playing every instrument in the orchestra.

The million-token window also deserves a realistic explanation. It is an enormous worktable, not a photographic memory. Google’s own MRCR test gives 3.1 Pro an 84.9% average at 128k but only 26.3% on the pointwise 1M setting. If the answer depends on one sentence buried near page 4,000, simply uploading everything and hoping is poor information architecture. Narrow the corpus, label the sources, ask for page-level evidence, and verify important claims.

What happened to Gemini 3.5 Pro?

The succession story has become more revealing than a normal missed launch date. On May 19, Google announced Gemini 3.5 Flash and said 3.5 Pro was already being used internally and would roll out “next month.” June ended without a release. On July 16, Bloomberg reported that the model was months behind schedule because Google was trying to improve its capabilities, particularly coding, and that a late-June training-data refresh had not met expectations.

Five days later, Google launched Gemini 3.6 Flash and 3.5 Flash-Lite. Its language about Pro had changed: 3.5 Pro was “testing with partners” and would become broadly available “as soon as it’s ready.” In the same announcement, Google said it had already begun its most ambitious pretraining run for Gemini 4. By August 14, the public Pro page still showed Gemini 3.1 Pro under a small promise that 3.5 Pro was “coming soon.”

There is now a credible industry report from SemiAnalysis that 3.5 Pro has been quietly canceled. The surrounding facts make the report plausible: the June window was missed, the model was reportedly struggling to meet coding goals, two newer Flash generations shipped, Gemini 4 training began, and there is still no public 3.5 Pro endpoint. But plausible is not proven. Google’s latest public statements have not announced a cancellation.

The most accurate conclusion is therefore conditional. The original Gemini 3.5 Pro plan has plainly slipped. The model carrying that exact name may be canceled, renamed, folded into Gemini 4, or held until a substantial rework is competitive. Users should not buy a subscription or design a system around an unreleased promise. If Google ships it, evaluate the actual endpoint; until then, “coming soon” is marketing, not availability.

The ecosystem did not wait

While Pro remained in the workshop, the Flash line took over much of Google’s practical roadmap. Gemini 3.7 Flash, released August 13, is now generally available and is Google’s recommended migration target even for 3.1 Pro workloads involving code generation, multimodal reasoning, multi-step agents, and design adherence. It supports a million-token context, 64k output, built-in tools, and adjustable low, medium, or high thinking effort.

The performance table explains why Google calls it a workhorse. It records 43.6% on FrontierCode 1.1, 65.3% on DeepSWE v1.1, 85.8% on Terminal-Bench 2.1, 30.4% on AutomationBench, and 34.0% on GDP.pdf. Its Artificial Analysis Intelligence Index is 56—one point below 3.1 Pro’s launch score, though cross-date scores should be compared cautiously—and independent testing cited by Google places output around 340 tokens per second. Through December 31, 2026, it costs $0.75 per million input tokens and $3.75 per million output tokens; the standard $1.50/$7.50 rate begins in January.

Gemini 3.5 Flash-Lite serves a different layer of the stack. It is stable, accepts the same broad set of input media, offers a 1,048,576-token input limit, and costs $0.30/$2.50. Google cites roughly 350 output tokens per second. Its natural jobs are document parsing, classification, receipt extraction, search subagents, and generating many candidates in parallel. Google reports 54% on Terminal-Bench 2.1 versus 31% for 3.1 Flash-Lite, 72.2% versus 60.1% on long-context MRCR, and 1140 versus 642 on GDPval-AA v2. That does not turn Lite into Pro; it makes Lite a remarkably capable conveyor belt.

Which Gemini should you actually use?

Use Gemini 3.1 Pro when your own evaluation shows that its deeper reasoning produces better scientific analysis, research synthesis, or handling of unfamiliar problems. Keep it if a carefully validated workflow depends on its particular behavior. It remains a serious model, not an antique.

Use Gemini 3.7 Flash as the default for new coding agents, web and UI generation, tool-heavy automation, and high-volume knowledge work. Its combination of current agent performance, speed, stable availability, and discounted price is difficult for 3.1 Pro to answer.

Use Gemini 3.5 Flash-Lite behind the scenes, where one orchestrating model delegates many small, well-specified tasks. A cheap model extracting invoice dates ten thousand times can create more business value than a flagship composing one magnificent paragraph.

And wait for evidence before treating Gemini 3.5 Pro as anything at all. The lesson of its delay is not that Google has stopped building powerful models. It is that model names are promises, while endpoints are products.

02

Strengths and honest limitations

Key Strengths

  • Excellent abstract and scientific reasoning: Gemini 3.1 Pro’s 77.1% ARC-AGI-2, 94.3% GPQA Diamond, and 44.4% Humanity’s Last Exam results made it one of February 2026’s strongest reasoning models. Benchmarks are maps rather than the territory, but these three point in the same direction: it is unusually comfortable with new logic, technical concepts, and research questions.
  • Multimodal evidence goes into one analysis: It can reason over text, images, audio, video, and PDFs together. A researcher can connect an interview’s tone with a chart in a report and frames from a product demonstration without manually converting everything into prose first. Its output is text; Veo, Nano Banana, and Lyria are separate media-generation models.
  • A large worktable for documents and code: The one-million-token input window can hold lengthy case files, research collections, or substantial repositories. Capacity is not the same as perfect recall—Google reports only 26.3% on its hardest pointwise 1M-token MRCR test—so focused evidence and requested citations still produce safer answers.
  • Strong tool use and coding for its generation: At launch it reached 80.6% on SWE-Bench Verified, 68.5% on Terminal-Bench 2.0, 85.9% on BrowseComp, and 69.2% on MCP Atlas. It supports function calling, structured output, Google Search, and code execution, making it more than a document-only thinker.
  • Available across Google’s serious-work surfaces: Google lists it in the Gemini app, AI Studio, the Gemini API, enterprise products, AI Mode, Antigravity, and NotebookLM. For an organization already using Google Workspace and Cloud, governance and access can be as important as a narrow leaderboard advantage.

Honest Limitations

  • The promised successor is stuck in limbo: Google said in May that 3.5 Pro would roll out in June. That date passed; Bloomberg reported the model was months behind and that attempts to improve coding had fallen short. Google then changed its language to partner testing and release ‘as soon as it’s ready.’
  • Cancellation is plausible, not confirmed: SemiAnalysis has reported that 3.5 Pro was quietly canceled, and Google’s focus on 3.6/3.7 Flash plus an already-started Gemini 4 pretraining run makes that interpretation credible. But Google’s live page still says ‘3.5 Pro coming soon.’ Treat canceled, renamed, and heavily reworked as possibilities—not facts.
  • The Flash family is a better production fit for many jobs: Gemini 3.7 Flash is generally available, much faster, stronger on many current coding and agent benchmarks, and temporarily costs $0.75/$3.75 per million input/output tokens. Gemini 3.5 Flash-Lite costs $0.30/$2.50 and targets high-volume subagents and document extraction.
  • Preview status and long-context falloff require caution: Preview endpoints have less lifecycle certainty than stable ones, and a one-million-token allowance does not guarantee that the model will retrieve every buried fact. Production systems need pinned model IDs, fallback tests, source citations, and independent checks for consequential answers.
03

Benchmark Snapshot

ARC-AGI-2 — 77.1%

Google's verified launch score on novel visual-logic puzzles, up from 31.1% for Gemini 3 Pro and ahead of every contemporary model in Google's comparison. It remains the clearest evidence for 3.1 Pro's leap in abstract reasoning.

GPQA Diamond — 94.3%

A no-tools test of difficult graduate-level science questions. The result supports using 3.1 Pro for scientific explanation and hypothesis work, but it does not make unsupported factual claims automatically trustworthy.

SWE-Bench Verified — 80.6%

A strong result on real software-repository issues, almost tied with Opus 4.6 in Google's launch table. Newer long-horizon and agent benchmarks are less flattering, showing that the model remains capable without remaining the coding frontier.

Artificial Analysis Intelligence Index — 57 at launch

Artificial Analysis ranked it first in February 2026 and measured about 114 output tokens per second. This is an important historical snapshot, not a claim that it still leads a fast-changing August leaderboard.

04

The Verdict

Gemini 3.1 Pro is still a good specialist for difficult reasoning, scientific work, and careful synthesis across many kinds of evidence. It is no longer the sensible default for every heavy task: 3.7 Flash is the stronger general production choice, while 3.5 Flash-Lite is the economical engine for extraction and parallel subagents. As for 3.5 Pro, assume nothing until an endpoint ships—its repeated delay, the reported cancellation, and Google’s Gemini 4 work all suggest that the product bearing that exact name may never arrive.

05

Frequently Asked Questions