Skip to content
Google Blog·

Google Launches Gemini 3.1 Ultra With Native 2-Million-Token Multimodal Context

Google Launches Gemini 3.1 Ultra With Native 2-Million-Toke…

TL;DR

Google launched Gemini 3.1 Ultra featuring a 2-million-token context window that operates natively across text, image, audio, and video without transcription i…

Google launched Gemini 3.1 Ultra featuring a 2-million-token context window that operates natively across text, image, audio, and video without transcription intermediaries. A new sandboxed Code Execution tool lets the model write, run, and test code mid-conversation.

Google Launches Gemini 3.1 Ultra With Native 2-Million-Token Multimodal Context — Google Blog

Key Points

1

2-million-token native multimodal context across text, image, audio, video

2

Trained from scratch to reason across modalities simultaneously, no transcription step

3

New sandboxed Code Execution tool runs and tests code in-conversation

4

Available via Gemini app and Vertex AI for enterprise customers

Why It Matters

True multimodal reasoning at 2M tokens collapses entire pipelines into a single model call, redefining what 'long-context' means for enterprise workflows.

googlegeminimultimodallong-context

Frequently Asked Questions

Why does this matter?

True multimodal reasoning at 2M tokens collapses entire pipelines into a single model call, redefining what 'long-context' means for enterprise workflows.

What happened?

Google launched Gemini 3.1 Ultra featuring a 2-million-token context window that operates natively across text, image, audio, and video without transcription i…

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,097 builders reading daily.

Also get