You can see the latest product updates for all of Google Cloud on the Google Cloud page, browse and filter all release notes in the Google Cloud console, or programmatically access release notes in BigQuery.
Anthropic's Claude Opus 5
Claude Opus 5 is available in Model Garden.
Gemini 3.6 Flash and 3.5 Flash-Lite are generally available (GA)
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now generally available (GA) and available for production use. These models are designed to improve upon their predecessors' capabilities, including improved token usage and improved document understanding. See the linked model information pages for more information.
This release includes some potentially breaking changes from previous Flash and Flash-Lite models:
Model are no longer supported and will return an error:
"type": "model_output" will fail."role": "model" will fail.Open model endpoint deprecations
The following open model endpoints are deprecated and will be retired on October 21, 2026. For more information, see Open model deprecations.
deepseek-ocr-maasdeepseek-r1-0528-maasdeepseek-v3.2-maasdeepseek-v3.1-maasglm-5-maasglm-4.7-maasgpt-oss-20b-maaskimi-k2-thinking-maasllama-3.3-70b-instruct-maasminimax-m2-maasmultilingual-e5-large-instruct-maasmultilingual-e5-small-maasqwen3-235b-a22b-instruct-2507-maasqwen3-coder-480b-a35b-instruct-maasqwen3-next-80b-a3b-instruct-maasqwen3-next-80b-a3b-thinking-maasSecurity update for Server-Side Request Forgery (SSRF) in Agent Studio
This release fixes a Server-Side Request Forgery (SSRF) vulnerability in the
auto-generated /api-proxy backend endpoint for web applications created
before July 1, 2026, using Agent Studio.
If you downloaded, generated, or deployed web application code from Agent Studio before July 1, 2026, regenerate the app from Agent Studio and deploy the new version. For more information, see Quickstart: Deploy your Agent Studio prompt as a web application
The updated backend code includes strict domain allowlist validation, ensuring
that destination hostnames for the /api-proxy endpoint end with allowed Google
Cloud domains, such as *-aiplatform.clients6.google.com
Gemini 3.1 Flash Image Preview and 3 Pro Image Preview are retired
The Nano Banana preview models gemini-3.1-flash-image-preview and
gemini-3-pro-image-preview have been retired and are no longer accessible.
Update your code to use either gemini-3.1-flash-image or
gemini-3-flash-image instead.
Memory Bank memory profiles are generally available (GA)
Memory profiles in Memory Bank are generally available (GA). Memory profiles allow you to generate structured profiles, which are data structures with static schemas populated and updated using LLMs. By defining a fixed schema, you ensure your agents have immediate, low-latency access to evolving information without the need for expensive search operations during a session.
For details, see Memory profiles.
Retirement for preview models for 2.5 Flash, 2.5 Flash-Lite, and 3.1 Flash-Lite
The following preview model endpoints have been retired and are no longer accessible:
gemini-2.5-flash-lite-preview-09-2025gemini-2.5-flash-preview-05-2025gemini-3.1-flash-lite-previewSee Migrate to the latest Google models for information on how to migrate your project.
Grok 4.1 models deprecation
The Grok 4.1 model family (including xai/grok-4.1-fast-reasoning and xai/grok-4.1-fast-non-reasoning) on the Gemini Enterprise Agent Platform is deprecated and will be shut down on August 20, 2026. After this date, Google Agent Platform Model as a Service (MaaS) will no longer serve these models. Specifically, API requests directed to the https://aiplatform.googleapis.com/v1/projects/<your-project>/locations/global/endpoints/openapi/chat/completions endpoint using the Grok 4.1 model IDs will fail and return a 400 error. To maintain service, migrate your applications to newer xAI models (such as Grok 4.2 or Grok 4.3) or choose an alternative model from the Google Cloud Model Garden.
Memory Bank support for Gemini Embedding 2
Memory Bank supports the gemini-embedding-2 model for similarity search configurations.
When you configure gemini-embedding-2 as the embedding model, you must use one of the global, us, or eu endpoints in the model's resource name (for example, projects/{project}/locations/us/publishers/google/models/gemini-embedding-2). Memory Bank does not support using regional locations (for example, us-central1) for gemini-embedding-2.
For details, see Similarity search configuration.
Memory Bank IngestEvents is generally available (GA)
The Memory Bank IngestEvents API is generally available. The IngestEvents API decouples event ingestion from memory generation, letting you continuously stream content to Memory Bank and configure when memory generation is triggered.
This GA release includes the following features:
overlap_event_count parameter to re-include already-processed events at the start of the next window to keep memories coherent.revision_labels, revision_ttl, or disable_memory_revisions to customize how generated revisions are managed.metadata and metadata_merge_strategy configuration parameters to store structured information alongside generated memories.For details, see Ingest events.
AlphaGenome released for Gemini Enterprise Agent Platform
AlphaGenome, Google DeepMind's state-of-the-art genomics foundation model, is now available for deployment and use with Gemini Enterprise Agent Platform.
Designed to decipher the functional regulatory code of the human genome, AlphaGenome analyzes large-scale DNA sequences at single-base resolution to predict how genetic variations affect molecular and biological mechanisms like gene expression, chromatin accessibility, and RNA splicing. For information on how to use AlphaGenome in Agent Platform, see the documentation.
Provisioned Throughput: Multiple pending new orders GA
The ability to submit multiple pending orders is generally available. you can submit up to seven Google model orders for the same model and region. See Multiple pending orders.
Anthropic's Claude Sonnet 5
Claude Sonnet 5 is available in Model Garden.
Gemini 3.5 Flash default model for Memory Bank
The default model used for Memory Bank generation is now Gemini 3.5 Flash instead of Gemini 2.5 Flash. For information, see Set up Memory Bank.
Semantic Governance Policies available in Public Preview
Semantic governance policies and the policy engine are now available in Preview. Semantic governance policy provides an intelligent security and compliance layer that evaluates an AI agent's proposed tool calls against user intent and organizational business rules at runtime.
Key capabilities include:
For more information, see Semantic governance policies overview.
Provisioned Throughput: Email notifications GA
The ability to get email notifications for Provisioned Throughput events is generally available. See Get email notifications.
New Provisioned Throughput features
The following Provisioned Throughput features are generally available:
Administrative correction to Gemini Online Inference API on Gemini Enterprise Agent Platform SLA
We made an administrative correction to the Gemini Online Inference API on Gemini Enterprise Agent Platform Service Level Agreement (SLA), addressing a clerical issue.
AI security findings in Agent Platform
Viewing AI security findings and posture management summaries in Gemini Enterprise Agent Platform is generally available.
With this release, the Security dashboard introduces the Top security findings widget.
Also, specific features within the AI security widgets are available in Preview, including the following:
For more information, see View security findings.
Model Armor for Agent Gateway in General Availability
You can enable Model Armor on an Agent Gateway resource to apply your organization's content security guardrails to prompts and responses that pass through the gateway. This feature is generally available (GA).
For more information, see Configure Model Armor on a gateway.
Supervised fine-tuning available for Gemini 3.1 Flash Lite and Gemini 3.5 Flash in Public Preview
Supervised fine-tuning is now available for the gemini-3.1-flash-lite and
gemini-3.5-flash models in
Preview. Model tuning
for Gemini 3.1 Flash Lite and Gemini 3.5 Flash is restricted to us-central1
and europe-west4 and tuned model serving is restricted to the us and eu
multi-region endpoints.
See About supervised fine-tuning for more information.
Provisioned Throughput support for supervised fine-tuned Gemini 3 model inference.
Provisioned Throughput can be used to assure supervised fine-tuned inference using the same quota. Supervised fine-tuned inference for Gemini 3 models incurs a higher burndown rate compared to base model inference. Learn more here.
Agent Gateway in General Availability
Agent Gateway is the networking component of the Gemini Enterprise Agent Platform ecosystem. It secures and governs connectivity for all agentic interactions, whether they occur between users and agents, agents and tools, or among agents themselves.
For details, see Agent Gateway overview.
Agent Observability is generally available (GA)
This release provides visibility into the performance, behavior, and health of deployed agents and Model Context Protocol (MCP) servers directly within the agent management workflow.
Key updates in this release include:
For more information, see the following:
Agent Registry is generally available (GA)
Agent Registry is generally available (GA). Agent Registry is a centralized catalog for discovering and registering agents and Model Context Protocol (MCP) servers.
The following features are available in Agent Registry for the GA launch stage:
v1 version of the Agent Registry API is available. Cloud client libraries are available in C#, Go, Java, Node.js, PHP, Python, and Ruby.1.0, letting you explicitly declare transport endpoints and bindings inside the supportedInterfaces array, in addition to the existing 0.3 schema support.Known limitations:
For more information, see the Agent Registry overview.
The Agent Identity API (agentidentity.googleapis.com) is
available in Preview.
This new API replaces the legacy IAM Connectors API
(iamconnectors.googleapis.com) for managing auth providers and agent
identities.
During the preview migration period, both APIs operate side-by-side. Existing
auth providers are automatically mirrored to the new V2 resource hierarchy
(authProviders/), allowing you to migrate your IAM policies, agent code, and
client applications without downtime.
Memory Bank and Sessions global and multi-regional endpoints GA
Memory Bank and Sessions support for multi-regional and global endpoints is now in General Availability (GA). For more information, see Supported locations for agents. Note that Customer-Managed Encryption Keys (CMEK) cannot be used if your Memory Bank or Sessions instance is configured to use the global endpoint.
Reinforcement Learning fine-tuning is available for Gemini 3.5 Flash in Public Preview
Reinforcement Learning fine-tuning is now available for the gemini-3.5-flash models in
Preview. Model tuning
for Gemini 3.5 Flash is restricted to us-central1 and europe-west4, and tuned
model serving is restricted to the us and eu multi-region endpoints.
See About reinforcement learning fine-tuning for more information.
Anthropic's Claude Fable 5
Claude Fable 5 is available in Model Garden.
Updates to abuse monitoring and zero data retention documentation
Documentation for abuse monitoring, zero data retention, and responsible AI has been updated to align with the Advanced AI Safety Addendum. These updates include new details regarding Advanced AI safety, partner-specific terms, and request-response logging for models like Claude Mythos and Opus.
For more information, see:
Gemini 2.0 Flash and Gemini 2.0 Flash-Lite are discontinued
Gemini 2.0 Flash and 2.0 Flash-Lite are discontinued and are no longer available. This includes both model serving and Provisioned Throughput. Use Gemini 3.1 Flash-Lite, Gemma 4, or more recent Gemini releases.
Agent Platform Gemini 3.1 Flash Image and Gemini 3 Pro Image
Gemini Enterprise Agent Platform Gemini 3.1 Flash Image and Gemini 3 Pro Image are Generally Available.
With this release, Gemini 3.1 Flash Image and Gemini 3 Pro Image support 4K image outputs in Preview.
Also supported in this release, Gemini 3.1 Flash Image supports video inputs in Preview. You can use video inputs to generate thumbnails or representative images of videos.
For more information, see the following:
Agent Platform Gemini 3.1 Flash Image Preview and Gemini 3 Pro Image Preview deprecation
Gemini Enterprise Agent Platform Gemini 3.1 Flash Image Preview and Gemini 3 Pro Image Preview are deprecated. We recommend that you update your model endpoints before July 17, 2026, to avoid service disruption.
The following are the discontinued endpoints and recommended endpoint migration:
| Discontinued endpoints | Recommended endpoint migration |
|---|---|
gemini-3.1-flash-image-preview |
gemini-3.1-flash-image |
gemini-3-pro-image-preview |
gemini-3-pro-image |
Anthropic's Claude Opus 4.8
Claude Opus 4.8 is available in Model Garden.
User ID logging now included with agent logs when you opt in to "Enable logging of prompt inputs and response outputs"
Prompt input and response output logging now includes the
user.id field. This addition allows better tracking of
anomalous tool interactions.
For details on configuration, see Write traces for an agent.
The Gemini Deep Research Agent released in Preview
The Gemini Deep Research Agent has been released in Preview. The Gemini Deep Research Agent is a managed AI agent that plans, executes, and synthesizes complex, multi-step research workflows across the public web and private enterprise data to generate comprehensive, cited reports.
For more information, see Use the Gemini Deep Research Agent.
Agent Platform Sandboxes
Additional Agent Platform sandbox features are now available:
Identify the agents with the most content security violations
The Security dashboard displays the top 10 agents with the most content violations detected by Model Armor. The list shows the agent ID of each agent and the number of violations detected for that agent. For more information, see Monitor content security.
Supervised fine-tuning available for Gemini 3.1 Flash Lite (Preview)
Supervised fine-tuning is now available for limited support for the
gemini-3.1-flash-lite model. During this period, model tuning for Gemini
3.1 Flash Lite is restricted to us-central1 and europe-west4 and tuned model
serving is restricted to the us and eu multi-region endpoints.
See About supervised fine-tuning for more information.
Set media resolution at a Part-level for data when using supervised fine-tuning
Supervised fine-tuning now supports Part-level mediaResolution declarations
for images, videos, and PDFs. Part-level media resolution declarations also
support the MEDIA_RESOLUTION_ULTRA_HIGH level.
See the following media type–specific pages for more information:
Gemini 3.5 Flash is generally available (GA)
For details, see the model specifications page.
Manage agent revisions and traffic splitting
Agent revisions and traffic splitting are now available in public preview. You can create immutable revisions of deployed agents, and split traffic between the different active revisions. This enables canary deployments and safe testing of new agent versions. For more information, see Manage revisions and traffic.
Manage and discover agent skills with Skill Registry
Manage and discover agent skills with the Skill Registry, in public preview. This secure, private, and low-latency repository stores skills as self-contained packages, including instructions, code, and documentation, to enhance agent abilities.
For more information, see:
Managed Agents API on Agent Platform released in Preview
The Managed Agents API on Agent Platform has been released in Preview.
This feature allows you to build and scale autonomous agents, including those built from configuration using the Antigravity harness. These agents run in a fully managed and isolated sandbox environment, equipped with tools and skills, and can be interacted with via a dedicated API.
For more information, see the following:
AI Content Detection API available
AI Content Detection API is available in Preview. For details see AI Content Detection.
Provisioned Throughput for Gemini now supports latency SLA
Provisioned Throughput now provides a tokens per second latency SLA, covering generation speed from the first returned token to the last.
For more information, see the Gemini Enterprise Agent Platform Online Inference Service Level Agreement (SLA).
Memory Bank and Sessions support for multi-regional and global endpoints is in Preview. For more, see Supported locations for agents in Agent Platform.
Priority PayGo is generally available (GA)
Priority PayGo is a consumption option that provides more consistent performance than standard PayGo without the upfront commitment of Provisioned Throughput. It is ideal for business-critical workloads with fluctuating or unpredictable traffic patterns.
For more information, see Priority PayGo.
You can purchase Provisioned Throughput for Gemma 4. To learn more, see the list of supported open models.
Gemini Distillation Service Early Access
We're introduction Gemini Distillation Service in Early Access. For information about requesting access, see Gemini Distillation Service.
Improvements to the Provisioned Throughput orders page have now made it possible to:
Gemini 3.1 Flash-Lite is now generally available
Our most cost-efficient Gemini model, 3.1 Flash-Lite, is out of preview and is now generally available. For technical information on this model, see the model information card.
You can purchase Provisioned Throughput for Gemini 3.1 Flash-Lite. To learn more, see the Provisioned Throughput overview.
Fixed an issue with Audio track extraction (Gemini Embedding 2 only) where the audio_track_extraction feature did not work. For more information, see Issue #504505771.
Agent Platform Gemini 3.1 Flash Image and Gemini 3 Pro Image
Gemini Enterprise Agent Platform Gemini 3.1 Flash Image Preview and Gemini 3 Pro Image Preview are introducing the following changes:
Improved transcription quality for Gemini Live API
You can now improve transcription quality for multilingual automatic speech
recognition (ASR) by using the
[input/output]_audio_transcription.language_codes field.
For more information, see Enable audio transcription for the session.
Asynchronous function calling with Live API
Asynchronous function calling is now available in public
preview in
Gemini Live API. You can run functions in parallel with conversation,
manage background processing, and handle function responses with policies
including SILENT, WHEN_IDLE, and INTERRUPT. For more information, see
Asynchronous function calling with
Gemini Live API.
Vertex AI to Gemini Enterprise Agent Platform naming changes
The table below lists all of the features that have been transitioned from Vertex AI and what their new names are in Agent Platform.
| Vertex AI name | Agent Platform name |
|---|---|
| Vertex AI Platform | Agent Platform |
| Generative AI on Vertex AI | Generative AI |
| Vertex AI Studio | Agent Studio |
| Vertex AI API | Agent Platform API |
| Vertex AI Model Garden | Model Garden |
| Vertex AI Models as a Service (MaaS) | MaaS |
| (Gemini/Veo) on Vertex AI | (Gemini/Veo) on Agent Platform |
| (Claude/Llama/DeepSeek/etc.) on Vertex AI | (Claude/Llama/DeepSeek/etc.), available on Agent Platform |
| Pre-trained APIs on Vertex AI | Pre-trained APIs on Agent Platform |
| (Provisioned Throughput/Pay-as-you-go/etc.) on Vertex AI | (Provisioned Throughput/Pay-as-you-go/etc.) on Agent Platform |
| Gemini Live API on Vertex AI | Gemini Live API on Agent Platform |
| Vertex AI Search | Agent Search |
| Vertex AI Search for Industry | Agent Search for Industry |
| Vertex AI Search for Commerce | Agent Search for Commerce |
| Recommendations from Vertex AI Search | Recommendations |
| Vertex AI Conversation | Agent Conversation |
| Vertex AI RAG Engine | RAG Engine |
| Vertex AI Vector Search | Vector Search |
| Vertex AI Vector Search 2.0 | Agent Retrieval |
| Vertex AI Agent Engine | Agent Runtime |
| Vertex AI Studio App Builder | App Builder in Agent Studio |
| Vertex AI Agent Engine Memory Bank | Agent Platform Memory Bank |
| Vertex AI Agent Engine Sessions | Agent Platform Sessions |
| Vertex AI Agent Engine Code Execution | Agent Platform Code Execution |
| Grounding with Google [...] in Vertex AI | Grounding with Google [...] in Agent Platform |
| Grounding with Google [...] in Vertex AI Search | Grounding with Google [...] in Agent Search |
| Grounding with Google [...] in Vertex AI Studio | Grounding with Google [...] in Agent Studio |
| Vertex AI Training | Agent Platform Managed Training |
| Vertex AI Serverless Training | Agent Platform Serverless Training |
| Vertex AI Training Clusters (VTC) | Managed Training Clusters |
| Ray on Vertex AI | Ray on Agent Platform |
| Reinforcement Learning from Human Feedback (RLHF)/Reinforcement Learning (RL) on Vertex AI | Reinforcement Learning on Agent Platform |
| Vertex AI Neural Architecture Search | Neural Architecture Search on Agent Platform |
| Vertex AI Prediction/Vertex AI Inference | Agent Platform Inference |
| Vertex AI Vision | Agent Platform Vision |
| Vertex AI Batch Inference | Agent Platform Batch Inference |
| Vertex AI Online Inference | Agent Platform Online Inference |
| Vertex AI Endpoints | Agent Platform Endpoints |
| Vertex AI Forecasting/Forecasting with AutoML | Forecasting on Agent Platform |
| Vertex AI Pipelines | Agent Platform Pipelines |
| Vertex AI Notebooks | Agent Platform Notebooks |
| Vertex AI Colab Enterprise | Agent Platform Colab Enterprise |
| Vertex AI Workbench | Agent Platform Workbench |
| Vertex AI Workbench Instances | Agent Platform Workbench Instances |
| Vertex AI Feature Store | Agent Platform Feature Store |
| Vertex AI Model Registry | Agent Platform Model Registry |
| Vertex AI Model Evaluation | Agent Platform Model Evaluation |
| Gen AI evaluation service on Vertex AI | Gen AI evals |
| Vertex AI AutoML (Vision/Video/Tables) | Agent Platform AutoML |
| Data Labeling on Vertex AI | Data Labeling |
| Vertex AI on GDC | Agent Platform on GDC |
| Vertex AI Experiments | Experiments on Agent Platform |
| Vertex AI Model Monitoring | Model Monitoring on Agent Platform |
| Vertex AI Media Studio | Agent Media Studio |
| Vertex AI | Agent Platform |
| Vertex AI Generative AI | Agent Platform Generative AI |
Initial release of Gemini Enterprise Agent Platform
This initial release includes (but is not limited to) the following releases or changes:
gemini-embedding-2) for General Availability.The following known issues affect Gemini Enterprise Agent Platform:
Audio track extraction (Gemini Embedding 2 only): Theaudio_track_extraction feature does not work. For more information, see Issue #504505771.
Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License, and code samples are licensed under the Apache 2.0 License. For details, see the Google Developers Site Policies. Java is a registered trademark of Oracle and/or its affiliates.
Last updated 2026-07-27 UTC.