Apple to Use Nvidia Blackwell B200 GPUs in Google Cloud for Gemini-Powered Siri
Apple will process some Siri queries using Nvidia's Blackwell B200 data center GPUs deployed in Google Cloud, according to The Information. The company plans to use Nvidia's confidential compute feature to encrypt data during processing on the chips.
Apple to Use Nvidia Blackwell B200 GPUs in Google Cloud for Gemini-Powered Siri
Apple will process some Siri queries using Nvidia's Blackwell B200 data center GPUs deployed in Google Cloud, according to a report from The Information. The company plans to use Nvidia's confidential compute feature to encrypt data during processing on the chips.
Technical Infrastructure
According to sources familiar with the matter, Apple will tap into Google's fleet of Nvidia Blackwell B200 chips specifically. The Nvidia B200 is one of Nvidia's flagship data center GPUs for large-scale AI training and inference, positioned for running trillion-parameter models with improvements in inference speed, memory bandwidth, and multi-GPU scaling compared to the previous Hopper architecture.
Apple has approved the use of Nvidia's confidential compute technology in this setup. This hardware-based security system encrypts data while it is actively being processed by the GPUs. According to Nvidia, the feature "preserves the confidentiality and integrity of AI models deployed on Rubin, Blackwell, and Hopper GPUs" and allows "sensitive AI workloads to run securely at scale with near-native performance, even in shared or cloud environments."
Partnership Structure
Under the agreement with Google, certain user queries to a new version of Siri will run in Google Cloud on a licensed version of Google's Gemini model. This arrangement represents a departure from Apple's traditional strategy of controlling all critical components of its products.
The report notes it remains unclear how Apple's previously launched Private Cloud Compute server system will integrate with the upcoming Siri product launch. Apple announced Private Cloud Compute as a privacy-focused server infrastructure for processing sensitive AI workloads.
Deployment Timeline
Apple is expected to detail its AI plans at WWDC, though specific dates for the Gemini-powered Siri rollout have not been disclosed.
What This Means
Apple's reliance on Google Cloud and Nvidia hardware marks a significant shift from its vertical integration strategy. The company is betting that Nvidia's confidential compute can provide sufficient privacy guarantees for processing user data on third-party infrastructure—a claim that will face scrutiny given Apple's privacy positioning. The architecture suggests Apple either cannot or chooses not to run these workloads on its own silicon, raising questions about the performance requirements of Gemini integration versus the capabilities of Apple's server chips.
Related Articles
Google Launches Pics, a Prompt-Based Design Tool Built Into Workspace
Google is launching Pics, an AI image creation and editing tool powered by its Nano Banana model, positioning it as a prompt-first alternative to Canva and Adobe Express. The tool rolls out first in Google Docs and Slides for Workspace and Google AI Pro/Ultra subscribers, with Drive support to follow.
Google Rolls Out Pics Image Editor to AI Pro, Ultra, and Business Workspace Subscribers
Google has begun broadly rolling out Pics, its AI image generation and editing app built on Nano Banana, to Google AI Pro and Ultra subscribers plus Business and Enterprise Workspace plans. The app supports object segmentation, text translation in 29 languages, and 4K upscaling, with direct integration into Docs and Slides.
Google Switches Gemini Notebook to Compute-Based Usage Limits Starting September 2
Google is switching Gemini Notebook from fixed daily feature limits to compute-based usage limits that factor in prompt complexity, chat length, and sources used. The change, mirroring a May update to the Gemini app, rolls out to consumer accounts on September 2, 2026.
OpenRouter Adds Auto-Updating Alias for Zhipu AI's GLM Flash Model Family
Z.ai has published GLM Flash Latest on OpenRouter, a routing alias that automatically points to the newest checkpoint in the GLM Flash lineup. It supports a 1.31M token context window and multimodal text, image, and video input at $0.07 per 1M input tokens and $0.25 per 1M output tokens.
Comments
Loading...