Nvidia
13 articles tagged with Nvidia
Nvidia Claims Groq 3 LPX Hits 3,400 Tokens/Sec, 4x Cerebras — But Needs 64 Chips to Do It
Nvidia's new Groq 3 LPX inference accelerator hit 3,400 tokens per second on Gemma 4 31B, a figure the company says is four times faster than Cerebras. Experts note the benchmark uses at least 64 LPX chips versus Cerebras' one or two accelerators, making the comparison far less clean than it appears.
Nvidia Reportedly Building Trillion-Parameter Nemotron 4 to Match Chinese Open Models
Nvidia is reportedly building Nemotron 4, an open-weight model with at least one trillion parameters — double the size of Nemotron 3 Ultra. The company has tripled its cloud spending on in-house training to $28 billion through 2031, with an earliest possible release this fall.
Nvidia Releases Nemotron 3.5 Lightning: A 31.6B-Parameter Open Model Built for Speed, Not Peak Intelligence
Nvidia's Nemotron 3.5 Lightning, a 31.6B-parameter open-weight model with only 3.6B active parameters, matches OpenAI's gpt-oss-120b on the Artificial Analysis Intelligence Index while delivering the fastest throughput in its class at nearly 670 tokens per second. The model posts especially large gains on agentic benchmarks, beating both gpt-oss-120b and the larger Nemotron 3 Super.
Mira Murati's Thinking Machines releases Inkling, 975B-parameter open-weight model trained on 45T tokens
Thinking Machines Lab released Inkling, a 975-billion-parameter mixture-of-experts model that uses 41 billion active parameters per task. The open-weight model was trained on 45 trillion tokens across text, image, audio, and video, marking the first public release from Mira Murati's AI startup.
Apple releases AFM 3 lineup: 20B-parameter on-device model and cloud AI running on Google's Nvidia infrastructure
Apple announced five third-generation foundation models at WWDC26, headlined by AFM 3 Core Advanced—a 20-billion-parameter sparse model that runs on-device by activating only 1-4 billion parameters at a time. For the first time, Apple extended Private Cloud Compute to third-party infrastructure, with AFM 3 Cloud Pro running on Nvidia GPUs in Google Cloud.
Apple's AFM Cloud Pro Model Runs on Nvidia GPUs in Google Cloud, Execs Confirm
Apple executives revealed at WWDC that its most advanced AI model, AFM Cloud Pro, runs on Nvidia GPUs deployed in Google's cloud infrastructure while maintaining Apple's privacy guarantees. The company disclosed it uses Google's Gemini frontier models to refine its own custom models built for Apple Silicon.
Apple deploys 1.2T-parameter Gemini model on Nvidia Blackwell GPUs for rebuilt Siri
Apple announced at WWDC 2026 that the rebuilt Siri runs on a custom 1.2-trillion-parameter model based on Google's Gemini technology, hosted on Google Cloud servers powered by Nvidia Blackwell B200 GPUs. The company unveiled a three-tier privacy architecture and five new Apple Foundation Models to handle queries across device, private cloud, and Google Cloud infrastructure.
Apple deploys Google-trained models in iOS 27 Siri via Private Cloud Compute on Nvidia GPUs
Apple's senior vice president Craig Federighi disclosed that iOS 27's Siri AI uses a family of third-generation Apple Foundation Models trained with outputs from Google's Gemini frontier models. The most capable model, AFM Cloud Pro, runs on Nvidia GPUs in Google's cloud infrastructure while maintaining Apple's Private Cloud Compute privacy architecture.
Nvidia releases Nemotron 3 Ultra: 550B-parameter MoE model with 1M context window for agentic workflows
Nvidia has released Nemotron 3 Ultra, a 550-billion parameter mixture-of-experts model with 55 billion active parameters and support for up to 1 million token context windows. The model uses a hybrid Transformer-Mamba architecture and is designed specifically for long-running agentic workflows including agent orchestration, coding agents, and complex enterprise tasks.
Apple to Use Nvidia Blackwell B200 GPUs in Google Cloud for Gemini-Powered Siri
Apple will process some Siri queries using Nvidia's Blackwell B200 data center GPUs deployed in Google Cloud, according to The Information. The company plans to use Nvidia's confidential compute feature to encrypt data during processing on the chips.
OpenAI reasoning model solves 80-year math problem as Anthropic hits $10.9B quarterly revenue
In a two-hour span Wednesday, OpenAI announced its reasoning model autonomously solved an 80-year-old geometry problem while Anthropic reported it's on track for $10.9 billion in Q2 revenue with $559 million in operating profit—two years ahead of internal projections. The developments came alongside Nvidia's $81.6 billion quarter, Anthropic's $1.25 billion monthly SpaceX compute deal, and a White House AI executive order signing.
Alibaba's Qwen AI integrates with BYD, Volkswagen and 8 other Chinese automakers for voice-controlled services
Alibaba announced Friday that its Qwen AI model will be integrated into vehicles from 10 Chinese automakers including BYD, Geely, Li Auto, and SAIC Volkswagen. The system runs on Nvidia's automotive chip platform and allows drivers to order food delivery, book hotels, and make payments through voice commands, even with limited network connectivity.
OpenAI raises $110B in largest private funding round ever, valued at $730B
OpenAI has raised $110 billion in what is now the largest private funding round in history, according to the company. The round values OpenAI at $730 billion and includes a $50 billion investment from Amazon alongside $30 billion each from Nvidia and SoftBank.