17 articles in ai infrastructure
OpenAI's custom inference chip Jalapeño shows industry-leading speed and efficiency in early benchmarks, signaling a new era of purpose-built AI hardware.
Groq announces it will be among the first adopters of NVIDIA Groq 3 LPX with Vera Rubin NVL72, deploying through Dell Technologies to power its inference cloud with 3,400 tokens/sec on agentic workloads.
Cerebras unveils CS-4, its fourth-generation system with three WSE-3 Turbo processors, delivering up to 30x faster inference than GPUs and 10x more throughput per watt than CS-3.
OpenAI partners with Cerebras to deliver GPT-5.6 Sol at 750 tokens per second — 14x faster than standard inference — enabling real-time AI applications.
Anthropic confirmed it is building an in-house chip design team, hiring engineers with silicon experience to co-design hardware and models so Claude runs faster and more efficiently. The move follows similar efforts by OpenAI, Google, and Meta to reduce dependency on Nvidia.
Google is expanding its Virginia footprint with new community investments that support thousands of local jobs, fund electrical apprenticeship training, and launch a $15 million Energy Impact Fund.
A year after declaring itself an AI maker, not an AI taker, the UK is scaling sovereign compute with doubled cloud providers and Isambard-AI on 5,400 GH200 Superchips.
NVIDIA and LG Group announced an AI factory on the NVIDIA DSX platform spanning robotics, autonomous driving, data centers and GPU cloud services.
Jensen Huang visited Seoul to meet South Korea's AI builders, unveiling gigawatt-scale AI factory plans with NAVER, LG, SK and Hyundai, plus robotics, memory and RTX Spark partnerships.
SoftBank announced a €75 billion investment to develop 5 gigawatts of AI data center capacity across three sites in France, the largest single AI infrastructure commitment in European history.
Live updates from NVIDIA GTC Taipei at COMPUTEX covering Jensen Huang's keynote, Vera Rubin in full production, Nemotron 3 Ultra, RTX Spark and the new AI factory era.
Google is deepening its Missouri roots with a new Montgomery County data center, over 500 MW of new capacity with Ameren, a $20 million Energy Impact Fund, and local workforce training.
NVIDIA and Google Cloud said their joint developer community has passed 100,000 members, adding JAX and Dynamo learning paths plus Gemma, Nemotron and SynthID resources.
At Dell Technologies World, Jensen Huang joined Michael Dell to unveil the Dell AI Factory refresh — Vera Rubin-based PowerEdge XE9812 servers with up to 10x lower cost-per-token for agentic inference.
Hugging Face shows how asynchronous continuous batching, using CUDA streams, events, and double-buffered slots, lifted GPU utilization from 76 to 99.4 percent and cut generation time by 22 percent.
AWS engineers detail the infrastructure building blocks for foundation model training and inference, from B300 GPUs and EFA networking to Slurm, Kubernetes and PyTorch.
Railway, the San Francisco cloud platform that amassed two million developers without marketing, raised a $100 million Series B led by TQ Ventures to fund its sub-second AI-native deployments and own data centers.