#inference

9 articles about inference

Hugging Face Releases @huggingface/kernels — 200+ WebGPU Kernels for Local AI in the Browser
Open Source AI Sep 1, 2026

Hugging Face Releases @huggingface/kernels — 200+ WebGPU Kernels for Local AI in the Browser

Hugging Face introduces @huggingface/kernels, a library of 200+ WebGPU kernels enabling local AI inference directly in the browser without server-side compute.

OpenAI's Jalapeño Chip Delivers Record-Breaking AI Inference Speed
AI Infrastructure Aug 25, 2026

OpenAI's Jalapeño Chip Delivers Record-Breaking AI Inference Speed

OpenAI's custom inference chip Jalapeño shows industry-leading speed and efficiency in early benchmarks, signaling a new era of purpose-built AI hardware.

Groq Among First to Deploy NVIDIA Groq 3 LPX and Vera Rubin NVL72 for Inference
AI Infrastructure Aug 24, 2026

Groq Among First to Deploy NVIDIA Groq 3 LPX and Vera Rubin NVL72 for Inference

Groq announces it will be among the first adopters of NVIDIA Groq 3 LPX with Vera Rubin NVL72, deploying through Dell Technologies to power its inference cloud with 3,400 tokens/sec on agentic workloads.

Cerebras Introduces CS-4 — Up to 30x Faster Inference, 1,000+ Tokens/sec for 10T-Parameter Models
AI Infrastructure Aug 18, 2026

Cerebras Introduces CS-4 — Up to 30x Faster Inference, 1,000+ Tokens/sec for 10T-Parameter Models

Cerebras unveils CS-4, its fourth-generation system with three WSE-3 Turbo processors, delivering up to 30x faster inference than GPUs and 10x more throughput per watt than CS-3.

OpenAI Launches Ultrafast Mode — GPT-5.6 Sol Runs 14x Faster on Cerebras
AI Infrastructure Aug 13, 2026

OpenAI Launches Ultrafast Mode — GPT-5.6 Sol Runs 14x Faster on Cerebras

OpenAI partners with Cerebras to deliver GPT-5.6 Sol at 750 tokens per second — 14x faster than standard inference — enabling real-time AI applications.

OpenAI Slashes GPT-5.6 Luna Prices by 80% in Aggressive Pricing Move
AI Business Jul 30, 2026

OpenAI Slashes GPT-5.6 Luna Prices by 80% in Aggressive Pricing Move

OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens, matching DeepSeek's Chinese pricing. The move signals intensifying competition in the AI API market and a race to the bottom on inference costs.

Cerebras Runs Trillion-Parameter AI Model Nearly 7x Faster Than GPU Clouds in Landmark Inference Test
AI Hardware May 21, 2026

Cerebras Runs Trillion-Parameter AI Model Nearly 7x Faster Than GPU Clouds in Landmark Inference Test

Cerebras Systems announced it runs Moonshot AI's trillion-parameter Kimi K2.6 model at 981 tokens per second, 6.7x faster than the fastest GPU cloud provider, in an independently verified benchmark.

Unlocking Asynchronicity in Continuous Batching
AI Infrastructure May 14, 2026

Unlocking Asynchronicity in Continuous Batching

Hugging Face shows how asynchronous continuous batching, using CUDA streams, events, and double-buffered slots, lifted GPU utilization from 76 to 99.4 percent and cut generation time by 22 percent.

vLLM V0 to V1: Correctness Before Corrections in RL
AI Research May 6, 2026

vLLM V0 to V1: Correctness Before Corrections in RL

ServiceNow details how its PipelineRL team moved RL rollout generation from vLLM V0 to V1 by fixing four backend gaps — logprob semantics, runtime defaults, inflight weight updates, and an fp32 lm_head — before changing the objective.