9 articles about inference
Hugging Face introduces @huggingface/kernels, a library of 200+ WebGPU kernels enabling local AI inference directly in the browser without server-side compute.
OpenAI's custom inference chip Jalapeño shows industry-leading speed and efficiency in early benchmarks, signaling a new era of purpose-built AI hardware.
Groq announces it will be among the first adopters of NVIDIA Groq 3 LPX with Vera Rubin NVL72, deploying through Dell Technologies to power its inference cloud with 3,400 tokens/sec on agentic workloads.
Cerebras unveils CS-4, its fourth-generation system with three WSE-3 Turbo processors, delivering up to 30x faster inference than GPUs and 10x more throughput per watt than CS-3.
OpenAI partners with Cerebras to deliver GPT-5.6 Sol at 750 tokens per second — 14x faster than standard inference — enabling real-time AI applications.
OpenAI cut GPT-5.6 Luna API prices by 80% to $0.20 per million input tokens, matching DeepSeek's Chinese pricing. The move signals intensifying competition in the AI API market and a race to the bottom on inference costs.
Cerebras Systems announced it runs Moonshot AI's trillion-parameter Kimi K2.6 model at 981 tokens per second, 6.7x faster than the fastest GPU cloud provider, in an independently verified benchmark.
Hugging Face shows how asynchronous continuous batching, using CUDA streams, events, and double-buffered slots, lifted GPU utilization from 76 to 99.4 percent and cut generation time by 22 percent.
ServiceNow details how its PipelineRL team moved RL rollout generation from vLLM V0 to V1 by fixing four backend gaps — logprob semantics, runtime defaults, inflight weight updates, and an fp32 lm_head — before changing the objective.