Reinforcement-learning agents — AI systems that learn by trial and error — can convert computation into new knowledge. That is the focus of a new engineering-level collaboration between NVIDIA and Ineffable Intelligence, the London-based AI lab founded by AlphaGo architect David Silver in the wake of its emergence from stealth.
The Collaboration
“The next frontier of AI is superlearners — systems that learn continuously from experience,” said Jensen Huang, founder and CEO of NVIDIA. “We are thrilled to partner with Ineffable Intelligence to codesign the infrastructure for large-scale reinforcement learning as they push the frontier of AI.”
Silver, one of the pioneers of reinforcement learning, framed the shift plainly: researchers have largely solved the easier problem of AI — building systems that know what humans already know — but the harder problem is building systems that discover new knowledge for themselves by learning from experience.
Why RL Infrastructure Is Different
That kind of learning needs a powerful, highly optimized pipeline. Unlike pretraining, where a fixed dataset of human data flows through the system, reinforcement learning workloads generate their data on the fly. The system must act, observe, score and update continuously in tight loops, which puts pressure on interconnect, memory bandwidth and serving in ways pretraining does not.
This is where NVIDIA and Ineffable are focusing their technical work: building a pipeline that can feed reinforcement learning systems at scale. Engineers from both companies have teamed up to explore the best way to create this training pipeline.
The Technical Focus
The effort is starting on NVIDIA Grace Blackwell and will be among the first to explore the upcoming NVIDIA Vera Rubin platform. The goal is to understand the next generation of hardware and software needed as the AI world shifts beyond human data toward models that learn through simulation and experience.
Later stages of training also introduce forms of experience quite distinct from human language and human data, which may require novel model architectures and training algorithms.
What This Means
Getting this infrastructure right could unlock reinforcement learning at unprecedented scale in highly complex, rich environments — allowing agents to discover breakthroughs across all fields of knowledge. For NVIDIA, the partnership extends its platform from serving static training runs to the continuous, experience-driven compute loop that next-generation AI research will demand. For Silver’s lab, it removes the infrastructure bottleneck standing between reinforcement learning and its theoretical promise.