AI Research

45 articles in ai research

AlphaGenome Atlas: A Predictive Map of Every Possible DNA Change
AI Research Sep 12, 2026

AlphaGenome Atlas: A Predictive Map of Every Possible DNA Change

Google DeepMind introduces AlphaGenome Atlas, a database predicting the effects of every possible single nucleotide variant in the human genome — a high-resolution map of human DNA.

WeatherNext 3: DeepMind's Most Accurate Weather AI Yet
AI Research Sep 12, 2026

WeatherNext 3: DeepMind's Most Accurate Weather AI Yet

Google DeepMind unveils WeatherNext 3, its most advanced and accurate global weather AI model, pushing further toward replacing traditional numerical forecasting pipelines.

OpenAI's AI Solves the Navier–Stokes Millennium Prize Problem — 10,000 Agents Resolve a 90-Year-Old Mathematics Mystery
AI Research Sep 8, 2026

OpenAI's AI Solves the Navier–Stokes Millennium Prize Problem — 10,000 Agents Resolve a 90-Year-Old Mathematics Mystery

An internal OpenAI model, significantly more capable than GPT-6 Astra, produced a verified proof that 3D fluid motion can develop singularities — resolving one of the seven Millennium Prize Problems using 10,000 concurrent agents over 88 hours.

Cohere Labs Maps the Agentic Task Ecosystem — Only 2.6% of 696,000 AI Tools Fully Automate Work Tasks
AI Research Sep 3, 2026

Cohere Labs Maps the Agentic Task Ecosystem — Only 2.6% of 696,000 AI Tools Fully Automate Work Tasks

Cohere Labs aggregates 696,000 AI tools across 123,000 MCP servers into the largest open dataset of agentic AI, revealing where agents are actually being built — and where they aren't.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
AI Research Sep 2, 2026

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

New research introduces Puffin-World, a unified multimodal model that works with native 3D world states — advancing AI's ability to understand and generate spatial environments.

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?
AI Research Sep 1, 2026

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?

New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.

Anthropic Expands Support for Scientists Working on AI Research
AI Research Aug 27, 2026

Anthropic Expands Support for Scientists Working on AI Research

Anthropic announces expanded support for scientists and research institutions working on AI safety, alignment, and capabilities research.

Allen AI and Hugging Face Expand Partnership for Open Science
AI Research Aug 6, 2026

Allen AI and Hugging Face Expand Partnership for Open Science

Allen AI (Ai2) expanded its partnership with Hugging Face to make its fully open models and datasets easier to find, customize, and build on. The Hugging Face Hub becomes the central platform for Ai2's open research work.

OpenAI Teases Astra Model Family After Internal Version Solves 10 Open Math Problems
AI Research Aug 5, 2026

OpenAI Teases Astra Model Family After Internal Version Solves 10 Open Math Problems

OpenAI teased a new Astra model family after an internal version reportedly solved 10 open math problems. The announcement signals OpenAI's next generation of models, coming as the company prepares its public IPO filing.

Meta Releases Muse Spark 1.2: Agentic Model Update Continues Proprietary Pivot
AI Research Aug 5, 2026

Meta Releases Muse Spark 1.2: Agentic Model Update Continues Proprietary Pivot

Meta released Muse Spark 1.2 on August 5, the latest update to its agentic AI model line. The release continues Meta's pivot from open-source Llama models to proprietary, metered API offerings.

DeepSeek V4-Flash Tops Global Token Usage: 7.1 Trillion Tokens Weekly, 9 of Top 10 Models Now Chinese
AI Research Aug 4, 2026

DeepSeek V4-Flash Tops Global Token Usage: 7.1 Trillion Tokens Weekly, 9 of Top 10 Models Now Chinese

DeepSeek's V4-Flash at $0.14/M input and $0.28/M output tokens topped global token usage at 7.1 trillion tokens weekly. A research firm found it over 100x cheaper to run than Claude Fable 5, with nine of the top ten models now Chinese.

Alibaba Unveils Qwen3.8-Max: 2.4 Trillion Parameters, Claims Parity With Western Frontier Models at One-Fifth Cost
AI Research Aug 3, 2026

Alibaba Unveils Qwen3.8-Max: 2.4 Trillion Parameters, Claims Parity With Western Frontier Models at One-Fifth Cost

Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter MoE model claiming parity with OpenAI and Anthropic flagships at one-fifth the cost. Priced at $2/M input and $6/M output with a 1M-token context, open weights arrive August 10.

Runway Brings AI-Generated Characters to SIGGRAPH's Biggest Stage
AI Research Jul 31, 2026

Runway Brings AI-Generated Characters to SIGGRAPH's Biggest Stage

Runway showcased its AI-generated characters at SIGGRAPH 2026, demonstrating how generative models can create consistent, expressive characters for film and interactive media. The presentation highlighted advances in character consistency and motion quality.

DeepSeek Releases V4-Flash With Stronger Agent Capabilities and Lower Costs
AI Research Jul 31, 2026

DeepSeek Releases V4-Flash With Stronger Agent Capabilities and Lower Costs

DeepSeek officially released V4-Flash with enhanced autonomous agent capabilities and reduced API costs, sustaining China's aggressive frontier pace in the AI price war. The model targets agentic workloads at significantly lower prices than US competitors.

Thinking Machines Releases Inkling-Small: 276B Open-Weight Model Beats Its 975B Sibling
AI Research Jul 30, 2026

Thinking Machines Releases Inkling-Small: 276B Open-Weight Model Beats Its 975B Sibling

Thinking Machines Lab released Inkling-Small, a 276B-parameter open-weight model with 12B active parameters that outperforms its 975B sibling on agentic coding. The model achieves 80.2% on SWE-Bench Verified and runs on consumer hardware.

MCP Goes Stateless: Anthropic's Bold Bet on Simplifying AI Tool Use
AI Research Jul 28, 2026

MCP Goes Stateless: Anthropic's Bold Bet on Simplifying AI Tool Use

Anthropic released a major update to the Model Context Protocol, moving to a stateless architecture. The change aims to simplify AI agent development but requires significant rewrites of existing integrations.

Moonshot Launches Kimi K3: Largest Open-Weight AI Model at 2.8 Trillion Parameters
AI Research Jul 26, 2026

Moonshot Launches Kimi K3: Largest Open-Weight AI Model at 2.8 Trillion Parameters

Moonshot AI released Kimi K3 as the largest openly available AI model with 2.8 trillion parameters. The model's open-weight release triggered a market sell-off in Chinese AI stocks, with Z.ai falling 30% and MiniMax dropping 16%.

Anthropic Releases Claude Opus 5: Matches Fable 5 Performance at Half the Price
AI Research Jul 24, 2026

Anthropic Releases Claude Opus 5: Matches Fable 5 Performance at Half the Price

Anthropic launched Claude Opus 5, scoring 43.3% on Frontier-Bench v0.1 — surpassing all competitors. Priced at $5/M input and $25/M output tokens, it delivers near-Fable 5 intelligence at half the cost per task, making it the new default on Claude Max.

OpenAI Introduces GPT-Red: A Self-Improving Safety Model for Robustness Testing
AI Research Jul 15, 2026

OpenAI Introduces GPT-Red: A Self-Improving Safety Model for Robustness Testing

OpenAI has published GPT-Red, a safety research model that autonomously probes and hardens other AI systems through self-improvement loops and adversarial testing.

NVIDIA Releases Nemotron Open Synthetic-Data Initiative With 10 Trillion Plus Pre-Training Tokens
AI Research Jul 3, 2026

NVIDIA Releases Nemotron Open Synthetic-Data Initiative With 10 Trillion Plus Pre-Training Tokens

NVIDIA's Nemotron open-data initiative releases over 10 trillion pre-training tokens and millions of post-training samples for building AI agents, plus region-specific personas and a Prompt Atlas.

Mistral Releases Leanstral 1.5: Saturates miniF2F at 100% for Formal Verification
AI Research Jul 2, 2026

Mistral Releases Leanstral 1.5: Saturates miniF2F at 100% for Formal Verification

Mistral released Leanstral 1.5 under Apache 2.0, a 119B model with 6B active parameters that hits 100 percent on miniF2F and finds real bugs in open-source code.

Anthropic Launches Claude Science: An AI Workbench for Researchers With NVIDIA BioNeMo Integration
AI Research Jun 30, 2026

Anthropic Launches Claude Science: An AI Workbench for Researchers With NVIDIA BioNeMo Integration

Anthropic launched Claude Science, an AI workbench for scientists with 60-plus curated skills, auditable artifacts, and on-demand compute, now in beta for Pro, Max, Team, and Enterprise plans.

Introducing North Mini Code: Cohere's First Model For Developers
AI Research Jun 9, 2026

Introducing North Mini Code: Cohere's First Model For Developers

Cohere releases North Mini Code, a 30B-parameter MoE model with 3B active parameters built for agentic software engineering, open-sourced under Apache 2.0.

How Anthropic Fable Compares to OpenAI Sora and Runway in AI Video and Storytelling
AI Research Jun 8, 2026

How Anthropic Fable Compares to OpenAI Sora and Runway in AI Video and Storytelling

A comprehensive comparison of Anthropic Fable, OpenAI Sora, and Runway Aleph — three competing approaches to AI-generated narrative and video.

Direct Preference Optimization Beyond Chatbots
AI Research Jun 3, 2026

Direct Preference Optimization Beyond Chatbots

Research shows Direct Preference Optimization can fix text degeneration in OCR models, cutting repetition-loop failures by an average 59.4 percent across five model families after fine-tuning.

OpenAI Launches Rosalind Biodefense: GPT-Powered AI for Pandemic Preparedness
AI Research Jun 1, 2026

OpenAI Launches Rosalind Biodefense: GPT-Powered AI for Pandemic Preparedness

OpenAI is expanding vetted, trusted access to GPT-Rosalind, its biodefense-focused model, giving government and research partners AI capabilities for genomics and pandemic preparedness.

Stability AI Releases Stable Video Diffusion 4K: Open-Source Video Generation at Ultra-High Resolution
AI Research May 30, 2026

Stability AI Releases Stable Video Diffusion 4K: Open-Source Video Generation at Ultra-High Resolution

Stability AI open-sources Stable Video Diffusion 4K, bringing high-resolution video generation to the open-source community with unprecedented quality.

Cohere Releases Command A+ as Open Source: 218B Parameter Enterprise Model Runs on Two H100 GPUs
AI Research May 22, 2026

Cohere Releases Command A+ as Open Source: 218B Parameter Enterprise Model Runs on Two H100 GPUs

Cohere released Command A+, a 218B parameter mixture-of-experts model, under Apache 2.0 license — targeting sovereign AI deployments with enterprise-grade agentic workflows on minimal hardware.

Specialization Beats Scale: A Strategic Variable Most AI Procurement Decisions Overlook
AI Research May 22, 2026

Specialization Beats Scale: A Strategic Variable Most AI Procurement Decisions Overlook

Dharma-AI research shows a specialized 3-billion-parameter model beat all commercial frontier APIs on structured OCR, scoring 0.911 versus 0.833 for Claude Opus 4.6 at about 52x lower cost.

OpenAI Model Disproves 80-Year-Old Erdős Geometry Conjecture in Landmark AI Research Breakthrough
AI Research May 21, 2026

OpenAI Model Disproves 80-Year-Old Erdős Geometry Conjecture in Landmark AI Research Breakthrough

An OpenAI general-purpose reasoning model has disproved a famous unsolved geometry problem posed by Paul Erdős in 1946, marking the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

A new experiment brings better group meetings to Google Beam
AI Research May 20, 2026

A new experiment brings better group meetings to Google Beam

Google Beam's new experiment renders remote participants in true-to-life size and sound on HP Dimension's immersive display, closing the hybrid meeting inclusion gap with spatial audio.

OlmoEarth v1.1: A more efficient family of Earth observation models
AI Research May 19, 2026

OlmoEarth v1.1: A more efficient family of Earth observation models

Ai2 released OlmoEarth v1.1, a new family of Earth observation models that cut compute costs by up to 3x while keeping v1 performance, using a redesigned Sentinel-2 token.

The Open Agent Leaderboard
AI Research May 18, 2026

The Open Agent Leaderboard

An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.

Google DeepMind Reimagines the Mouse Pointer for the AI Era
AI Research May 14, 2026

Google DeepMind Reimagines the Mouse Pointer for the AI Era

Google DeepMind unveils an AI-powered cursor that understands context, letting users point and speak instead of writing long prompts.

NVIDIA, Ineffable Intelligence Team Up to Build the Future of Reinforcement Learning Infrastructure
AI Research May 13, 2026

NVIDIA, Ineffable Intelligence Team Up to Build the Future of Reinforcement Learning Infrastructure

NVIDIA and Ineffable Intelligence, the London lab founded by AlphaGo architect David Silver, are codesigning reinforcement-learning infrastructure for the new era of superlearners.

Perceptron Mk1 Shocks AI Industry: Video Analysis Model 80-90% Cheaper Than Rivals
AI Research May 12, 2026

Perceptron Mk1 Shocks AI Industry: Video Analysis Model 80-90% Cheaper Than Rivals

Startup Perceptron Inc. releases Mk1, a video analysis reasoning model that matches frontier AI performance at 80-90% lower cost than OpenAI, Anthropic, and Google.

vLLM V0 to V1: Correctness Before Corrections in RL
AI Research May 6, 2026

vLLM V0 to V1: Correctness Before Corrections in RL

ServiceNow details how its PipelineRL team moved RL rollout generation from vLLM V0 to V1 by fixing four backend gaps — logprob semantics, runtime defaults, inflight weight updates, and an fp32 lm_head — before changing the objective.

Adding Benchmaxxer Repellant to the Open ASR Leaderboard
AI Research May 6, 2026

Adding Benchmaxxer Repellant to the Open ASR Leaderboard

To fight test-set contamination, the Open ASR Leaderboard adds private evaluation datasets from Appen and DataoceanAI, keeping the default average WER on public sets while exposing optional private metrics.

DeepSeek V4 Outperforms Claude Opus 4.6 in Benchmarks
AI Research Apr 26, 2026

DeepSeek V4 Outperforms Claude Opus 4.6 in Benchmarks

Chinese AI company DeepSeek launches V4 model with superior efficiency, available under MIT License.

NVIDIA and Google Cloud Unveil Next-Gen AI Infrastructure
AI Research Apr 25, 2026

NVIDIA and Google Cloud Unveil Next-Gen AI Infrastructure

The Vera Rubin Stack aims to tackle latency, cost, and security bottlenecks for agentic and physical AI at scale.

Google Cloud Launches TPU 8t and 8i AI Chips
AI Research Apr 22, 2026

Google Cloud Launches TPU 8t and 8i AI Chips

New custom AI chips offer 3x faster training and 80% better performance per dollar, directly competing with Nvidia.

UN Global Dialogue on AI Governance Faces Critical Decisions
AI Research Apr 19, 2026

UN Global Dialogue on AI Governance Faces Critical Decisions

World leaders meet to decide on international AI rules as April 2026 deadline approaches for unified AI governance framework.

Nous Research's NousCoder-14B is an open-source coding model landing right in the Claude Code moment
AI Research Jan 7, 2026

Nous Research's NousCoder-14B is an open-source coding model landing right in the Claude Code moment

Nous Research released NousCoder-14B, a 14B open coding model that scores 67.87 percent on LiveCodeBench v6 after just four days of training on 48 Nvidia B200 GPUs.

Qwen3Guard: Real-Time Safety for Your Token Stream
AI Research Sep 23, 2025

Qwen3Guard: Real-Time Safety for Your Token Stream

Alibaba's Qwen team releases Qwen3Guard, its first safety guardrail model — fine-tuned from Qwen3 for precise prompt and response classification with risk levels, SOTA on major safety benchmarks.

Google DeepMind's Gemini 2.0 Beats GPT-4 on Key Benchmarks
AI Research Apr 17, 2025

Google DeepMind's Gemini 2.0 Beats GPT-4 on Key Benchmarks

Google's Gemini 2.0 Ultra outperforms GPT-4 on reasoning, coding, and multimodal tasks in independent evaluations.