45 articles in ai research
Google DeepMind introduces AlphaGenome Atlas, a database predicting the effects of every possible single nucleotide variant in the human genome — a high-resolution map of human DNA.
Google DeepMind unveils WeatherNext 3, its most advanced and accurate global weather AI model, pushing further toward replacing traditional numerical forecasting pipelines.
An internal OpenAI model, significantly more capable than GPT-6 Astra, produced a verified proof that 3D fluid motion can develop singularities — resolving one of the seven Millennium Prize Problems using 10,000 concurrent agents over 88 hours.
Cohere Labs aggregates 696,000 AI tools across 123,000 MCP servers into the largest open dataset of agentic AI, revealing where agents are actually being built — and where they aren't.
New research introduces Puffin-World, a unified multimodal model that works with native 3D world states — advancing AI's ability to understand and generate spatial environments.
New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.
Anthropic announces expanded support for scientists and research institutions working on AI safety, alignment, and capabilities research.
Allen AI (Ai2) expanded its partnership with Hugging Face to make its fully open models and datasets easier to find, customize, and build on. The Hugging Face Hub becomes the central platform for Ai2's open research work.
OpenAI teased a new Astra model family after an internal version reportedly solved 10 open math problems. The announcement signals OpenAI's next generation of models, coming as the company prepares its public IPO filing.
Meta released Muse Spark 1.2 on August 5, the latest update to its agentic AI model line. The release continues Meta's pivot from open-source Llama models to proprietary, metered API offerings.
DeepSeek's V4-Flash at $0.14/M input and $0.28/M output tokens topped global token usage at 7.1 trillion tokens weekly. A research firm found it over 100x cheaper to run than Claude Fable 5, with nine of the top ten models now Chinese.
Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter MoE model claiming parity with OpenAI and Anthropic flagships at one-fifth the cost. Priced at $2/M input and $6/M output with a 1M-token context, open weights arrive August 10.
Runway showcased its AI-generated characters at SIGGRAPH 2026, demonstrating how generative models can create consistent, expressive characters for film and interactive media. The presentation highlighted advances in character consistency and motion quality.
DeepSeek officially released V4-Flash with enhanced autonomous agent capabilities and reduced API costs, sustaining China's aggressive frontier pace in the AI price war. The model targets agentic workloads at significantly lower prices than US competitors.
Thinking Machines Lab released Inkling-Small, a 276B-parameter open-weight model with 12B active parameters that outperforms its 975B sibling on agentic coding. The model achieves 80.2% on SWE-Bench Verified and runs on consumer hardware.
Anthropic released a major update to the Model Context Protocol, moving to a stateless architecture. The change aims to simplify AI agent development but requires significant rewrites of existing integrations.
Moonshot AI released Kimi K3 as the largest openly available AI model with 2.8 trillion parameters. The model's open-weight release triggered a market sell-off in Chinese AI stocks, with Z.ai falling 30% and MiniMax dropping 16%.
Anthropic launched Claude Opus 5, scoring 43.3% on Frontier-Bench v0.1 — surpassing all competitors. Priced at $5/M input and $25/M output tokens, it delivers near-Fable 5 intelligence at half the cost per task, making it the new default on Claude Max.
OpenAI has published GPT-Red, a safety research model that autonomously probes and hardens other AI systems through self-improvement loops and adversarial testing.
NVIDIA's Nemotron open-data initiative releases over 10 trillion pre-training tokens and millions of post-training samples for building AI agents, plus region-specific personas and a Prompt Atlas.
Mistral released Leanstral 1.5 under Apache 2.0, a 119B model with 6B active parameters that hits 100 percent on miniF2F and finds real bugs in open-source code.
Anthropic launched Claude Science, an AI workbench for scientists with 60-plus curated skills, auditable artifacts, and on-demand compute, now in beta for Pro, Max, Team, and Enterprise plans.
Cohere releases North Mini Code, a 30B-parameter MoE model with 3B active parameters built for agentic software engineering, open-sourced under Apache 2.0.
A comprehensive comparison of Anthropic Fable, OpenAI Sora, and Runway Aleph — three competing approaches to AI-generated narrative and video.
Research shows Direct Preference Optimization can fix text degeneration in OCR models, cutting repetition-loop failures by an average 59.4 percent across five model families after fine-tuning.
OpenAI is expanding vetted, trusted access to GPT-Rosalind, its biodefense-focused model, giving government and research partners AI capabilities for genomics and pandemic preparedness.
Stability AI open-sources Stable Video Diffusion 4K, bringing high-resolution video generation to the open-source community with unprecedented quality.
Cohere released Command A+, a 218B parameter mixture-of-experts model, under Apache 2.0 license — targeting sovereign AI deployments with enterprise-grade agentic workflows on minimal hardware.
Dharma-AI research shows a specialized 3-billion-parameter model beat all commercial frontier APIs on structured OCR, scoring 0.911 versus 0.833 for Claude Opus 4.6 at about 52x lower cost.
An OpenAI general-purpose reasoning model has disproved a famous unsolved geometry problem posed by Paul Erdős in 1946, marking the first time AI has autonomously solved a prominent open problem central to a field of mathematics.
Google Beam's new experiment renders remote participants in true-to-life size and sound on HP Dimension's immersive display, closing the hybrid meeting inclusion gap with spatial audio.
Ai2 released OlmoEarth v1.1, a new family of Earth observation models that cut compute costs by up to 3x while keeping v1 performance, using a redesigned Sentinel-2 token.
An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.
Google DeepMind unveils an AI-powered cursor that understands context, letting users point and speak instead of writing long prompts.
NVIDIA and Ineffable Intelligence, the London lab founded by AlphaGo architect David Silver, are codesigning reinforcement-learning infrastructure for the new era of superlearners.
Startup Perceptron Inc. releases Mk1, a video analysis reasoning model that matches frontier AI performance at 80-90% lower cost than OpenAI, Anthropic, and Google.
ServiceNow details how its PipelineRL team moved RL rollout generation from vLLM V0 to V1 by fixing four backend gaps — logprob semantics, runtime defaults, inflight weight updates, and an fp32 lm_head — before changing the objective.
To fight test-set contamination, the Open ASR Leaderboard adds private evaluation datasets from Appen and DataoceanAI, keeping the default average WER on public sets while exposing optional private metrics.
Chinese AI company DeepSeek launches V4 model with superior efficiency, available under MIT License.
The Vera Rubin Stack aims to tackle latency, cost, and security bottlenecks for agentic and physical AI at scale.
New custom AI chips offer 3x faster training and 80% better performance per dollar, directly competing with Nvidia.
World leaders meet to decide on international AI rules as April 2026 deadline approaches for unified AI governance framework.
Nous Research released NousCoder-14B, a 14B open coding model that scores 67.87 percent on LiveCodeBench v6 after just four days of training on 48 Nvidia B200 GPUs.
Alibaba's Qwen team releases Qwen3Guard, its first safety guardrail model — fine-tuned from Qwen3 for precise prompt and response classification with risk levels, SOTA on major safety benchmarks.
Google's Gemini 2.0 Ultra outperforms GPT-4 on reasoning, coding, and multimodal tasks in independent evaluations.