10 articles about LLM
Google DeepMind introduces Gemini 3.8 Flash alongside Flash Cyber, a cybersecurity-specialized variant, continuing the rapid Gemini release cadence into fall 2026.
DeepSeek releases V4.1-Flash, a 552B asymmetric MoE that outperforms V4-Pro on benchmarks at lower cost, with a KV cache needing just 1/4 the HBM of the previous generation.
New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.
Meta releases Muse Glimmer, a multimodal AI model designed to run on personal computers with 16GB+ RAM, marking a return to fully open weights.
xAI's Grok 4.6 launches with a focus on long-running agents and more ambitious interactive and visual work, introduced first through Cursor's platform.
Moonshot AI releases Kimi K3, its latest flagship model — landing as the largest open-weight model family competing with DeepSeek, Qwen, and GLM on agentic and reasoning benchmarks.
Meta launches Muse Spark 1.1, an AI model designed specifically for agentic tasks like tool usage, code execution, and multi-step workflows. The model targets enterprise automation and competes directly with Claude and GPT-5.6.
Grok 4.5 targets coding, agentic tasks, and knowledge work at $2 per million input and $6 per million output tokens — roughly half the price of leading rivals.
OpenAI announces GPT-5.5, positioning it as the most advanced model for real-world work and agentic AI.
OpenAI's latest model GPT-5 introduces advanced chain-of-thought reasoning, outperforming humans on complex benchmarks.