#LLM

10 articles about LLM

Gemini 3.8 Flash and Flash Cyber: Google's Fast Models Get a Security Sibling
AI Security Sep 12, 2026

Gemini 3.8 Flash and Flash Cyber: Google's Fast Models Get a Security Sibling

Google DeepMind introduces Gemini 3.8 Flash alongside Flash Cyber, a cybersecurity-specialized variant, continuing the rapid Gemini release cadence into fall 2026.

DeepSeek V4.1-Flash: Smarter, Faster, Cheaper — and It Beats Their Own Flagship
Open Source AI Sep 10, 2026

DeepSeek V4.1-Flash: Smarter, Faster, Cheaper — and It Beats Their Own Flagship

DeepSeek releases V4.1-Flash, a 552B asymmetric MoE that outperforms V4-Pro on benchmarks at lower cost, with a KV cache needing just 1/4 the HBM of the previous generation.

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?
AI Research Sep 1, 2026

BenchMIRT Asks: What Are LLM Benchmarks Actually Measuring?

New research examines whether LLM benchmark scores reflect real model capability, revealing how measurement artifacts and benchmark design choices distort our picture of model performance.

Meta Returns to Open Weights with Muse Glimmer
Open Source AI Aug 14, 2026

Meta Returns to Open Weights with Muse Glimmer

Meta releases Muse Glimmer, a multimodal AI model designed to run on personal computers with 16GB+ RAM, marking a return to fully open weights.

Grok 4.6: Built for Long-Running Agents and Visual Work
AI Models Aug 12, 2026

Grok 4.6: Built for Long-Running Agents and Visual Work

xAI's Grok 4.6 launches with a focus on long-running agents and more ambitious interactive and visual work, introduced first through Cursor's platform.

Kimi K3: Moonshot AI's Flagship Joins the Open-Weights Arena
Open Source AI Jul 16, 2026

Kimi K3: Moonshot AI's Flagship Joins the Open-Weights Arena

Moonshot AI releases Kimi K3, its latest flagship model — landing as the largest open-weight model family competing with DeepSeek, Qwen, and GLM on agentic and reasoning benchmarks.

Meta Releases Muse Spark 1.1: A Purpose-Built Model for Agentic Tasks and Tool Use
AI Models Jul 9, 2026

Meta Releases Muse Spark 1.1: A Purpose-Built Model for Agentic Tasks and Tool Use

Meta launches Muse Spark 1.1, an AI model designed specifically for agentic tasks like tool usage, code execution, and multi-step workflows. The model targets enterprise automation and competes directly with Claude and GPT-5.6.

xAI Releases Grok 4.5: Opus-Class Performance at Half the Price of Competitors
AI Models Jul 8, 2026

xAI Releases Grok 4.5: Opus-Class Performance at Half the Price of Competitors

Grok 4.5 targets coding, agentic tasks, and knowledge work at $2 per million input and $6 per million output tokens — roughly half the price of leading rivals.

OpenAI Launches GPT-5.5 as Most Advanced Model Yet
AI Models Apr 26, 2026

OpenAI Launches GPT-5.5 as Most Advanced Model Yet

OpenAI announces GPT-5.5, positioning it as the most advanced model for real-world work and agentic AI.

OpenAI Launches GPT-5 With Breakthrough Reasoning
AI Models Apr 18, 2025

OpenAI Launches GPT-5 With Breakthrough Reasoning

OpenAI's latest model GPT-5 introduces advanced chain-of-thought reasoning, outperforming humans on complex benchmarks.