12 articles in open source ai
DeepSeek releases V4.1-Flash, a 552B asymmetric MoE that outperforms V4-Pro on benchmarks at lower cost, with a KV cache needing just 1/4 the HBM of the previous generation.
Hugging Face introduces @huggingface/kernels, a library of 200+ WebGPU kernels enabling local AI inference directly in the browser without server-side compute.
Alibaba's Qwen team releases an open-weight coding model achieving GPT-5.6 level performance on a single consumer GPU with 24GB VRAM.
Meta releases Muse Glimmer, a multimodal AI model designed to run on personal computers with 16GB+ RAM, marking a return to fully open weights.
MiniMax launches H3, a general-purpose multimodal generation model that unifies text, image, video, and audio in one context — 15-second 2K video with native stereo sound, with open weights promised within days.
Moonshot AI releases Kimi K3, its latest flagship model — landing as the largest open-weight model family competing with DeepSeek, Qwen, and GLM on agentic and reasoning benchmarks.
OpenEnv, the agentic RL environment library, is now coordinated by a committee including PyTorch, NVIDIA, Microsoft, and Hugging Face as a common protocol layer for RL environments.
Hugging Face rebuilt the hf CLI so the same commands serve humans and coding agents, cutting token use up to 6x on multi-step Hub tasks and trimming agent tool calls by roughly 30 percent.
Holo3.1 is H company's open computer-use model family with first quantized checkpoints for fast local agents, lifting AndroidWorld scores to 79.3% and cutting step times.
PaddleOCR 3.5 adds Hugging Face Transformers as a document parsing inference backend, letting PP-OCRv5 and PaddleOCR-VL 1.5 models run wherever the engine parameter is set.
IBM Granite releases two Apache 2.0 multilingual embedding models on ModernBERT — a 97M model scoring 60.3 and a 311M model scoring 65.2 on MTEB Multilingual Retrieval.
Hermes Agent from Nous Research crossed 140,000 GitHub stars in under three months, bringing self-improving AI agents to NVIDIA RTX PCs, workstations, and DGX Spark.