Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context — Best Sub-100M Retrieval Quality

IBM Granite releases two Apache 2.0 multilingual embedding models on ModernBERT — a 97M model scoring 60.3 and a 311M model scoring 65.2 on MTEB Multilingual Retrieval.

Thursday May 14, 2026 Source: huggingface.co
TL;DR — Quick Answer

IBM Granite released two Apache 2.0 multilingual embedding models on May 14, 2026: the 97M-parameter granite-embedding-97m-multilingual-r2, which scores 60.3 on MTEB Multilingual Retrieval — the best result for any open model under 100M parameters, 9.4 points ahead of the next competitor — and the 311M granite-embedding-311m-multilingual-r2, which scores 65.2 with Matryoshka dimension support. Both are built on ModernBERT, cover 200+ languages with enhanced retrieval for 52 plus code across 9 languages, and extend context from 512 to 32,768 tokens — a 64x increase. Both ship with ONNX and OpenVINO weights for CPU inference.

Key Takeaways

Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context — Best Sub-100M Retrieval Quality — AI news article illustration

Multilingual embedding models face a persistent tension: broad language coverage usually costs model size, and small models usually sacrifice languages. IBM’s Granite Embedding Multilingual R2 release, published May 14, 2026, narrows that gap considerably — with two Apache 2.0 models that push small-model quality up and long-document retrieval way up.

Two Models, One Family

What Changed from R1

The R1 models were XLM-RoBERTa encoders with a 512-token window. R2 is a ground-up rebuild on ModernBERT: alternating attention lengths improve long-sequence throughput, rotary position embeddings enable the 32K window without interpolation hacks, and Flash Attention 2.0 speeds up GPU encoding. The headline number is context: 32,768 tokens, a 64x jump, driving +31.3 to +34.0 point gains on LongEmbed — R1 could only judge a legal contract by its first page.

Training and Governance

The 311M model trains through knowledge distillation from Granite and Mistral instruct teachers, contrastive fine-tuning across 52 languages plus code, model merging, and Matryoshka objectives. The 97M model adds vocabulary pruning — 180K tokens versus 262K — and distillation from a Granite 4.1 8B teacher. Training uses IBM-curated data with governance review, deliberately avoiding MS-MARCO and non-commercial-licensed datasets for enterprise-safe use.

Performance and Deployment

The 97M model encodes over 2,500 documents per second on a single H100; the 311M runs about 1,800 docs/sec. Matryoshka truncation on the 311M model cuts 768 to 256 dimensions with only a 0.5-point quality drop. Both cover 200+ languages with enhanced retrieval in 52, plus code retrieval across 9 programming languages, and ship with ONNX and OpenVINO weights, vLLM endpoints, and drop-in one-line swaps for LangChain, LlamaIndex, Haystack, and Milvus.

What This Means

The framework-default embedder most RAG apps silently use is now 23.7 points behind a model that is smaller, covers 200+ languages, and reads whole documents. For framework maintainers, the Granite team is explicitly inviting adoption as a default — one line of code, no API changes, and every user gains multilingual retrieval overnight.

Frequently Asked Questions

What is Granite Embedding Multilingual R2?

A pair of Apache 2.0-licensed multilingual embedding models from IBM Granite built on ModernBERT: a compact 97M-parameter model with 384-dimensional embeddings and a full-size 311M-parameter model with 768-dimensional embeddings and Matryoshka support.

How good is the 97M Granite embedding model?

It scores 60.3 on MTEB Multilingual Retrieval — the highest score IBM found for any open multilingual embedding model under 100M parameters, 9.4 points ahead of multilingual-e5-small at 50.9, and 23.7 points ahead of the widely used paraphrase-multilingual-MiniLM-L12-v2 default.

What context length do the Granite R2 embedding models support?

Both support up to 32,768 tokens — a 64x increase over their R1 predecessors' 512-token window — enabling retrieval over full legal contracts, technical manuals, and research papers.

Are the Granite R2 embedding models open source?

Yes, both are released under Apache 2.0, trained without MS-MARCO, and ship with ONNX and OpenVINO weights for CPU-optimized inference, plus drop-in compatibility with sentence-transformers, LangChain, LlamaIndex, Haystack, and Milvus.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news