Hugging Face has released the Ettin Reranker family — six new CrossEncoder rerankers from 17M to 1B parameters that are state-of-the-art at their respective sizes on retrieval benchmarks. Each is built on the Ettin encoder suite from Johns Hopkins University and ships under the Apache 2.0 license with its full training recipe.
The Release
- 17M, 32M, 68M, 150M, 400M, and 1B parameter rerankers
- Backed by the Ettin ModernBERT-style encoders (unpadded attention, RoPE, GeGLU, 2T tokens of pre-training, 8K-token context)
- Drop-in
CrossEncodermodels for Sentence Transformers, usable in three lines of code
The headline: the 17M — the smallest model beats the 33M ms-marco-MiniLM-L12-v2 by +0.051 NDCG@10 on MTEB at roughly half the parameters, and the 68M model nearly matches the 596M Qwen3-Reranker-0.6B using a ninth of the compute.
What a Reranker Does
A cross-encoder takes a (query, document) pair and lets both texts attend to each other through every transformer layer — far more accurate than separate embedding vectors, but too expensive to run over a whole corpus. The standard pattern is retrieve-then-rerank: a fast embedder pulls the top-K candidates, then the reranker re-orders just those K. These models make that final re-ranks stage both cheaper and better.
Architecture
All six share a modular head: a Transformer module with sequence unpadding so Flash Attention 2 never computes on padding tokens, CLS pooling, and a two-dense-layer scoring head. Ablations showed CLS pooling edges out mean pooling despite ModernBERT’s local-window attention in most layers.
Speed
The 17M model processes 7,517 pairs per second on an H100 — nearly double the throughput of the MiniLM-L6-v2 it outperforms. The 150M model hits 3,237 pairs per second, a 2.3x gap over the two peer 150M ModernBERT rerankers that load through padded AutoModel paths. Combined bfloat16 plus unpadded FA2 delivers speedups from 1.71x on the 17M to 8.26x on the 1B.
The 1B Flagship
The 1B model scores 0.6114 on MTEB(eng, v2) Retrieval — within 0.0001 of its 1.54B teacher mxbai-rerank-large-v2 — while running at 928 pairs per second to the teacher’s 387. Distilling from a stronger teacher than the current one is, the author notes, the likely path to closing the remaining gap to Qwen3-Reranker-4B.
What This Means
Small, fast, open rerankers quietly decide the quality of every search, RAG pipeline, and recommendation system on the web. A 17M model that beats a 2x-larger legacy incumbent costs next to nothing to self-host — and with Apache 2.0 weights and a published recipe, this line of work makes best-in-class ranking a commodity.