Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI

NVIDIA released Nemotron 3.5 Content Safety, a 4B model that unifies multimodal, multilingual and custom-policy moderation with auditable reasoning traces.

Thursday June 4, 2026 Source: huggingface.co
TL;DR — Quick Answer

NVIDIA released Nemotron 3.5 Content Safety on June 4, 2026, a 4B-parameter guard model that evaluates a prompt, an optional image and an optional assistant response in a single pass. It covers 12 trained languages with zero-shot reach across about 140, supports custom enterprise policies, and can emit auditable reasoning traces via a THINK mode. NVIDIA reports roughly 85 percent average accuracy across multimodal benchmarks and about 92.7 percent combined on Multilingual Aegis and RTP-LX, with latency unchanged from Nemotron 3.

Key Takeaways

Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI — AI news article illustration

NVIDIA released Nemotron 3.5 Content Safety on June 4, 2026, the latest in a line of guard models that has grown from an English text classifier into a family spanning modalities, languages and inference modes. This version unifies multimodal input, multilingual reach, custom enterprise policy and auditable reasoning in a single call.

Unified Multimodal Evaluation

Earlier versions scored text and images separately. Nemotron 3.5 takes the user prompt, an optional image and an optional assistant response as one context window and issues one verdict, catching violations that only emerge from the interaction between request, image and reply.

Global Language Coverage

The model keeps the 12-language explicit training coverage of its predecessors while inheriting strong zero-shot generalization to roughly 140 languages from the Gemma 3 base. That matters for deployments in markets where safety training data is sparse.

Custom Policy and THINK Mode

Enterprise deployments rarely share one safety taxonomy, so Nemotron 3.5 accepts a custom policy specification and reasons over it. An optional THINK mode emits a concise reasoning trace before the verdict, which regulators and reviewers can audit. The taxonomy follows Aegis 2.0: 13 core categories plus 10 subcategories.

Benchmarks and Latency

Default mode latency is unchanged from Nemotron 3, so teams can keep real-time moderation on the fast path and run THINK-mode evaluation asynchronously.

Open Weights and an Open Dataset

The model is available on Hugging Face under the NVIDIA Open Model License, supports transformers, vLLM and SGLang, and ships with a multimodal, multilingual safety dataset that includes reasoning traces. It is also offered as a production-grade NIM microservice for teams that want a pre-packaged deployment.

Frequently Asked Questions

What is Nemotron 3.5 Content Safety?

It is NVIDIA's 4B-parameter multimodal, multilingual safety model released June 4, 2026, built on Gemma 3 4B IT and fine-tuned to classify prompts, images and assistant responses for policy violations.

Which languages does it support?

It is explicitly trained on 12 languages including English, French, Spanish, German, Chinese, Japanese, Korean and Arabic, and inherits zero-shot coverage of roughly 140 languages from its Gemma base model.

What is THINK mode?

THINK mode is an optional setting that makes the model output a short step-by-step reasoning trace before issuing its safe or unsafe verdict, which helps with compliance and human review.

Is Nemotron 3.5 Content Safety open?

Yes, the model is on Hugging Face under the NVIDIA Open Model License for research and commercial use, and NVIDIA also released the accompanying safety dataset.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news