Skip to content

NVIDIA Nemotron 3 Nano Omni Arrives on Amazon SageMaker JumpStart

NVIDIA Nemotron 3 Nano Omni, a compact 30B multimodal mixture of experts model reading video, audio, images, and text, is now one-click deployable on Amazon SageMaker JumpStart.

NVIDIA Nemotron 3 Nano Omni hybrid MoE architecture with audio, vision, and text encoders
Nemotron 3 Nano Omni · Credit: NVIDIA

NVIDIA Nemotron 3 Nano Omni landed on Amazon SageMaker JumpStart on April 28, 2026, giving builders a one-click way to deploy a model that reads video, audio, images, and text and returns text. The pitch is broad input coverage without a large serving footprint, aimed at teams that want a single model to handle several modalities but do not want to stand up heavyweight infrastructure to run it.

The model carries 30 billion total parameters with 3 billion active, a 30B A3B mixture of experts design NVIDIA builds on a Mamba2 Transformer hybrid architecture, pairing the Nemotron 3 Nano language backbone with the CRADIO v4-H vision encoder and the Parakeet speech encoder. It supports a 131,000-token context window, chain of thought reasoning, tool calling, JSON output, and word-level transcription timestamps, and runs in FP8 precision under the NVIDIA Open Model Agreement for commercial use. On the input side it takes MP4 video up to two minutes and 256 frames, WAV or MP3 audio up to one hour, and JPEG or PNG images. On SageMaker JumpStart, AWS lists ml.p4d.24xlarge or ml.p5.48xlarge as the deployment instances, launched from the JumpStart catalog.

AWS says the model delivers up to nine times higher throughput than alternative open omni models, a claim the announcement attributes to the architecture but does not benchmark against named competitors, so it is worth treating as a vendor figure until independent tests weigh in. The more durable point is the packaging: an NVIDIA multimodal model folded into a managed AWS deployment flow lowers the setup cost of trying omni models at all, which matters more for adoption than any single throughput number.

Share this story

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.