NVIDIA Nemotron 3 Nano Omni landed on Amazon SageMaker JumpStart on April 28, 2026, giving builders a one-click way to deploy a model that reads video, audio, images, and text and returns text. The pitch is broad input coverage without a large serving footprint, aimed at teams that want a single model to handle several modalities but do not want to stand up heavyweight infrastructure to run it.
The model carries 30 billion total parameters with 3 billion active, a 30B A3B mixture of experts design NVIDIA builds on a Mamba2 Transformer hybrid architecture, pairing the Nemotron 3 Nano language backbone with the CRADIO v4-H vision encoder and the Parakeet speech encoder. It supports a 131,000-token context window, chain of thought reasoning, tool calling, JSON output, and word-level transcription timestamps, and runs in FP8 precision under the NVIDIA Open Model Agreement for commercial use. On the input side it takes MP4 video up to two minutes and 256 frames, WAV or MP3 audio up to one hour, and JPEG or PNG images. On SageMaker JumpStart, AWS lists ml.p4d.24xlarge or ml.p5.48xlarge as the deployment instances, launched from the JumpStart catalog.
AWS says the model delivers up to nine times higher throughput than alternative open omni models, a claim the announcement attributes to the architecture but does not benchmark against named competitors, so it is worth treating as a vendor figure until independent tests weigh in. The more durable point is the packaging: an NVIDIA multimodal model folded into a managed AWS deployment flow lowers the setup cost of trying omni models at all, which matters more for adoption than any single throughput number.













