Nvidia has rolled out yet another model and promised it will make multimodal agents look clever without burning the budget.
Dubbed the Nvidia Nemotron 3 Nano Omni, the open multimodal model is billed as bringing these capabilities together into a single system, and it is meant to cover video, audio, images, and text in one go.
Nvidia said: “This best-in-class model gives enterprises and developers a production path for more efficient and accurate multimodal AI agents with full deployment flexibility and control.”
The outfit claims Nemotron 3 Nano Omni sets a fresh efficiency bar for open multimodal models, with leading accuracy and low cost. It says it tops six leaderboards for complex document intelligence, plus video and audio understanding.
AI and software firms already adopting the model include Aible, Applied Scientific Intelligence, Eka Care, Foxconn, H Company, Palantir and Pyler. Dell, DocuSign, Infosys, K-Dense, Lila, Oracle and Zefr are still poking at it.
Nemotron 3 Nano Omni bundles vision and audio encoders into a 30B-A3B hybrid mixture-of-experts setup, so it can bin the old “one model per sense” mess. Nvidia says that drives inference efficiency at scale without turning responsiveness into a slideshow.
The outfit claims the model achieves nine times the throughput of other open omni models with the same interactivity, which is a tidy claim if it survives contact with real workloads. Lower costs and better scalability are the sales pitch, with quality supposedly staying put.
In agent systems, Nemotron 3 Nano Omni is pitched as a sub-agent workhorse that can sit next to proprietary cloud models or other Nvidia Nemotron open models. Nvidia namechecks Nemotron 3 Super for high-frequency execution and Nemotron 3 Ultra for complex planning, then nods at proprietary models from other providers.
For computer-use agents, Nvidia says the model can run the perception loop to navigate graphical user interfaces, reason over on-screen content, and track UI state over time. H Company’s latest computer usage agent uses a native input resolution of 1920 x 1080 pixels, and Nvidia says early OSWorld results show a significant leap in handling fiddly interfaces.
For document intelligence, Nvidia is aiming at charts, tables, screenshots and mixed media, where the agent has to keep visual structure and text aligned. It is pitching this for enterprise analysis and compliance work.
For audio and video understanding, Nvidia wants a single reasoning stream that ties what was said, shown and documented together instead of dumping disconnected summaries everywhere, which is where enterprise dashboards go to die.

TOPICS:
Raspberry Pi shrugs off RAMageddon
Raspberry Pi has shrugged off soaring memory prices to post record first-half revenue and profit…
Musk stuffs more Nvidia kit into Colossus 2
Elon Musk plans to roughly double Colossus 2’s already enormous collection of advanced Nvidia AI…
ASML’s European sales fall to zero
Europe wants a stronger chip industry, but its biggest chipmaking equipment supplier says customers on…
Oracle invokes force majeure at Stargate data centre
Oracle has sent a force majeure notice to the developer of its New Mexico Stargate…
Hot topics
Latest reviews
Pixel Watch 4 45mm A nice evolution
Review: Faster charging, longer battery life, and a larger screen. A long time ago, I set…
GEEKOM A5 Pro Mini PC is a great budget mini PC
Pre-“RAMpocalypse” Pricing and Solid Everyday Performance Mini PCs have always been close to my heart.…
Baseus 65W Charger 2 PRO affordably charges notebook and two more device
Review: Baseus 65W GaN Charger 2 PRO Quick Charge 4.0 3.0 Type C PD USB Charger…






