Skip to main content

Live Demo: Running Gemma 3 on NVIDIA Jetson Orin Nano

Google for DevelopersApril 2, 202518 min11,214 views
27 connections·32 entities in this video→

Jetson Orin Nano: Powerful Embedded AI

  • πŸš€ The Jetson Orin Nano is highlighted as a powerful, compact device capable of running state-of-the-art AI models.
  • ⚑ Performance has been significantly boosted, with trillions of operations per second and increased memory bandwidth, while maintaining stability and low power consumption (25W, with 15W and 7W modes available).
  • πŸ’‘ The device's capabilities are compared to earlier models, emphasizing its suitability for running large AI models like Gemma 3.

Setting Up Your Jetson Device

  • πŸ› οΈ The setup process involves unboxing the Jetson, connecting an SSD (e.g., 2TB), and using the SDK Manager on a Ubuntu computer to flash the operating system.
  • πŸ”Œ The device needs to be put into recovery mode and connected via USB to a host computer for flashing, which can take 30 minutes to an hour.
  • πŸ’» After flashing, the Jetson runs Ubuntu (JetPack) and requires connection to an external monitor and keyboard to complete setup and updates.

Running Gemma 3 on Jetson Orin Nano

  • 🧠 The demo showcases Gemma 3 (4 billion parameters) running on the Jetson Orin Nano, demonstrating its capability for generative AI tasks.
  • πŸ–ΌοΈ The model successfully performed image analysis, including identifying objects in a selfie (tangerine and banana) and describing an image of a mother cat and kitten.
  • ⏱️ Performance metrics showed approximately 14 tokens per second for text generation, highlighting its efficiency for an embedded device.

Advanced Use Cases and Models

  • πŸ—£οΈ The speaker demonstrated Gemma's ability to read text from an image (a timetable) and perform tasks like answering questions about a selfie.
  • πŸ”Š An example of using Gemma with voice and RAG (Retrieval Augmented Generation) was presented, with code available on GitHub, adaptable for Gemma 3.
  • πŸ€– AI agent simulations were also discussed, where two agents with different personalities could converse on the Jetson device, creating a continuous podcast-like output.
  • 🀯 Testing the 12 billion parameter Gemma model was possible, though slow, highlighting its potential for long inference times with patience, and suggesting the 27 billion parameter model might also work with significant swap memory.

Developer Engagement

  • 🀝 NVIDIA encourages developers to share their projects built with Jetson and Gemma, offering exposure and opportunities to connect with the research lab.
  • πŸ”— Resources like NVIDIA Jetson documentation and Google's AI platform documentation are recommended for integration and further learning.
Knowledge graph32 entities Β· 27 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover Β· drag to explore
32 entities
Chapters7 moments

Key Moments

Transcript66 segments

Full Transcript

Topics12 themes

What’s Discussed

Gemma 3Jetson Orin NanoNVIDIAGenerative AILarge Language ModelsEmbedded SystemsAI DemoSDK ManagerImage AnalysisAI AgentsRAGTokens per second
Smart Objects32 Β· 27 links
ProductsΒ· 15
MediasΒ· 3
CompaniesΒ· 3
ConceptsΒ· 10
EventΒ· 1