Skip to main content

Running Gemma 3 Models Locally on Mac & Windows with Ollama

Google for DevelopersApril 5, 20258 min17,086 views
22 connections·23 entities in this video→

Ollama: Easy Local AI Model Deployment

  • πŸ’‘ Ollama is presented as the simplest method for running AI models locally on your machine.
  • πŸš€ The project is open-source, allowing users to download binaries or build from source on GitHub.
  • πŸ“ˆ Ollama has seen significant adoption, with tens of millions of developers using it and over 150 million models downloaded.

Under the Hood: Ollama's Inference Engine

  • βš™οΈ At its core, Ollama functions as a portable inference engine that manages model execution.
  • πŸ’» It handles scheduling models across various hardware like CPUs, GPUs, and TPUs, optimizing memory usage.
  • πŸ› οΈ Significant effort has been made to ensure Ollama runs smoothly on Mac and Windows environments, as well as single-GPU Linux systems.

Gemma 3 Integration and Demo

  • ✨ Ollama has collaborated with DeepMind to bring Gemma 3 models to its platform, available in four parameter sizes.
  • πŸ’¬ A demonstration shows running the Gemma 3 4B model on a MacBook Pro, showcasing fast inference without needing API keys.
  • πŸ–ΌοΈ Gemma 3's multimodal capabilities are highlighted, successfully interpreting images and even deciphering difficult handwriting from a doctor's note.
  • πŸ’¨ The 1B parameter model offers even faster inference and requires minimal VRAM (2GB), making it accessible on most machines.
  • πŸ“ˆ The 27B model, while requiring more VRAM, is still usable on medium to high-end MacBooks and other GPU-enabled devices.

Beyond Local: Ollama on Google Cloud

  • ☁️ Ollama extends beyond local installations, partnering with Google Cloud Run for serverless deployments.
  • ⚑ This allows users to access GPUs on Google Cloud Run, scaling from zero to hundreds of GPUs with rapid startup times (under five seconds).
  • βœ… Ollama is available across Windows, Mac, and Linux, enabling powerful AI access both locally and in the cloud.
Knowledge graph23 entities Β· 22 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover Β· drag to explore
23 entities
Chapters4 moments

Key Moments

Transcript32 segments

Full Transcript

Topics12 themes

What’s Discussed

OllamaGemma 3Local AIMacWindowsInference EngineOpen SourceAI ModelsMultimodal AIGoogle Cloud RunServerlessGPU
Smart Objects23 Β· 22 links
ProductsΒ· 10
PeopleΒ· 2
CompaniesΒ· 2
ConceptsΒ· 4
LocationsΒ· 4
MediaΒ· 1