Skip to main content

Exploring Gemma 2 Model Differences: 2B vs 9B Behavior

[HPP] Neel NandaApril 20, 20252h 34min
26 connections·40 entities in this video

Research Objective & Methodology

  • 🎯 The primary goal is to understand behavioral differences between Gemma 2 2B and 9B language models.
  • 🔬 The approach involves calculating per-token KL divergence between the two models on the Pile 10K dataset.
  • 🗣️ The session also explores "vibe coding," using voice and cursor for coding, and narrating the thought process.

Hypotheses on Model Differences

  • 💡 Investigating whether 9B's superior performance stems from sparse, interpretable structures (e.g., specific facts/circuits) or diffuse refinement.
  • 🧠 Considering if larger models primarily offer increased precision on existing knowledge or introduce entirely new, albeit sometimes noisy, facts.
  • ⚠️ Noted instances where the small model sometimes performs better, potentially due to overfitting or random noise.

Exploratory Data Analysis

  • 📊 Visualizing KL divergence and log probability differences using 2D histograms and scatter plots to understand the distribution of differences.
  • 🔍 Focusing on the exploration stage of research, aiming to gain "surface area" by examining qualitative examples and identifying weird phenomena.
  • ✅ Filtering for tokens where the big model is highly confident (rank 1) and the small model is significantly confused (rank worse than 50).

Key Findings from Specific Examples

  • 📚 Identified cases of factual recall, where the 9B model knew niche information (e.g., "Gabra Wad," "Gordon Capes") that the 2B model did not.
  • 🗣️ Demonstrated the big model's superior ability in language recognition and completion, specifically with Norwegian sentences.
  • 💻 Observed the 9B model's knowledge of standard technical commands and outputs (e.g., Docker system info), suggesting memorization of common syntax.

Interpreting Performance Gains

  • 📈 The performance difference is a mix of small and significant changes, with a decent amount driven by "big things" (large log probability differences).
  • 🧠 The "Norwegian" example suggests a generalizable capability (language inference) improved in the larger model.
  • 🧩 Speculated that open-ended tokens (e.g., after full stops) are where larger models might show the most significant benefits.
Knowledge graph40 entities · 26 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover · drag to explore
40 entities
Chapters20 moments

Key Moments

Transcript442 segments

Full Transcript

Topics13 themes

What’s Discussed

Gemma 2Language ModelsKL DivergenceModel BehaviorVibe CodingExploratory ResearchFactual RecallLanguage RecognitionTechnical CommandsData AnalysisLog ProbabilityModel PerformanceLarge Language Models
Smart Objects40 · 26 links
People· 2
Products· 11
Concepts· 22
Companies· 2
Event· 1
Medias· 2