Skip to main content

Building Multilingual AI Applications with Gemma 3

Google for DevelopersApril 2, 20259 min1,205 views
17 connections·23 entities in this video→

Enhancing Multilingual Capabilities in Gemma 3

  • πŸ’‘ Gemma 3 has significantly improved its multilingual performance, supporting 140 languages in pre-training and over 35 languages out-of-the-box in instruction tuning.
  • 🎯 This enhancement is crucial for developers aiming to reach global audiences, as 80% of Gemma users and over 64% of users come from non-English speaking countries.
  • πŸš€ The model's capabilities are comparable to GPT-4.0 in internal evaluations for instruction-tuned performance across several languages.

Training Strategies for Multilinguality

  • πŸ”‘ Gemma 3 utilizes a new tokenizer from Gemini that prioritizes non-English text, forming an excellent basis for multilingual applications.
  • πŸ“Š During pre-training, data quality was carefully curated both within and across languages, and the multilingual share of content was doubled.
  • πŸ› οΈ Post-training involved creating a curated set of instruction-following prompts and datasets using human and synthetic methods, ensuring coverage of diverse use cases, languages, and cultural contexts.
  • βœ… A key emphasis was placed on not compromising the quality of English language capabilities or other functions like reasoning and coding.

Real-World Applications and Examples

  • πŸ’¬ Gemma 3 can power diverse global applications, such as chatbot support platforms for financial institutions or healthcare providers needing to support hundreds of languages.
  • πŸ’‘ A practical example demonstrated Gemma's ability to identify languages in a sign in Berlin, translate it, and provide historical context from the Cold War.
  • 🌍 The model's multilingual capabilities, combined with multimodal understanding, long context, and reasoning, highlight its potential for global use cases.

Preserving Cultural Context with Specialized Models

  • 🎭 The Seine model, a collaboration between Google and AI Singapore, is based on Gemma 29b and tailored for Southeast Asian languages.
  • πŸ—£οΈ This model was continually pre-trained and instruction-tuned specifically for over a thousand native languages in Southeast Asia, outperforming similarly sized models on regional benchmarks.
  • 🀝 High-quality training data for Seine was curated by linguists and native speakers, with Project Aquarium set to make this data available to the public soon.

Tips for Developers Fine-Tuning Gemma 3

  • πŸ“ˆ Evaluation is critical; having a quality dataset, even a translated one, is better than none for fine-tuning.
  • ⚑ Gemma's strong tokenizer and pre-training can lead to high gains in post-training, sometimes outperforming models trained monolingually from scratch.
  • 🧠 Even a small amount of instruction tuning, as few as 40 samples, can significantly improve a model's instruction-following ability, even in languages it hasn't seen before.
Knowledge graph23 entities Β· 17 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover Β· drag to explore
23 entities
Chapters4 moments

Key Moments

Transcript32 segments

Full Transcript

Topics12 themes

What’s Discussed

Gemma 3Multilingual AINatural Language ProcessingMachine TranslationCross-lingual TransferInstruction TuningTokenizerData CurationSoutheast Asian LanguagesAI SingaporeProject AquariumGoogle AI for Developers
Smart Objects23 Β· 17 links
ProductsΒ· 6
CompaniesΒ· 4
PersonΒ· 1
ConceptsΒ· 10
MediaΒ· 1
LocationΒ· 1