Skip to main content

State of Reasoning Models, Building LLMs from Scratch and 7 Years of Scaling GPT | Sebastian Raschka

[HPP] Sebastian RaschkaMarch 27, 202556 min
26 connections·40 entities in this video

Learning AI Effectively

  • 🎯 Stay focused amidst abundant resources and avoid constant switching between learning materials.
  • 🌱 Pick a project that excites you to apply learned concepts and maintain motivation, balancing theory with practical application.
  • 🧠 Understand that mathematics and statistics are foundational but should be integrated with practical work, not as a dry starting point.

Navigating Scientific Literature

  • 🔍 Approach papers with a specific goal or question in mind, rather than general reading, to make note-taking more effective.
  • 💡 Use tools like e-readers to minimize distractions and focus on the content.
  • 📚 For general learning, textbooks are often more effective than papers, which are better for project-specific updates and new research.

Understanding Reasoning Models

  • 🚀 Reasoning models represent the next stage of LLMs, excelling at complex tasks like code and math by providing step-by-step solutions.
  • 💡 They act as a "study partner," explaining "why" and "how," which can enhance user learning, unlike regular LLMs that might foster laziness.
  • ✅ Modern LLMs are integrating the ability to toggle reasoning capabilities on and off, making them more versatile for different query types.

Evolution of Model Architectures

  • ⚙️ The Transformer architecture continues to be refined with innovations like group query attention and sliding window attention for efficiency and longer contexts.
  • 🧠 Alternatives like Mamba (state space models) aim to reduce the quadratic cost of attention and handle longer sequences, though Transformers remain dominant in large-scale applications.
  • 📈 There's a perceived saturation in Transformer scaling for pre-training, suggesting a need for new architectural paradigms or increased focus on post-training methods.

Multi-GPU Training & Open Source

  • ⚠️ Large-scale multi-GPU training is complex, expensive, and requires significant resources, often a team effort, making it impractical for individuals.
  • 💡 Post-training approaches offer significant opportunities for improving models, especially in specialized domains, without the immense cost of pre-training from scratch.
  • 🤝 Open-source contributions create a synergy, benefiting both companies (by building useful tools like PyTorch Lightning) and the broader community through collaboration and shared learning.
Knowledge graph40 entities · 26 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover · drag to explore
40 entities
Chapters20 moments

Key Moments

Transcript208 segments

Full Transcript

Topics15 themes

What’s Discussed

Large Language Models (LLMs)Reasoning ModelsTransformer ArchitectureState Space Models (Mamba)Multi-GPU TrainingOpen-Source SoftwarePre-trainingPost-trainingReinforcement LearningSupervised Fine-tuningAgent-based SystemsLong ContextData QualityLow-Rank Adaptation (LoRA)Compute Resources
Smart Objects40 · 26 links
People· 2
Concepts· 23
Companies· 3
Products· 5
Medias· 7