Skip to main content

Unlocking Scalable Robot Learning in the Real World – Karl Pertsch (UC Berkeley and Stanford)

[HPP] Sergey LevineApril 2, 202558 min
30 connections·40 entities in this video→

Challenges in Robot Learning

  • πŸ’‘ While machine learning has advanced in areas like language and vision, its impact on the physical world and robotics is limited due to a lack of real-world interaction.
  • πŸ€– Traditional robots operate in structured environments with pre-programmed behaviors, which is not viable for open-world manipulation.
  • ⚠️ Existing robot learning methods, often based on teleoperation, produce brittle and task-specific models that require extensive and repeated data collection, hindering generalization.

Building Scalable Robot Learning Pipelines

  • 🎯 The goal is to combine generalist machine learning models with dexterous robot control by developing diverse robot learning datasets and high-capacity model architectures.
  • πŸ”‘ Key elements for scalable ML in robotics include data, models, and evaluations, each presenting unique challenges compared to other ML domains.
  • 🚧 Robotics faces difficulties in scraping large datasets, adapting model architectures for continuous control, and conducting scalable, reproducible evaluations.

Advancing Robot Data Collection

  • 🀝 Community collaborations like DROID and Open X-Embodiment were crucial for creating large, diverse robot learning datasets.
  • πŸ“Š The Open X-Embodiment project aggregated over 1 million real robot episodes spanning 22 different robot embodiments, fostering data reuse within the community.
  • 🌱 This shift from bespoke, task-specific datasets to general-purpose, reusable datasets is essential for supporting large-scale robot learning research.

Developing Generalist Robot Policies

  • 🧠 The research transitioned from skill-based learning to Vision-Language Action (VLA) models to create flexible and scalable policy architectures.
  • πŸ“‰ Initial VLA models struggled with dexterous control due to low-information density in action tokens from naive binning tokenization, leading to slow and imprecise performance.
  • πŸš€ The introduction of the FAST (DCT-based compression) tokenizer significantly improved training efficiency and enabled effective learning for high-frequency, complex tasks like laundry folding.
  • βœ… The pi0 FAST model demonstrated zero-shot generalization by allowing deployment on new robot platforms and environments, responding to natural language prompts.

Scalable Evaluation for Robotics

  • πŸ”¬ Addressing the lack of scalable and reproducible evaluations is critical for advancing robotics research.
  • πŸ§ͺ Developed simulated evaluation platforms that are realistic enough for policies trained on real robot data to provide indicative success rates.
  • πŸ‘οΈ Found that visual realism is a key factor for the reliability of simulated evaluations, more so than accurate physics for simpler manipulation tasks.

Future Directions and Open Challenges

  • 🧩 Future work aims to improve generalization to novel objects by unifying embodied and non-embodied foundation models.
  • πŸ’ͺ Enhancing dexterity for complex manipulation tasks is a major focus, potentially by combining scalable learning with skill-based learning or human video data.
  • πŸ’¬ Making policies easier for humans to teach and interact with through more natural interfaces, such as natural language feedback, is an important area for unlocking broader utility.
Knowledge graph40 entities Β· 30 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover Β· drag to explore
40 entities
Chapters15 moments

Key Moments

Transcript214 segments

Full Transcript

Topics15 themes

What’s Discussed

Robot LearningGeneralist Robot PoliciesMachine Learning PipelinesVision-Language Models (VLMs)Continuous ControlRobot Data SetsOpen X-EmbodimentImitation LearningAction TokenizationDiscrete Cosine Transform (DCT)Dexterous ControlSimulated EvaluationVisual RealismNatural Language InstructionFoundation Models
Smart Objects40 Β· 30 links
PeopleΒ· 3
MediasΒ· 6
ConceptsΒ· 24
CompaniesΒ· 3
ProductsΒ· 3
EventΒ· 1