Skip to main content

Cerebras SUPERNOVA 2025: Andrew Feldman on AI Speed and Reasoning Models

[HPP] Andrew FeldmanMay 16, 202517 min
27 connections·26 entities in this video→

Cerebras' Latest AI Innovations

  • πŸš€ Cerebras has launched the Qwen 32B model, which is the fastest in the industry, performing 20 to 30 times faster than GPU offerings.
  • 🧠 Qwen 32B is a reasoning model that behaves like leading closed-source models, offering smart, open-source capabilities.
  • πŸ’‘ The company is pushing the envelope on speed to meet market demand for faster AI processing.

The Critical Role of Speed in AI

  • ⏱️ Reasoning models require more computational work and more tokens as they iterate through developing, reviewing, and improving answers.
  • ⚑ Fast performance is crucial for agentic workflows, enabling complex tasks to be completed in seconds rather than minutes.
  • 🎯 While time to first token is important for voice applications, output speed is generally the best measure for most other applications.

Benchmarking AI Performance

  • πŸ“Š Artificial Analysis is highlighted as a reliable benchmarking firm that pings APIs daily to report real-world performance results.
  • πŸ” This approach provides a realistic measure of what customers experience, unlike historic benchmarks that could be tuned for months.
  • βœ… The goal is to make benchmarking as realistic as possible, mirroring actual customer performance rather than special, optimized versions.

Enabling Diverse AI Applications

  • πŸ₯ Large enterprises like GlaxoSmithKline are using Cerebras's technology for agentic flows in drug discovery, generating hypotheses, and linking to automated labs and robotics.
  • 🌱 A new class of AI entrepreneurs and startups are leveraging Cerebras's horsepower to develop innovative use cases.
  • πŸ€– Computer vision is identified as a significant consumer of inference compute, requiring substantial processing power.

Strategic Partnerships and Market Expansion

  • 🀝 Cerebras achieved a hyperscaler win with Meta, powering their Llama 4 API service for fast token delivery.
  • πŸ’Ό IBM also partnered with Cerebras to provide fast inference to their large enterprise clients, expanding reach from developers to major financial institutions.
  • 🌍 Partnerships with companies like Perplexity, Mistral, G42, and AlphaSense are delivering blazing fast models globally.

Empowering AI Developers

  • πŸ› οΈ Cerebras is offering one million free tokens per day to developers, with no waitlist, to encourage innovation.
  • πŸ’‘ The message to developers is to use inference to solve specific, real-world problems, moving beyond "AI for AI's sake."
  • πŸš€ This initiative aims to foster an ecosystem where developers can create niche, domain-specific SaaS applications powered by Cerebras's fast engine.
Knowledge graph26 entities Β· 27 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover Β· drag to explore
26 entities
Chapters2 moments

Key Moments

Transcript63 segments

Full Transcript

Topics15 themes

What’s Discussed

CerebrasAI PerformanceReasoning ModelsQwen 32BAgentic WorkflowsInference ComputeBenchmarkingAI DevelopersMemory BandwidthHyperscalersEnterprise AIDrug DiscoveryRoboticsComputer VisionSaaS Applications
Smart Objects26 Β· 27 links
CompaniesΒ· 11
ProductsΒ· 7
PeopleΒ· 3
ConceptsΒ· 2
MediasΒ· 3