Skip to main content

AI Agents, MCP, and the Problems with AI Benchmarks with Matt Carey

[HPP] Simon WillisonApril 19, 202547 min
31 connections·40 entities in this video→

Understanding MCP and AI Agents

  • πŸ’‘ Multi-Client Protocol (MCP) is presented as a universal plug-in system for AI clients, akin to Chrome extensions, allowing users to extend functionality in AI-powered applications like Cursor, Notion, and Canva.
  • πŸ› οΈ MCP defines a standardized protocol for API specifications, primarily supporting "tools" (JavaScript/Python snippets) that enable AI clients to interact with external systems like databases or 3D modeling software.
  • πŸ”„ Unlike proprietary client-side tool calling, MCP offers a standardized integration method for services, allowing any AI-powered app to connect using the same protocol, with remote MCP servers becoming increasingly common.
  • 🧠 The core philosophy of AI agents involves a loop where the model continuously decides what to do next, including calling tools, rather than a single-shot prompt, giving the AI agency over the workflow.
  • βš–οΈ Google's Agent-to-Agent approach differs from MCP by focusing on specialized sub-agents for specific tasks orchestrated by a main agent, contrasting with MCP's emphasis on user control and client-dictated inference.

StackOne's Role in AI Integrations

  • 🏒 StackOne specializes in B2B integrations, providing a unified API for connecting various enterprise applications like HR and applicant tracking systems.
  • πŸš€ StackOne's data unification capabilities are proving highly beneficial for AI agents, helping to manage context windows and clean data for AI applications.
  • 🎯 Matt Carey's current work at StackOne focuses on evolving their API into a toolset for B2B agents and MCP, enabling more sophisticated AI-driven integrations.

Building and Evaluating AI Systems

  • βœ… When developing AI applications, it's crucial to start simple with direct LLM calls and only transition to an agentic system when the problem's complexity or unbounded nature necessitates giving the model more control.
  • πŸ“Š Robust evaluation is paramount, requiring 20 to 100+ "gold standard" inputs and outputs to accurately test and compare model performance.
  • 🀝 Involving domain experts is essential for creating meaningful evaluations and ensuring that AI outputs align with real-world requirements and quality standards.

The Problem with AI Benchmarks

  • ⚠️ Many public AI benchmarks are saturated and often do not accurately reflect real-world use cases, frequently serving more as marketing tools than reliable performance indicators.
  • πŸ“‰ Claims of massive context windows (e.g., 1 million tokens) are often misleading, as models typically struggle with effective reasoning beyond 10,000-20,000 tokens, with Gemini 2.5 Pro being a notable exception.
  • πŸ”¬ To obtain reliable performance metrics, it's recommended to create proprietary benchmarks using internal, never-before-seen data (e.g., company reports, internal documents) to test a model's reasoning capabilities.

Navigating the AI Information Landscape

  • πŸ” Filtering noise from signal in the rapidly evolving AI space is a significant challenge, with much content being hype-driven marketing.
  • πŸ“š Recommended sources for reliable AI information include curated social media feeds (e.g., Simon Willison on X), long-form blogs, podcasts, research papers, and in-person events.
  • πŸ§‘β€πŸ’» Leveraging other people's curation, such as trusted newsletters or event organizers, can help individuals find high-quality, relevant insights without having to sift through excessive information themselves.
Knowledge graph40 entities Β· 31 connections

How they connect

An interactive map of every person, idea, and reference from this conversation. Hover to trace connections, click to explore.

Hover Β· drag to explore
40 entities
Chapters20 moments

Key Moments

Transcript176 segments

Full Transcript

Topics12 themes

What’s Discussed

AI AgentsMulti-Client Protocol (MCP)AI BenchmarksLLM Context WindowsTool CallingAI System EvaluationB2B IntegrationsAgentic SystemsOAuth AuthenticationProprietary Data BenchmarkingSignal-to-Noise RatioLarge Language Models (LLMs)
Smart Objects40 Β· 31 links
ConceptsΒ· 13
PeopleΒ· 3
ProductsΒ· 15
CompaniesΒ· 9