Brendan Foody on Mercor's AI Training: Hiring Experts for Economic Value and the Future of Knowledge Work
Conversations with Tyler
Summary:
Brendan Foody, CEO of Mercor, discusses how his company hires experts to train AI models, offering poets $150/hour to create rubrics for poetry evaluation. Key takeaways from the conversation include:
- Mercor addresses the gap between academic AI evaluations and real-world economic value by hiring top experts in various fields.
- AI models are showing a rapid 25-30% annual improvement in economically valuable tasks, with GPT-5 scoring 64%.
- The future of knowledge work will involve professionals building reinforcement learning (RL) environments and training AI agents, akin to software development.
- "Vibes-based" hiring is inefficient; Mercor emphasizes skill-based project assessments and leverages AI for better talent matching.
- The Thiel Fellowship could scale massively by using AI for interviewing and candidate selection, promoting economic mobility.
- Dyslexia is linked to entrepreneurship, potentially fostering unconventional thinking and early delegation skills.
- Mercor's next goal is to scale super-realistic evaluations measuring AI's ability to use tools for long-horizon enterprise tasks.
Mercor's Approach to AI Training: Expert Evaluation [00:00:00]
Brendan Foody, CEO and co-founder of Mercor, discusses his company's unique approach to training advanced AI models.
- Hiring Experts for AI Evaluation [00:00:00]
- Mercor employs domain experts, including poets, to teach leading AI models.
- Poets create rubrics and examples to measure and guide AI performance in tasks like poetry generation, explaining the high pay ($150/hour) by the broad application of learned skills.
- The process aims for a balance of consensus and diverse perspectives among expert graders, even using AI to detect human grading inconsistencies.
- The AI Productivity Index (APEX) [00:03:53]
- Mercor developed APEX to bridge the gap between academic AI benchmarks and real-world economic value.
- They collaborated with renowned experts like Larry Summers (finance/economics), Cass Sunstein (law), and Eric Topol (medicine) to define evaluation methodologies focused on economically valuable tasks.
- This involves surveying experts on how they spend their time, then creating prompts and rubrics based on those tasks (e.g., medical diagnosis, legal drafting, financial analysis).
Measuring AI Progress and Capabilities [00:05:29]
- Rapid Model Improvement [00:05:29]
- AI models demonstrate significant improvement (25-30% annually) in economically valuable tasks.
- For example, GPT-5 scores around 64% on these real-world evaluations.
- This rapid progress suggests a profound impact on the economy in the near future.
- Current AI Limitations [00:07:51]
- Models still struggle with:
- Long-horizon tasks: Activities requiring 50-100 hours of continuous work.
- Tool integration and human interaction: Effectively using multiple tools or interacting with people.
- Foody anticipates significant advancements in these areas within the next 6-12 months as evaluation methods improve.
- AI Stumping Human Experts [00:09:31]
- It's predicted that within 2-3 years, models will make it very difficult for experts like Cass Sunstein to find errors in their responses, especially in less taste-driven domains.
- Human experts excel in areas with uncodified "taste" or niche, undocumented knowledge, where models lack sufficient pre-training or post-training data.
- Value of Data for AI Advancement [00:13:42]
- The most valuable data for AI training comes from detailed rubrics, test questions with answers, and unit tests for code.
- These "measurement of success" data allow models to attempt problems repeatedly, score responses, and learn efficiently.
- For creative fields like poetry, defining good taste for AI is complex; relying on power users' preferences versus top experts' opinions is a strategic decision for AI labs.
- In the long run, AI could learn and personalize taste across different historical eras for individual users.
The Future of Knowledge Work [00:22:38]
- Society as an RL Machine [00:22:38]
- Foody believes a significant portion of the economy will transform into a reinforcement learning (RL) environment.
- Knowledge workers will transition from performing repetitive tasks to building RL environments and training AI agents to perform those tasks once, then leveraging the agents repeatedly (similar to software development).
- This shift is driven by economic incentives where initial fixed-cost investments in AI training yield scalable, reusable solutions.
- Impact on Employment and Education [00:30:46]
- A new job category will emerge: people dedicated to training AI agents and building RL environments.
- These roles will primarily require domain expertise to identify model mistakes, rather than deep technical AI knowledge.
- The price elasticity of demand for different professions will dictate job displacement or growth; software engineering, for example, is highly elastic, potentially leading to more engineers creating more software.
- Education is poised for transformation with personalized AI tutors (e.g., "Sal Khan for everyone"), providing 24/7 access to information and tailored explanations. Human teachers will likely shift focus to personal relationships and emotional development in smaller class settings.
Optimizing Hiring and Labor Markets [00:35:05]
- Mercor's Hiring Philosophy [00:35:05]
- Mercor utilizes its own technology to automate resume review, conduct interviews, and make hiring decisions, focusing on objective skill assessment.
- Traditional "vibes-based" interviewing, which over-indexes on personal rapport rather than job-specific skills, is deemed inefficient.
- Future hiring will involve project-based assessments, gathering extensive contextual data (e.g., manager notes, interview recordings) for AI analysis to predict performance.
- Improving Labor Market Efficiency [00:39:55]
- Current labor markets are inefficient due to disaggregation and a difficult matching problem (e.g., LinkedIn having distribution but lacking deep performance insight).
- AI can facilitate a more efficient, aggregated, and globalized matching system, especially as work becomes more fractional and remote.
- While AI might optimize cover letters and interview performances, potentially increasing reliance on nepotism in some industries, data-driven AI could also act as a counter-signal by accurately predicting performance.
- AI-run labor markets could offer more second chances by identifying suitable roles for individuals where they can excel.
- AI and Assessment Tools [00:43:38]
- Instead of prohibiting AI cheating tools in assessments, the approach should shift to evaluating what individuals can achieve while using these tools, as this reflects real-world productivity. Mercor has developed methods to discern genuine skill even with AI assistance.
Entrepreneurial Journey and Personal Insights [00:45:01]
- Scaling the Thiel Fellowship [00:45:01]
- As a Thiel Fellow dropout, Foody suggests the fellowship, which faces a candidate matching problem, could use AI-powered interviews and transcript analysis to scale its selection process.
- This would allow identifying unconventional talent globally and significantly increase economic mobility by providing opportunities to more aspiring entrepreneurs.
- Hypothetical Gap Year [00:48:11]
- Given a magical, consequence-free gap year, Foody would travel extensively, particularly to Japan, to gain diverse cultural and global perspectives on AI and life, drawing parallels to AI models learning different "tastes."
- Early Entrepreneurial Experience and Dyslexia [00:50:31]
- Foody recounts his 8th-grade "donut dynasty," where he arbitraged Safeway donuts to sell at school, learning early lessons in business, scaling, and competitive pricing.
- His dyslexia, he notes, is statistically correlated with entrepreneurship, possibly fostering unconventional thinking and an early understanding of the need to delegate tasks where others have comparative advantages.
- He emphasizes the importance of understanding and leveraging one's strengths.
- Culture and Dining [00:56:15]
- Foody feels less connected to the general culture of 22-year-olds due to his intense work schedule, but observes a "dating crisis" in San Francisco due to gender imbalance, viewing dating apps as efficient matching tools.
- He shares San Francisco dining recommendations (El Metate for Mexican, Cotogna, Quince for higher-end) and suggests using the "Beli" app for reliable food ratings.
Mercor's Vision and Future Learning [00:59:01]
- Company Name and Mission [00:59:01]
- Mercor (Latin for "marketplace") reflects the company's goal to build the world's largest marketplace for talent, facilitating the training and improvement of AI models.
- Next Goals for Mercor [01:00:12]
- The primary goal is to scale "super-realistic evaluations" that measure AI's ability to use various tools for long-duration, multi-day or multi-week tasks, especially for enterprise applications.
- The focus is on bridging the gap between raw AI intelligence and practical usefulness for businesses.
- Personal Learning Focus [01:00:54]
- Foody is most interested in further learning about advancements in AI research, particularly how human talent and labor can be most effectively applied to frontier AI problems, and which specific rubrics or data types drive the most significant model improvements.