Building and Scaling AI Startups: Lessons from Casetext's $650M Exit in Legal AI
Y Combinator
Summary:
This talk outlines key strategies for AI startup success, drawing from Casetext's journey to a $650 million acquisition.
- Idea Selection: Identify problems where people pay others to perform tasks, categorizing AI solutions into assisting, replacing, or enabling previously unthinkable tasks. This significantly expands the total addressable market by focusing on the combined salaries currently spent on those tasks.
- Reliable Product Development: Focus on deeply understanding professional workflows, breaking tasks into specific steps, and implementing these steps using a combination of prompts and traditional software engineering.
- Rigorous Evaluation (Evals): Crucially, prioritize building reliable AI products, not just flashy demos. Develop objective evaluation frameworks, relentlessly test prompts, and iterate based on real-world customer feedback to achieve high accuracy (e.g., 97%+) and robustness.
- Marketing and Sales: Emphasize that superior product quality drives organic growth (word-of-mouth, news) more effectively than aggressive marketing. Price based on value, listen to customer preferences, and actively build trust through comparisons, pilots, and dedicated customer success efforts.
- Founder Focus: Maintain an unwavering focus on achieving product-market fit at every stage of the company's growth, as all other business functions should serve this primary goal.
How We Built a $650M AI Company [00:00]
Jake Heller, co-founder and CEO of Casetext, shares insights into building an AI company that led to a $650 million acquisition by Thomson Reuters.
- The journey involved three core ideas:
- Picking the right ideas: Deciding which problems to pursue.
- Actually building it: Developing the AI product effectively.
- Marketing and selling it: Successfully bringing the product to market.
- Casetext's Background:
- Founded in 2013, Casetext focused on applying AI (initially NLP/Machine Learning) to law.
- Early access to GPT-4 in summer 2022 prompted a complete pivot to build CoCounsel, an AI assistant for lawyers.
- CoCounsel, recognized as the first and best AI assistant for lawyers, led to the acquisition for $650 million.
- Future Market Potential: The speaker believes future AI companies will achieve even greater valuations, unlocking immense potential for innovation.
Picking the Right Idea in the AI Era [01:00]
The challenge of "making something people want" is simplified in the AI era by observing what people currently pay others to do.
- Identifying Market Needs: Look at tasks people are paying for, such as customer support, insurance adjusters, paralegals, personal trainers, or executive assistants. These represent existing desires and value.
- AI's Role in Problem Solving:
- Large Language Models (LLMs) can solve many current human-performed tasks.
- Robotics can address physical world tasks.
- Three Categories of AI Startups:
- Assist: Help professionals accomplish tasks (e.g., CoCounsel assisting lawyers with document review and research).
- Replace: Completely take over tasks (e.g., an AI-powered law firm or accountant).
- Do the Unthinkable: Enable tasks previously impossible or prohibitively expensive (e.g., AI reviewing hundreds of millions of documents, a task human lawyers would never attempt due to cost).
Three Types of AI Startups: Assist, Replace, or Do the Unthinkable [04:45]
- Total Addressable Market (TAM) Expansion:
- Traditionally, TAM was calculated by the number of professionals (seats) multiplied by a monthly fee (e.g., $20/month SaaS).
- With AI, the TAM is now the combined salaries of all people currently performing the job, a number potentially 1,000x larger. Instead of $20/month, solutions can command $5,000-$20,000/month if they replace a professional's salary.
- Societal Impact of AI:
- Unlocking new futures: AI frees humans from antiquated tasks, similar to how electricity replaced lamp lighters, enabling new, unimaginable possibilities.
- Democratizing access: AI makes expensive or hard-to-access services (like legal, financial, or personal assistance) affordable for a wider population. For instance, addressing the 85% of low-income individuals lacking legal services due to cost.
How to Build Reliable AI Products (Not Just Demos) [09:25]
Building reliable AI requires a deep understanding of the problem space and methodical execution, contrasting with many flashy but unreliable AI demos.
- 1. Understand Professional Workflows:
- Deep domain expertise: Founders must intimately know "what people actually do" in the target profession. This can be achieved through personal experience (like Jake as a lawyer), embedded observation, or partnering with domain experts.
- Specificity is key: Break down tasks into minute, specific steps (e.g., for legal research: understand the request, ask clarifying questions, make a research plan, execute searches, read results, filter irrelevance, make notes, synthesize an essay, verify citations).
- 2. Translate Steps to Code/Prompts:
- Prompts for human-level intelligence: Many steps requiring human-like intelligence become AI prompts (e.g., "rate relevance 0-7," "write essay from notes," "verify citation accuracy").
- Traditional software for deterministic tasks: Use conventional software engineering for deterministic or mathematical calculations where possible, as prompts are slower and more expensive.
- 3. Workflow vs. Agentic Systems:
- Workflows for deterministic tasks: If a task consistently follows the same steps, build a simple, sequential Python workflow.
- Agentic systems for dynamic tasks: If the approach varies by circumstance, a more agentic system is needed, which is harder to make reliable.
The Importance of Evals and Testing [16:30]
Reliability is paramount for AI products to move beyond demos and be successful in practice.
- The "Reliability Gap": LLMs, like humans, can be inconsistent. Demos often show 60-70% accuracy, which is insufficient for real-world applications.
- Defining "Good" with Domain Expertise:
- Establish clear criteria for what constitutes a "good" or "correct" output for both the overall task and each microtask.
- Involve actual professionals to define these benchmarks.
- Creating Objective Evaluations:
- Design evals that yield objectively gradable answers (e.g., true/false, a numerical score from 0-7).
- Use evaluation frameworks (like Promptfoo) to automate testing.
- Start with a small set of tests (e.g., a dozen), perfect them, then scale to 50, then 100, continuously tweaking prompts.
- Utilize a hold-out set to prevent overfitting prompts to the training evals.
- Iterative Prompt Engineering:
- AI often fails predictably (ambiguous instructions, consistent biases). Prompts should be refined with direct instructions and examples to guide the AI away from errors.
- This "grind" involves spending weeks sleeplessly refining a single prompt to achieve high accuracy (e.g., 97% pass rate, where remaining 3% are human-like judgment calls).
- For pre-production/beta, aim for 100 tests per prompt and per overall task, targeting 99% accuracy.
- Continuous Improvement:
- Customer complaints are invaluable for creating new tests and refining the product. Customers will use the app in unexpected, "dumb" ways, revealing gaps in testing.
- Regularly test new models as they emerge and continuously iterate on prompts. Even a 1% increase in accuracy can be crucial in fields like finance, medicine, or law.
Why Product Quality Beats Marketing and Hype [24:20]
- Product as Primary Marketing:
- Contrary to some VC advice, exceptional product quality is the most powerful marketing tool.
- A truly "awesome product" generates word-of-mouth referrals and attracts news coverage, which are forms of free marketing.
- Sales teams transition from "sellers" to "order takers" when the product's value is self-evident.
- Avoid the "Demo Trap": Many startups focus on flashy demos for fundraising but fail to build reliable products that work in practice, leading to pilot program failures and a "mass extinction event" for companies with unproven revenue.
How to Price and Sell AI Products [26:00]
- 1. Sell Service/Employment, Price Accordingly:
- Shift from selling traditional software to selling a service or "employment" (e.g., an AI lawyer).
- Price based on the value delivered, not traditional SaaS models. Example: $500 per contract review (compared to $1,000 by a human lawyer) is a significant price step-up from typical $20/month SaaS tools.
- 2. Listen to Customer Payment Preferences:
- Ask customers how they prefer to pay. Casetext found customers preferred predictable budgeting ($6,000 per seat/year) over per-usage pricing, even if it meant paying more overall.
- 3. Build Trust with Customers:
- AI is new and can be scary for businesses. Build trust through:
- Head-to-head comparisons: Allow customers to use your AI alongside their existing human service providers and compare results (speed, quality, differences).
- Studies and pilots: Conduct structured tests to demonstrate effectiveness.
- Conscious onboarding and training: Actively train and roll out the product, ensuring users understand and adopt it.
- "Deployed engineers": Hire people to sit with customers ("boots on the ground") to ensure the product works for them and address issues.
- 4. Customer Adoption and Retention:
- The sale doesn't end with a check or a pilot. Ensure everyone uses and understands the product. High pilot conversion to recurring revenue is critical to avoid "pilot recurring revenue" (PRR) turning into "mass extinction."
Product Isn’t Just Pixels, It’s Everything Around It [29:30]
- Holistic Product Experience: The product encompasses more than just the user interface or feature set.
- It includes human interactions with support, customer success teams, and even founders.
- Training, onboarding, and overall customer experience are integral to the product.
- Companies that invest in these surrounding aspects will outperform those focused solely on "pixels."
What Founders Should Really Focus On [33:00]
- Unwavering Focus on Product-Market Fit: At every stage (seed, Series A, B, C), the primary focus should be on building a great product that achieves product-market fit.
- Supporting Functions as Means, Not Ends: Other functions like HR, finance, and fundraising should be viewed as means to support product development and market fit, not as goals in themselves. Founders often fall into the trap of focusing on these in the abstract, rather than their instrumental role.
- CEO's Role: As CEO, your focus on product-market fit will naturally lead to addressing HR (who to hire for the product), marketing/sales (how to tell the world about the product), and culture (what environment fosters product love).
Q&A: Picking Markets, Focus, and Defensibility [36:00]
- Picking Markets and Competitors:
- Don't be scared of competitors. The market for AI-driven automation is so vast (trillions of dollars currently spent on human tasks) that multiple companies can succeed. Many existing competitors are often found to be "bad" and easily outbuilt.
- Target roles that are already being outsourced to other countries, as this indicates a willingness to externalize the work.
- Avoid roles where tasks are deeply tied to a company's core identity (e.g., Pixar's storytelling), as these are less likely to be outsourced to AI.
- Focus on large, painful problems that AI can solve, leveraging existing knowledge or easily accessible information.
- Defensibility in AI (Beyond GPT Wrappers):
- The defensibility comes from the sheer difficulty and complexity of building a truly reliable AI product.
- This includes the numerous small pieces, data integrations, rigorous checks, fine-tuned prompts, and careful model selection required.
- The effort (e.g., two years of dedicated work) creates a moat that others cannot easily replicate, transforming a simple "GPT wrapper" idea into a robust, defensible solution.