Course overview

Build vs. buy decision factors: why Meta, OpenAI, and Anthropic build internal tools

Why annotation infrastructure suddenly became a strategic decision

Here's the thing, when OpenAI and Google both exited Scale AI as customers in mid-2025, it wasn't because the platform got worse. It was because Meta acquired 49% of Scale for $14B, and suddenly sharing your proprietary training data with a competitor-controlled vendor created unacceptable risk. From what I've seen, this moment crystallized something many AI leaders had been quietly debating: is annotation infrastructure a commodity you rent, or a strategic asset you control?

The build vs. buy decision isn't purely technical. Data sensitivity, competitive dynamics, and strategic importance to core AI competency drive the choice as much as engineering capacity. Dario Amodei at Anthropic contracts with external providers like Surge AI for RLHF work, but maintains tight control over annotation workflows and model training pipelines. Meanwhile, Meta built internal tools like Hatch, sandboxed web environments where their shopping agent trains on simulations of real websites (DoorDash, Etsy, Reddit), because the annotation process is inseparable from the product itself.

The pattern I keep seeing: frontier labs buy infrastructure for everything that isn't a competitive wedge (OpenAI uses Salesforce CRM, Linear, Greenhouse for hiring), but build when the tooling touches proprietary model training data or encodes unique product logic. The question is always the same, does sharing this with a vendor create leverage for a competitor?

Check your understanding

After Meta acquired 49% of Scale AI in 2025, both Google and OpenAI stopped using Scale for annotation. What does this tell you as a PM evaluating a new annotation vendor?

The four factors that tip the decision toward 'build'

From what I've observed, data sensitivity is the first trigger. If your annotation work exposes proprietary model training data, competitive product strategy, or customer data that could leak strategic intent, you're already halfway to a build decision. Meta's Hatch project trained shopping agents on simulated transactions, the kind of data you can't share with a vendor who might work with Amazon or Shopify tomorrow.

The second factor is workflow control requirements. If your annotation process is tightly coupled to your model training loop, think active learning pipelines where model predictions continuously inform which examples get labeled next, then off-the-shelf platforms add friction. You need the ability to iterate on labeling schemas, retrain models mid-project, and adjust task instructions without waiting for a vendor's feature roadmap.

Third: technical capacity to maintain tooling. Building isn't just writing code once, it's ongoing maintenance, UI/UX iteration, workforce management, and quality monitoring. Anthropic has the engineering bandwidth; a 20-person AI startup probably doesn't. Edwin Chen founded Surge AI precisely because most teams don't want to build this infrastructure. The hidden cost is always in the long tail of edge cases and workforce operations.

Finally, strategic importance to core AI competency. If annotation tooling is a competitive wedge (like Scale AI's Thunderforge contract for DoD military planning AI), you build. If it's table stakes, you buy. Alexander Ratner's work on Snorkel AI suggests a middle path: programmatic weak supervision lets technical teams encode domain knowledge as labeling functions, reducing dependence on large-scale human annotation without building a full platform.

The hidden build cost

Most teams underestimate the operational complexity of running an annotation workforce. Quality monitoring, annotator onboarding, task instruction iteration, and payment infrastructure are not one-time engineering problems, they're ongoing ops work. Surge AI scaled to $1B+ revenue bootstrapped in part because most AI labs would rather pay a premium than build this infrastructure themselves.

Check your understanding

What the vendor landscape tells you about when to buy

The market has segmented pretty cleanly around different build/buy personas. Scale AI's full-stack managed service (platform + 240,000+ annotators) appeals to teams that want to treat annotation as an API call, you send data, you get labels back, you don't manage people. That model worked brilliantly until competitive dynamics made vendor neutrality a concern. After Meta's acquisition, Surge AI scaled past $1B revenue by positioning as the independent, quality-first alternative.

Labelbox took the opposite bet: platform-only, bring-your-own-labor. They provide SOC2-certified enterprise workflow tooling, API flexibility, VPC/on-prem deployment, but customers must source and manage their own annotators. This appeals to teams with existing internal annotation operations or those who want full control over data security and workforce management. It's a 'build the workforce, buy the tooling' model.

Then there's Snorkel AI's programmatic labeling wedge, encode domain rules as labeling functions (code-based heuristics, rules, model signals) and generate training labels with minimal human annotation. Alexander Ratner's research on weak supervision resonates with technical teams working on text/tabular data (finance, healthcare, insurance). Five of the top 10 US banks use Snorkel because classification schemas change constantly, and re-annotating millions of records manually isn't viable. It's a 'reduce the need to build or buy annotation at scale' play.

For me, the tell is this: if annotation quality is a bottleneck to model performance and you're scaling frontier models, you lean toward curated expert networks like Surge AI or Mercor (which uses AI-driven recruiting to match domain experts to tasks). If annotation is table stakes and you need cheap throughput, you lean toward managed platforms. If you're encoding complex domain logic that changes fast, programmatic labeling makes sense. And if you're training on data that reveals competitive strategy, you build.

Check your understanding

You're a PM at a mid-sized AI startup building a legal document analysis product. Your ML team needs 50,000 contract clauses annotated for entity extraction and clause classification. You have two ML engineers and no existing annotation workforce. Should you build, buy, or use programmatic labeling? Walk through your reasoning.

Scale AI's Thunderforge contract

Scale won prime contractor role for the DoD's flagship AI program for military planning and operations, then expanded the contract ceiling to $500M in 2025. Scale positioned specifically around the data-quality bottleneck in operational military AI deployment, hired former White House CTO Michael Kratsios as managing director, and leveraged Meta's investment to expand public-sector business at speed. This illustrates annotation infrastructure becoming a strategic build, the DoD couldn't risk vendor-hopping for mission-critical AI systems.

Key takeaways

  • The build vs. buy decision is driven as much by competitive dynamics and data sovereignty as by technical capability.
  • Vendor neutrality became a purchasing criterion after Meta acquired 49% of Scale AI, sharing proprietary training data with a competitor-controlled platform creates unacceptable risk.
  • Build when annotation touches proprietary model training data, encodes unique product logic, or is a competitive wedge; buy for table-stakes throughput.
  • Even frontier AI labs buy most infrastructure (CRM, project management, hiring tools) and only build what creates competitive advantage.
  • The hidden cost of building is ongoing ops work: quality monitoring, workforce management, and tooling maintenance, not just the initial engineering effort.

Your product check-in

Apply “Build vs. buy decision factors: why Meta, OpenAI, and Anthropic build internal tools” to a product or workflow you know. What would you try, what could go wrong, and what evidence would help you decide?

Ask AI
AI Learning Assistant