Surge AI, Mercor, SuperAnnotate: expert-first, niche vertical, and self-serve wedges
Why the specialist providers won customers away from Scale
Here's the thing, after Meta acquired 49% of Scale AI in June 2025, Google and OpenAI both pulled their annotation work off the platform. Google's internal research teams reportedly said Scale's data quality was lower than what they got from Surge AI and Mercor. That wasn't a technical failure on Scale's part. It was a strategic wedge opening.
From what I've seen, the specialist providers, Surge AI, Mercor, SuperAnnotate, win by saying no to parts of the market Scale tried to own. Surge says no to commodity crowdsourcing and focuses exclusively on expert-tier RLHF and language annotation at $50-100+/hour. Mercor says no to generic labeling and matches domain experts (lawyers, analysts, engineers) to annotation tasks using AI-driven recruiting. SuperAnnotate says no to one-size-fits-all workflows and builds fully customizable annotation interfaces with a drag-and-drop workflow builder.
The pattern: when your wedge is narrow enough, you can deliver quality that makes the full-stack provider look like a compromise. And when a competitor acquires that full-stack provider, data sovereignty concerns suddenly make your neutrality a feature, not a bug.
Check your understanding
After Meta acquired 49% of Scale AI in 2025, Google and OpenAI shifted their annotation work to other providers. What was the PRIMARY reason these companies exited?
Surge AI's expert-first wedge: quality over volume
Edwin Chen founded Surge AI in 2020 after getting frustrated with crowdsourced annotation quality. His bet: frontier AI labs would pay premium rates for curated specialists rather than settle for commodity crowds. He was right. Surge scaled to ~1M annotators by 2025 and crossed $1B in revenue, all while staying bootstrapped until mid-2025 valuation talks at $15-25B.
The product wedge is narrow: RLHF, language tasks, and expert-tier work, the stuff that breaks when you route it through Mechanical Turk. Surge pays annotators $50-100+/hour and wins OpenAI, Google, and Anthropic as customers by prioritizing quality over throughput. MIT CSAIL found ImageNet had a 6% error rate that skewed model rankings for years; annotation error rates average 10% across search relevance tasks. When you're training a frontier model, those errors compound.
Here's what I find interesting: Surge's smaller pool of ~1M curated experts outperformed Scale's 240,000+ contractors on quality-sensitive work. For expert-tier tasks (legal, medical, code evaluation), smaller specialist networks often deliver better results than large crowd platforms. You're not buying headcount, you're buying judgment.
The quality paradox
Bigger annotation workforces don't always mean better throughput or quality. Surge AI's 1M curated experts outperformed Scale's 240K+ contractors on RLHF work because expert-tier tasks reward judgment and domain knowledge, not volume. When you're training models that will be deployed to millions of users, annotation errors compound, a 6% error rate in ImageNet skewed model rankings for years.Check your understanding
Match each specialist provider to its primary competitive wedge:
Mercor and SuperAnnotate: domain experts and workflow control
Mercor's wedge is even narrower than Surge's. They use AI-driven recruiting to match domain experts, lawyers, analysts, engineers, to annotation tasks that require real subject-matter knowledge. They raised at a $10B valuation and hit ~$500M annualized revenue by late 2025 by positioning as a quality-first alternative to commodity labeling. When you need legal document annotation or financial contract labeling, you can't fake domain expertise with a training manual.
SuperAnnotate takes a different angle: workflow control and computer vision depth. Their drag-and-drop workflow builder lets teams customize annotation interfaces without writing code. They serve 400+ vetted teams across 18 languages with backing from NVIDIA and Databricks. The bet here is that annotation workflows are too heterogeneous for a one-size-fits-all platform, teams want to configure task routing, quality control, and labeler UI to match their specific model training pipeline.
What ties these players together: they all win by owning a dimension Scale chose not to prioritize, expert matching, workflow customization, or quality-over-volume RLHF. The build vs. buy decision factors that matter most are data sensitivity (proprietary training data creates competitive risk when shared with vendor-controlled platforms), workflow control requirements, and strategic importance of annotation infrastructure to core AI competency. If annotation quality is your moat, you don't outsource it to a Meta-controlled vendor.
Check your understanding
A frontier AI lab is deciding whether to build internal annotation tooling or contract with Surge AI for RLHF work. What are the key factors they should weigh in this build vs. buy decision, and what would push them toward 'buy' rather than 'build'?
Don't confuse vendor size with quality
It's tempting to pick the annotation vendor with the largest workforce or biggest customer list. But from what I've seen, quality varies dramatically by task type. MIT CSAIL found a 6% error rate in ImageNet that skewed model rankings for years. For expert-tier work like RLHF, legal document labeling, or medical image annotation, smaller specialist networks often outperform large crowd platforms because judgment and domain knowledge matter more than throughput.Key takeaways
- Specialist providers win by saying no to parts of the market Scale tried to own, Surge focuses on expert-tier RLHF, Mercor on domain expert matching, SuperAnnotate on workflow customization.
- After Meta acquired 49% of Scale AI, Google and OpenAI exited due to competitive risk, not technical degradation, data sovereignty became a purchasing criterion.
- Bigger annotation workforces don't always mean better quality; Surge's 1M curated experts outperformed Scale's 240K+ contractors on quality-sensitive work.
- The build vs. buy decision is driven by data sensitivity, vendor neutrality, technical capacity, and whether annotation infrastructure is core to your AI moat.
- Even frontier labs like Anthropic contract with external providers like Surge for RLHF, they focus eng resources on model training, not annotation platform maintenance.
Your product check-in
Apply “Surge AI, Mercor, SuperAnnotate: expert-first, niche vertical, and self-serve wedges” to a product or workflow you know. What would you try, what could go wrong, and what evidence would help you decide?