The three sides: AI labs/requesters, the expert workforce, and ops/QA, who they are and what they want
Three sides, not two, why annotation is a triple-sided marketplace
Here's the thing, most people hear "annotation platform" and picture a simple two-sided marketplace: companies need labels, workers provide them. From what I've seen, the platforms that win treat annotation as a three-sided ecosystem where AI labs/requesters, the expert workforce, and ops/QA reviewers all have distinct needs and pain points. Miss any one of these sides and the whole system tips over.
On the demand side, AI labs and ML teams are burning 80% of their engineering effort on data prep, not model innovation. They want high-quality training data fast, but they're juggling cost versus quality, worrying about vendor lock-in with proprietary tools, and often flying blind on who's actually doing the work. That lack of transparency becomes a regulatory risk when GDPR or HIPAA compliance matters.
The supply side isn't a faceless crowd anymore. The market has stratified from basic labelers at $15-20/hour to domain experts with PhDs earning $40-100/hour for medical, legal, or RLHF work. These contributors want reliable payment, clear rejection criteria, and a sense of career progression, not gig-style churn. Then there's the third side: ops and QA reviewers who maintain quality at scale, adjudicate edge cases, and prevent dataset drift over months-long projects. They're balancing a 15-25% QA budget while keeping inter-annotator agreement above 0.80 (the Kappa threshold Ron Artstein and Massimo Poesio established as "strong" back in 2008).
Check your understanding
An AI lab director tells you their annotation vendor delivered a dataset on time and under budget, but when they fine-tuned their LLM, the model's performance on edge cases was terrible. They have no visibility into which annotators worked on the project or how disagreements were resolved. Which requester pain point does this scenario best illustrate?
The workforce side: from anonymous crowds to expert marketplaces
Let's get real, the idea that annotation is just low-skill click work hasn't been true for years. RLHF for LLMs requires understanding nuance, intent, and complex logic; it's cognitive knowledge work, not factory piecework. Platforms like Braintrust have placed 25,000+ contributors in 18 months by vetting for domain expertise, not just speed. They onboard 4,000+ vetted contributors monthly with under 72-hour project ramp-up, addressing the gap between general crowd workers and the medical, legal, and technical specialists requesters actually need.
From what I've observed, workforce pain points cluster around three themes: payment reliability, unexplained rejections, and lack of career path. Pure gig platforms struggle with high turnover, which kills consistency on long projects. Sama (founded by Leila Janah) demonstrated an alternative: impact-sourcing centers in Kenya, Uganda, and India where annotators stay for years, becoming highly skilled. Microsoft, Google, and Nvidia use Sama not just for quality but to address ethical supply chain concerns, something that matters more as AI regulation tightens.
The career progression piece is underrated. Platforms that let annotators level up from basic labeling to QA reviewer to expert adjudicator see better retention and dataset consistency. Anonymous crowds might work for commodity image tagging, but for the growing RLHF and domain-specialist tiers, contributors want scorecards, skill development, and recognition. (You'll see this split when we map the workflow, the "expert marketplace" model is eating share from pure crowd platforms.)
Check your understanding
Match each workforce model to the scenario where it provides the strongest competitive advantage.
The stratification insight
The annotation market isn't a single labor pool, it's stratified by skill tier. Basic labeling ($15-20/hr) versus domain expert annotation ($40-100/hr for PhDs) represents a 3-5× price and quality gap. Your platform architecture, QA workflow, and pricing model all depend on which tier you're optimizing for. Trying to serve both with the same tooling is where most platforms leak margin or quality.The ops/QA side: quality control as a budget balancing act
Here's where the system gets real, ops and QA reviewers are the unsung third side of the platform, and they're juggling constraints most people don't see. They need to maintain quality at scale, keep inter-annotator agreement above 0.80 Kappa, adjudicate edge cases without creating bottlenecks, and do it all within a 15-25% QA budget envelope. Go over 25% and you're signaling process inefficiency, not quality improvement.
The workflow mechanisms they rely on are worth naming. Gold-set monitoring uses pre-labeled test samples injected into the annotation stream to catch drift in real time. Consensus-based review runs multiple annotators per item and flags disagreements. Expert adjudication brings in senior reviewers to resolve conflicts. Platforms like SuperAnnotate report 60% faster annotation cycles and 35% higher reviewer throughput by integrating all these layers in a single system with real-time consensus tracking, versus the fragmented tool chaos most ops teams inherit.
The misconception I see most often: more QA reviewers means better quality. Reality? Quality comes from clearer standards, better tools, and consistent application, not headcount. Model-in-the-loop pre-labeling (where AI suggests labels and humans validate) can collapse QA budget by catching obvious errors early. The ops challenge is knowing when to invest in tooling versus process versus people, and that trade-off shifts dramatically depending on whether you're annotating commodity images or nuanced RLHF preferences.
Check your understanding
Scale AI's Meta contract as strategic moat
In June 2025, Meta invested ~$14.3B for a 49% stake in Scale AI, one of the largest in the industry. Scale's hybrid human-plus-model pipeline had already processed billions of annotations for OpenAI, Lyft, and Toyota. The contract demonstrated that annotation infrastructure is now a strategic competitive advantage, not a commodity service. The catch? Premium pricing at 1.5-2× generalist vendors and vendor lock-in remain pain points for requesters. For a director-level PM, the lesson is this: the winning platform isn't the cheapest, it's the one requesters can't afford to leave because it's embedded in their model training loop.Key takeaways
- Annotation platforms are three-sided ecosystems: AI labs/requesters need quality and transparency, the expert workforce wants reliable pay and career progression, and ops/QA reviewers balance quality at scale within a 15-25% budget envelope.
- The workforce has stratified from basic labelers ($15-20/hr) to domain experts earning $40-100/hr for medical, legal, and RLHF work, your platform model and tooling must match the tier you serve.
- Requesters' biggest pain points are lack of transparency into who does the work, vendor lock-in with proprietary tools, and the 80% of ML engineering effort spent on data prep instead of model innovation.
- Quality comes from clearer standards and better workflow design, not just more QA reviewers, exceeding 25% QA budget signals process inefficiency, not quality commitment.
- The platforms winning today (Scale, Sama, Braintrust) differentiate on workforce model, domain specialization, and embedding themselves into the requester's training loop so deeply that switching costs become prohibitive.
Your product check-in
Apply “The three sides: AI labs/requesters, the expert workforce, and ops/QA, who they are and what they want” to a product or workflow you know. What would you try, what could go wrong, and what evidence would help you decide?