AI data operations: What product teams outsource and what they keep

- Data preparation and annotation consume an estimated 80% of the time spent on AI projects, leaving product teams with far less bandwidth for model architecture and deployment than project timelines assume.
- The functions that can be outsourced in AI data operations (annotation, evaluation, red teaming, content moderation) are labor-intensive and require human judgment at scale, not just processing capacity.
- What product teams keep in-house is not the work itself, but the decisions: what the model should optimize for, how evaluations will be designed, and what constitutes a failure.
- Hugo provides AI data operations services across annotation, evaluation, red teaming, speech and audio processing, agentic and coding tasks, and Trust and Safety, with Google and Meta among its clients.
Most AI product teams underestimate how much of their engineering capacity gets absorbed by data operations.
Building a model is one phase. Keeping it accurate, safe, and improving across deployment is continuous work, and most of it is not engineering work. It is labeled data, structured evaluation, adversarial testing, and content moderation.
The build-it-in-house instinct is understandable. Data operations touch the model directly, and teams are protective of that.
But the instinct conflates two different things: the work of generating training data and evaluations, and the strategic decisions about what those evaluations are designed to measure.
The first can be outsourced. The second cannot.
Understanding where that line sits is what determines whether a product team’s AI capability scales or stalls.
Deloitte’s 2024 Global Outsourcing Survey found that 83% of executives are already leveraging AI as part of their outsourced services, a figure that reflects how quickly AI work has moved from internal-only to distributed execution.
For product teams actively building AI systems, that distribution extends to the human data operations layer.
Most teams get this wrong because they have never defined what “data operations” actually covers. So that is where we start, before working out what to keep and what to hand off.
What AI data operations actually includes
AI data operations are not a single function. It is a set of distinct workstreams that sit between raw data and a trained, deployed model.
Data annotation and labeling
The foundational layer. Annotation includes image labeling, text classification, bounding box drawing for computer vision, audio transcription, sentiment tagging, and any other task that attaches structured metadata to raw inputs so a model can learn from them.
Data preparation and annotation consume the majority of total AI project time, by some estimates as much as 80%, a reality highlighted by the data-centric AI movement led by Andrew Ng. The volume of labeled data required for competitive model performance makes in-house annotation at scale impractical for all but the largest AI organizations.
Model evaluation (Evals)
Evals assess model behavior against defined benchmarks. They require human raters to judge whether model outputs are accurate, helpful, safe, or aligned with intended behavior, judgments that automated metrics cannot reliably capture for open-ended generation tasks.
Red teaming is a specialized form of evaluation where human testers actively try to elicit harmful, dangerous, or policy-violating outputs. Agentic evals test AI systems that take real-world actions, not just generate text.

Trust and safety and content moderation
Deployed models generate content. Some of that content requires human review to determine whether it violates policy, law, or community standards.
Trust and Safety operations include reviewing flagged outputs, training classifiers that automate moderation at scale, and managing the evolving edge cases that classifiers miss.
This is high-volume, psychologically demanding work that product teams consistently understaff relative to the actual moderation burden their deployment creates.
Pro Tip: Before scoping an AI project, map the full data operations workload across annotation, evals, and moderation for the first 12 months post-launch. Most teams scope for model development and discover data operations at 2x to 3x their initial estimate only after deployment.
What product teams outsource
The functions that scale with human labor hours rather than engineering decisions are the right candidates for outsourcing.
Annotation at volume is the clearest case. A team can design the annotation schema and quality standards internally, then hand execution to a specialized provider with trained annotators, QA infrastructure, and throughput that matches the dataset size required.
The same logic applies to evaluation: product teams define what good looks like, write the rubric, and set the pass/fail criteria. Trained human raters do the rating.
Red teaming is less intuitive as an outsourced function, but the adversarial mindset required for effective red teaming benefits from an external perspective.
Internal teams know what the model was built to do and tend toward blind spots in areas outside that intent. External red teamers approach the model without that framing.

Trust and Safety moderation is the third category. At deployment scale, the volume of content requiring human review exceeds what internal teams can sustain.
Moderation at volume, with appropriate category training and escalation protocols, is a specialist function, not a general support role.
Pro Tip: When handing off annotation or eval work to an external provider, the most important document is the rubric, not the data. Spend time on the rubric before handing off the task. Annotator disagreement almost always traces back to ambiguous instructions, not annotator error.
| Function | Outsource | Keep In-House |
|---|---|---|
| Data annotation and labeling | Yes — scales with labor hours, not engineering headcount | |
| Eval rating and human feedback | Yes — execute under the team’s rubric and pass/fail criteria | |
| Red teaming | Yes — external perspective surfaces blind spots internal teams miss | |
| Trust and Safety moderation | Yes — deployment volume exceeds what internal teams can sustain | |
| Eval design and rubric authorship | Yes — defines what the model optimizes for | |
| Model architecture decisions | Yes — sets capability ceilings and tradeoff profiles | |
| Training strategy | Yes — data selection, proportions, RLHF signals | |
| Policy and safety thresholds | Yes — what the model refuses and why; not delegable |
What product teams keep in-house
The decisions that define what the model is optimizing for cannot be outsourced without losing control of the product. These include:
- Eval design: what the model will be measured against, what metrics define success, and what failure modes are prioritized for detection
- Model architecture decisions: the structural choices that determine capability ceilings and tradeoff profiles
- Training strategy: what data gets used, in what proportion, and what RLHF or preference optimization signals are incorporated
- Policy and safety threshold decisions: what the model will refuse to do, what content categories require special handling, and what constitutes acceptable risk in deployment
The distinction matters because the line between “design” and “execution” in AI data operations is meaningful. A provider can execute annotation at scale under the team’s rubric. They cannot set the rubric’s values.
A provider can red team within a defined threat model. They cannot determine what constitutes a threat worth modeling for a specific product context. That strategic layer stays in-house.
How Hugo supports AI product teams
Hugo provides AI data operations services built around the Omni Evals framework, a six-layer evaluation and annotation system covering the full range of tasks that AI product teams need at scale.
Hugo’s client base includes Google and Meta, reflecting the level of operational maturity required to work within enterprise AI development pipelines.
- Perception and Annotation: image, video, and text annotation across classification, bounding box, segmentation, and entity tagging tasks at production volume
- Speech and Audio Processing: transcription, speaker identification, audio classification, and language-specific annotation for speech model training and evaluation
- Multimodal Evaluation: human evaluation of models that process multiple input types simultaneously, including image-text, audio-text, and video-text interactions
- Agentic and Coding Evaluation: testing AI systems that take real-world actions (code generation, tool use, multi-step reasoning) against structured benchmarks requiring human expert judgment
- Red Teaming: adversarial testing designed to surface harmful, policy-violating, or unsafe model outputs before deployment
- Trust and Safety: content moderation and policy enforcement at the scale and category depth required by consumer-facing AI deployments
See how Hugo supports AI product teams at scale. Book a call with Hugo.
Key takeaways
- Data preparation and annotation consume approximately 80% of AI project time, making data operations the largest labor category in AI development regardless of whether it appears in engineering headcount.
- Annotation, model evaluation, red teaming, and Trust and Safety moderation are the four AI data operations functions that scale with human labor and are candidates for outsourcing.
- What product teams keep in-house is not the execution of data operations work, but the decisions: eval design, training strategy, model architecture, and policy and safety thresholds.
- Hugo provides AI data operations services across the full Omni Evals framework for clients including Google and Meta, covering annotation, speech, multimodal, agentic, red teaming, and Trust and Safety functions.







Independent




