title: "跨职能团队如何做 AI Evals:DoorDash 的质量运营闭环"
source_url: "https://www.youtube.com/watch?v=bMjlRrWjdT0"
author: "AI Engineer"
excerpt: "公共早报 DoorDash 认为,可靠的 AI 评测是一套跨职能运营闭环:共享的遥测与工作流让领域专家、运营人员、产品团队和工程师能够共同创建、校准并持续改进质量判断。"
Brief Description
DoorDash's GenAI platform team explains why evaluation is not merely an engineering harness but a cross-functional operating system for shipping reliable AI. The talk describes a platform that balances accuracy, latency, and cost while enabling domain experts, product teams, operations, and engineers to build, annotate, calibrate, and monitor AI quality together.
Table of Contents
A shared platform for AI quality
Evaluation as a cross-functional loop
Tracing, annotation, and flexible workflows
Calibrating judges without engineering bottlenecks
Results and the continuing quality loop
A shared platform for AI quality
Speaker 1: DoorDash's GenAI platform team is a horizontal team serving product teams that build on its infrastructure and primitives. Its value is helping those teams balance three forces: accuracy, latency, and cost. That began with model choices, but the same trade-offs now apply to agents as well.
The team provides an LLM gateway that lets teams switch models and try new ones, plus an agent gateway for connecting tools and other agents. The gateway centralizes authentication and agent identity so that security teams can approve the common solution instead of each team reinventing it. The platform also pairs model access with open-weights model hosting, an investment that has already reduced costs. Evaluation is the fourth pillar and the focus of the talk.
Product teams had very different needs. A consumer discovery and shopping assistant needed session-level quality judgments; personalization ML needed a way to scale human judgment; and multi-agent systems required trajectory-based evaluation. The platform question was how to serve those distinct needs without fragmenting the infrastructure.
Evaluation as a cross-functional loop
The answer was to empower domain experts, not only engineers. At DoorDash, these experts include strategy and operations, product managers, and labeling partners. The initial direction was UI-first so non-engineers could contribute. The platform then became API-first so engineers could build their own systems without waiting on a central team, and workflow-first so strategy-and-operations staff and product managers could run operations with coding agents.
Evals are therefore a team sport. Domain knowledge must enter the quality system through traces, datasets, and scoring mechanisms. Strategy and operations set priorities and the desired quality bar. Product people turn requirements into rubrics and workflows. Operations teams run annotations. Engineering provides APIs, telemetry, datasets, and judges. Combining these responsibilities is how the organization ships quality AI products.
The team frames the process as a continuous loop: capture and inspect traces and sessions, sample a manageable set, annotate it with domain expertise, review the annotations, create golden datasets, calibrate against them, monitor performance over time, and repeat. This makes quality improvement a recurring operational practice instead of a one-time model test.
Tracing, annotation, and flexible workflows
At the platform level there are two main surfaces. The telemetry layer contains traces, scores, and observations, and it exposes that data through MCP, SDK, and APIs. The workflow layer is where strategy, operations, and product teams create annotation tasks, review golden datasets, create judges, and calibrate those judges.
Tracing and sampling make visible what agents and LLMs actually produce. Stable table APIs power the platform's UIs and provide access to scores and datasets through the same common plane. The intended lifecycle begins by capturing traces and sessions and measuring scores, then adds human judgment, context, and domain knowledge before calibrating judges.
Annotation is where teams inspect agent behavior and sessions to find what went well and what failed. A horizontal platform cannot practically build a bespoke UI for every use case. The common pattern is still recognizable: a strategy-and-operations person decides what to annotate, an annotator labels the dataset, and the platform provides the APIs.
Because coding agents are available to the users of the platform, DoorDash doubled down on the API-first approach. Strategy-and-operations teams can use coding tools to build their own annotation UIs for different tasks, including image annotation and manual testing. The underlying interaction patterns are similar, so a simple, locally useful UI can be created by a partner while the platform retains the shared primitives. The key outcome is that the workflow is placed in the hands of operators, who can build and evolve the interfaces they need.
Calibrating judges without engineering bottlenecks
Once annotations exist, teams need to improve the prompts behind LLM-as-a-judge metrics. The process starts with a judge prompt and a clear definition of what should be measured. Teams run the judge against traces to establish baseline scores, then use an optimization loop. DoorDash uses the JPEA library for prompt optimization; once partners are satisfied, they promote the prompt into their judge.
Prompt calibration may sound straightforward, but it remains a new and evolving field. The team wanted to remove the friction of going back and forth with engineers, so it hid complicated logic behind a self-service UI. A product manager or operator can set platform configurations and run a calibration loop without having to understand every tuning setting, and can choose a model for that loop.
Making calibration reviewable was equally important. The platform gives partners visibility into what changed by showing the original system prompt beside the calibrated prompt. This helps users understand why a change was made and build trust in the process. It also supports different organizational arrangements: in some teams strategy and operations owns a prompt, in others a product manager or engineering owns it. The platform gives teams room to design and evolve those responsibilities while the organization continues learning.
The broader aim is self-service. People should not always be blocked on the central platform team. The team began with UIs, has become API- and workflow-first, and tries to reuse existing DoorDash infrastructure rather than create parallel systems.
Results and the continuing quality loop
The team reports that self-service annotation reduced per-annotation spending while increasing velocity. At DoorDash scale, thousands of rows may need annotation each week, so the cost can become substantial. A platform that lets teams work directly with annotation and judge calibration has reduced the cost spent on annotators and accelerated iteration.
Teams can now iterate faster and calibrate their own judges in a self-service way. The closing message returns to the same loop: inspect traces and sessions, sample to a manageable size, annotate datasets with human and domain knowledge, improve the data, calibrate workflows, agents, and LLM judges against golden datasets, and repeat over time to ship reliably with high quality.
The speakers emphasize that the loop is not a fixed engineering feature. As product teams, operators, and engineers learn, the platform and even the organizational design improve together. Evaluation succeeds when the platform makes that continuous, shared learning practical.