返回知识库
0

title: "深入 Clay 评测技术栈:3 亿次智能体运行与一条 LangSmith 管线"
source_url: "https://www.youtube.com/watch?v=Uny6LpmjraI"
author: "LangChain"
excerpt: "公共早报 Clay 说明如何通过分层 LangSmith 评测框架,把本地检查、生产信号、漂移控制和面向智能体的数据基础连接起来,评估高规模 GTM 智能体。"


Brief Description

Clay explains how its go-to-market agents reached a scale where manual trace review is no longer possible, and how the company built a layered evaluation stack around LangSmith. The talk covers local, staging, and production evaluation; the need to feed production learning back into offline tests; and Clay's move toward an agent-first data foundation that lets internal and external agents use the same tools.

Table of Contents

  • Claygent, Sculptor, and the scale problem

  • Why evaluations became non-negotiable

  • A layered, extensible evaluation harness

  • Coverage, drift, and production feedback loops

  • An agent-first interface and shared tool flywheel

  • Building a unified data foundation for agents

Claygent, Sculptor, and the scale problem

Clay has worked on agents for some time. Its first agent, Claygent, launched in 2023 as a go-to-market research agent: it searches the public web, helps research and qualify companies, and can also search a customer's first-party datasets. Claygent now runs at substantial scale, with more than 300 million runs per month.

Clay later launched Sculptor, a go-to-market engineering agent. Sculptor helps users build and orchestrate go-to-market workflows in the product, analyze data, and use Clay's company and contact database to find leads or prospects. It has become one of the main ways people interact with Clay, with more than 100,000 messages sent to it each week.

At that volume, Clay cannot inspect every trace, workspace, or user interaction. Production quality for capabilities such as Sculptor for Search took a long journey, because many things can go wrong. That scale is why Clay turned its attention to the ways it evaluates and controls quality.

Why evaluations became non-negotiable

When Clay first built agentic products, its evaluations were not very strong. As Claygent reached billions of runs and Sculptor began performing end-to-end and long-running tasks, evaluations became non-negotiable.

A good evaluation suite makes agentic development safer. Teams can let tools such as Claude, Codex, Devin, or LangChain Engine propose prompt changes while retaining confidence that nothing harmful will be shipped to production. Clay has therefore spent significant time building a more comprehensive suite and reconsidering its philosophy for agent evaluations.

The first requirement is to match the environment to the purpose. Local-development evaluations should be low lift, cheap, and fast. Evaluations that run in CI or staging should resemble the production harness as closely as possible. Locally, Clay deliberately avoids production concerns such as its sandbox and VFS so developers can run a command-line evaluation suite where they already work, rather than provisioning a managed agent or creating a separate experiment.

A layered, extensible evaluation harness

Clay wants evaluations to persist and be versioned. Although they can run locally, the results are written to LangChain so they can be stored and compared. The same harness must also be extensible across the product: as Sculptor takes on more work in Clay, different teams should be able to bring their own evaluation suite and evaluators, including their own LLM judges, while reusing the underlying harness as a plug-and-play foundation.

The team thinks about evaluation coverage in several categories. For deterministic offline checks, it uses golden examples. Goldens work well for simple tasks, such as a demonstrated query language, but become too static for complicated tasks: changing keyword order or node order can make a valid result fail. Noisy evaluations are ignored, so Clay moved toward structured checks that examine only the parts of a query that matter and tolerate irrelevant variation.

Another category is trajectory or tool assertions. For example, when an agent answers a pricing question, the evaluation can confirm that it actually read the pricing scale. Offline, non-deterministic work can use an LLM as a judge. Multi-turn evaluation can use a simulated user driven by historical traces, or a deterministic version with hardcoded user turns.

Clay found deterministic multi-turn evaluations especially useful during development. Simulated agents were often noisy and became another agent to maintain, update, and evaluate, so the added complexity was not always worthwhile.

Coverage, drift, and production feedback loops

Moving toward online and deterministic evaluations, Clay tracks objective A/B-testing metrics such as latency, cost, whether users move from chat into other parts of the product, whether they get stuck, and whether they rage quit. For online, non-deterministic signals, it uses LangChain's primitives and evaluators such as NPS, user-satisfaction measures, and perceived-quality evaluators that detect whether users correct or guide an agent. LangChain Engine can bulk-analyze production traces, while humans still manually inspect traces as well.

The difficult but essential part is the feedback loop: what the team learns in production must inform offline evaluations. Evaluation and production drift remain unsolved problems in the agent space. Data drift can mean production use cases differ from the cases a team tested and bug-bashed. Judge drift can arise because model families have their own biases; optimizing too aggressively for one LLM judge can overfit to it. A small evaluation set can similarly cause a prompt to mirror only those examples.

To counter this, Clay brings examples from online evaluators into its evaluation work. Customer-support tickets provide high-signal feedback, human-annotated goldens help judge drift, and use-case classifiers and internal tagging help verify that the suite covers real production use cases.

An agent-first interface and shared tool flywheel

Clay is becoming an agent interface, not only a web UI. The company wants every capability available in the interface to be available through a CLI and a public API as well, so both internal and external agents can use it.

This creates a flywheel. The same tools exposed through the API and CLI are also available to internal agents such as Sculptor. As agents invoke those tools, failures in tools or trajectories create user signals that can improve the agent harness and the tools themselves. Those improvements benefit both internal and external users. Clay uses Engine to examine such traces, alongside human, vibe-based evaluation.

Building a unified data foundation for agents

Scaling these learning loops is difficult when data primitives are fragmented. Clay is moving toward a data-lake architecture that brings first-party and third-party data into one platform for agents. In this model, agents are the first-class user, so the system needs guardrails from the start and safe shadow builds. Agents can build new data models and deploy them on S3 while serving and development compute remain separated, allowing experiments without bringing down production.

Clay is investing heavily in skills and a CLI so agents can access this data natively. This enables proof-of-concept work to emerge quickly, lets large models work with data at scale through systems such as Athena, and supports long-running goals such as building a data model over one or two hours.

The timing reflects a step change in model capability. Clay has seen agents handle large-scale context, including requests to examine 10,000 examples and find trends, where earlier workflows were largely vibe-based and relied on inspecting only a few examples. New subagents, goals, and harnesses also make it possible to iterate quickly by setting up evaluations first and driving toward them.

Clay's data had previously been spread across LangChain traces, Snowflake analytics, Postgres, and first-party data in ClickHouse. Bringing those sources into one platform removes the need for agents to tie together multiple databases themselves. The intended end state is a self-iterating loop: customer and third-party data flow into a unified foundation, agents reason over it and execute work, results feed back into the system, and the product can build better iterations of itself.

AI知识库 / 深入 Clay 评测技术栈:3 亿次智能体运行与一条 LangSmith 管线 0 字 0 行 iliuqi
2026-09-04T08:55:45.332932914Z 2026-09-04T09:40:28.589753547Z