title: "RLM 如何重塑长上下文处理:从 Token 到可执行状态"
source_url: "https://www.youtube.com/watch?v=xo68uCibfm8"
author: "AI Engineer"
excerpt: "公共早报 Kevin Madura 介绍的 RLM 将长上下文保留为 REPL 中可操作对象,让模型编写代码、递归拆分任务,再取回下一步需要的结果。它为文档、日志和大型代码库提供不同于堆叠提示词的思路:给模型可执行环境与目标,让它决定检索、计算和委派路径。对处理密集长输入的团队,这是一条值得评估的工程路线。"
Brief Description
Kevin Madura of AlixPartners introduces recursive language models (RLMs): an approach in which a language model works with long context as a symbolic object in its execution environment, can write code over that context, and can delegate portions of a problem to submodels. He contrasts the approach with ordinary tool calling, RAG, and coding agents; reviews reported long-context results; and walks through practical examples involving data frames, documents, traces, and large codebases.
Table of Contents
- What recursive language models are
- Why RLMs help with long context
- A practical mental model for RLM systems
- RLMs compared with RAG, tools, and agents
- When to use RLMs
- Benchmarks and production-system trade-offs
- Libraries and a data-frame example
- Real-world applications and the outlook
What recursive language models are
Kevin Madura introduces RLMs, or recursive language models. Their first important difference from a normal tool call is that they treat context as an object which the model can interact with symbolically in its environment. In a typical tool call, JSON or another string is sent to another program and a string is returned. With an RLM, the model works inside a symbolic execution environment, typically a REPL or Python REPL.
The second difference is delegation. An RLM can call another language model, often itself, with particular parameters. Because that call also lives in the execution environment, the system can recursively decompose a problem. The model can decide how to apply logic, interpret information, or write code to solve a subproblem; then a submodel can make the same decisions at the next level.
Madura adds a design principle shaped by hard-won experience: as models improve, the implementation should defer more to the model's ability to determine what it needs to do. The RLM is intended to give the model a capable environment and clear objectives rather than force every intermediate decision into a rigid hand-authored flow.
Why RLMs help with long context
One early precursor was an idea from Omar, an adviser to the RLM work, using DSPy and related techniques to accept inputs of arbitrary length. A representative use case was summarizing an arbitrarily long document and producing a table of contents and summary. For Madura, the key implication was that context windows might not always need to be managed directly: there may be ways to work beyond them with the right techniques.
He points to reported results on long-context tasks. The ULong benchmark measures a model's ability to answer questions about excessively long context, while BrowseComp requires iterating through a large body or corpus of text to answer specific questions. In the displayed results, RLM performance leads several other models. He also notes that a tool-calling approach using GPT-5 with a BM25 tool appeared more expensive while performing worse than the RLM in the comparison shown.
The core behavior is that the RLM receives the input and the model determines how to decompose it, what processing is required, and what code to write. It is tightly integrated with the REPL and can define the work it needs to do rather than simply consuming the entire input as a flat prompt.
A practical mental model for RLM systems
Madura describes a relatively deterministic outer shell. A developer defines the intent and the task to accomplish, then specifies inputs, desired outputs, and guidance for the model. The instruction can be simple: these are the expected inputs, this is the desired result, and the model should determine the remaining implementation.
He sees this pattern in DSPy and in RLMs. The developer does not need to prescribe every detail of the implementation in the middle. Instead, the system makes guarantees about inputs and outputs, while the model determines how to bridge them. This shifts effort away from micromanaging an implementation and toward defining an outcome and its constraints.
That model becomes especially valuable when a problem naturally decomposes into smaller investigations. An RLM can retain a large source object in the environment, choose a relevant slice, perform computation against it, delegate a focused task to a submodel, and bring back only the result that matters for the next decision.
RLMs compared with RAG, tools, and agents
Madura connects the idea to context engineering and the problem sometimes called context rot: after a context window becomes sufficiently full, model performance begins to degrade. RLMs partly avoid that issue because the main model does not have to keep the entire source context exposed in its own token window. The context can live as a variable in the REPL, while the model chooses how to access it and sends subtasks to submodels. The main model receives the pieces that are useful rather than every token of the underlying source.
He contrasts this with RAG, which often places retrieved context directly into the context window. Coding agents and tool calls also commonly pass strings back and forth, without the same tight coupling between the logic, execution, and intermediate results. That can recreate the context-management problem. In an RLM, the language model interacts with context and results as variables in the REPL, allowing additional computation instead of requiring the model to attend directly to every relevant token at once.
The distinction is also relevant to coding agents. Tool calls are generally string exchanges, whereas an RLM works through variables and computation in the environment. Madura observes that Anthropic workflows appear to use a related idea: intermediate results live in script variables. He mentions a conference discussion that identified the RLM paper as an influence on workflows, suggesting that these techniques are also informing broader work from model labs.
When to use RLMs
RLMs are most useful for large or dense input context. Madura also calls outputs an underexplored opportunity: if a task needs to generate hundreds of thousands of lines, an RLM may be a good candidate. The approach fits tasks that can be decomposed.
For example, a system might need to inspect an entire tax code for potential loopholes. It cannot place all of that material into context at once. An RLM could work through the material iteratively, use subagents to examine interesting sections, return those sections or findings, and reason over them. The same logic applies to longer-horizon sessions.
He also explains when not to use the approach. It is less appropriate for work that already fits in context, needs very low latency, or is better served by a model that is already a strong enough coder on its own.
Benchmarks and production-system trade-offs
Madura references performance testing by Raymond on a long chain-of-thought benchmark. The reported overall accuracy rose from 2.6 to 45.4 percent on many tasks. The approach performed particularly well on work amenable to code, including logic puzzles, chess, and chemistry: the system can write code, pull in only relevant context, compute over it, and return a result for the main model to use.
He gives a simple example: finding the sum of twelve numbers buried across thirty thousand tokens. A base LLM attempting to attend to everything and calculate the answer by itself may be less reliable than a method that writes a regular expression or comparable code. Data frames make the contrast clearer. If the model can interact with a data frame inside the REPL, it can understand and iterate through the content more efficiently than repeatedly sending JSON-string tool calls back and forth.
He also tested similar experiments with a coding agent. He has not studied every detail and cautions that the comparison may contain unfair assumptions, but the coding-agent approach appeared bloated in how it tried to solve the tasks. More experiments are needed to compare base models, RLMs, and coding agents. For production workloads, his conclusion is that it is risky to rely on an oversized prompt and hope for a good result. A structured pipeline with defined inputs and outputs can lower cost, complexity, and bloat, which is where RLMs can be valuable.
Libraries and a data-frame example
Several open-source libraries implement RLM-like ideas. Some are more directly RLM-focused, such as PredictRLM; others integrate the technique into a broader framework. Madura mentions DSPy, AIX, PredictRLM for knowledge work involving spreadsheets and PDFs, and FastRLM. He also cites an example from Sam Hogan of inference.net, where an RLM extracts insights from production workload traces and determines what may be worth deferring to another model, iteratively and automatically as traffic flows through the system.
The practical promise is that a developer can spend less effort on context engineering and allow the RLM to determine how to work through the data. Madura then sketches a cohort-retention analysis. A data scientist might receive three data frames, a description of what to examine, and the desired output types. That can be enough for the RLM to begin.
The RLM has its own REPL and can directly interact with the data frames. It reasons about the next step, writes code, and executes it without the additional bloat of repeated tool calls. It behaves more like it is working in its own Jupyter notebook. It can inspect outputs, iterate, and, where useful, delegate a large subset of the data to a submodel for analysis before incorporating the returned result.
Madura notes that the Compound platform is RLM- and DSPy-native and exposes the reasoning breakdown: generated code, key findings, recommendations, and the final typed output defined up front. The model itself decides when to stop. A developer can set a maximum number of iterations, such as ten or one hundred, but the model explores the data and submits a final result when it judges the task complete. As models improve, he expects developers to be able to defer more of this work and operate at higher levels of abstraction, as long as they can define the objective.
Real-world applications and the outlook
Madura turns to case studies. He mentions PredictRLM and work by Trampoline AI on knowledge work involving PDFs and spreadsheets. Consider consolidating a directory of invoices into one inventory. Invoices can be long, inconsistent, and complicated. Without an RLM-like framework, processing them can quickly require complicated strategies. With an RLM, a two-hundred-page invoice or contract can be processed directly, producing a result without requiring the developer to manage chunking, embeddings, and other context-window workarounds.
The benefit is that teams can focus on the abstraction and the outcome they want rather than context engineering. For DSPy users, he notes that PredictRLM uses DSPy to determine schemas between the main language model and submodel calls. That provides readability and maintainability because the intended goal is clearer, and it can make the model more precise about the data types expected from submodels. Madura expects this kind of enforced structure may also improve performance for cheaper models, because inputs and outputs are specified and typed handoffs are enforced.
He also describes an experiment from an AWS engineer using arbitrary log data to surface useful results. Another project, Halo, uses an RLM to inspect traces from agent tasks. Rather than optimizing only a particular workflow or framework structure, Halo iterates on the harness itself: a kind of meta-optimization. Because traces can be long and complicated, an RLM can take in the context and use the trace structure to recommend a better harness.
Finally, Madura describes a deliberately vulnerable web application used in a small experiment. With a short amount of code, an agent can work through roughly five hundred thousand lines of code and generate a security report. The point is not the particular example, but the ability to feed an arbitrarily sized codebase into the method and obtain insights without extensive context engineering or additional surrounding structure.
He closes by acknowledging that the presentation moved quickly and inviting questions. The larger promise, in his view, is a future in which models are post-trained to understand RLM methodology natively. When models know how to use this approach directly, their ability to take advantage of recursive decomposition and symbolic context could advance rapidly.