Recursive Language Models (RLMs) have proven a powerful harness architecture for Large Language Models, enabling dynamic agent workflows (Anthropic), extremely long context processing (MIT), SOTA results on “long chain-of-thought” benchmarks (RAW.works), and even best-in-class performance on benchmarks such as ARC-AGI 3 (Prime Intellect).
“Recursive Decision Models” (RDMs) is an exploration of how the principles of RLMs can be applied to “decision models” (aka “System One models” aka “classifiers”) such as Jev.

Briefly, a note on the naming: I don’t particularly care what term people choose for the “genre of Jev”. Diogo officially prefers “System One models”, and has explicitly argued against the alternative: “Some people are trying to call them decision models… I wouldn’t do that, because I think there’s other types that are machine-native that are not decisions.” Nevertheless, the zeitgeist seems to be trending towards “decision models” and I’m not willing to wait any longer for people to make up their mind.
In developing this concept of RDMs, I’ve been working from three “non-negotiables” of how the system must behave:
-
The RDM system must externalize context and be able to programmatically decompose slices of that context.
-
The decision model must have the choice to call another decision model, up to some maximum depth of recursion.
-
The RDM system must be able to work with and without the presence of any generative model or function. In other words, it must stand alone as “pure decision models” and also compose with LLMs and agents.
I initially started exploring some of these ideas in a repo that I dubbed “one-system”, which started as a “LiteLLM for Jev” and then I quickly realized could support recursion by composing requests through the gateway. Lucky for me, the DSPy team (in particular my esteemed peers at cmpnd.ai) has rapidly embraced Jev. Now that the TypeSafe System One API is supported in DSPy, my task of conveying the essence of Recursive Decision Models is greatly simplified.
A Simple RDM
Let’s start with a dspy.Module that recursively makes decisions up to its max depth.
import dspy
class RDM(dspy.Module):
def __init__(self, signature, tools, max_depth=3):
super().__init__()
self.predict = dspy.Predict(signature)
self.tools = {tool.__name__: tool for tool in tools}
self.max_depth, self.depth = max_depth, 0
def forward(self, **inputs):
answers = self.predict(**inputs) # one request to the decision model
observations = {}
if self.depth < self.max_depth:
self.depth += 1 # anything a tool calls is one level deeper
for name, tool in self.tools.items():
if answers[name].value: # a yes runs the tool of the same name
observations[name] = tool(**inputs)
self.depth -= 1
return dspy.Prediction(observations=observations, **answers)
As a simple example, we write a signature that asks if the answer to a question is in a chunk of text, and a function to divide and conquer by recursively halving the content.
from dspy.experimental import Noul, TypeSafe
class Search(dspy.Signature):
"""Search a text for the answer to a question."""
text: str = dspy.InputField()
question: str = dspy.InputField()
look: Noul = dspy.OutputField(desc="Is the answer to `inputs.question` stated in `inputs.text`?")
def look(text, question):
lines = text.splitlines()
if len(lines) <= 1:
return lines # down to one line, and the model said yes to it
middle = len(lines) // 2
found = []
for half in lines[:middle], lines[middle:]:
found += rdm(text="\n".join(half), question=question).observations.get("look", [])
return found
To finish the specific example, we can search a handbook recursively to find the specific line that answers a query.
handbook = open("handbook.md").read() # 71 lines
rdm = RDM(Search, tools=[look], max_depth=10)
rdm.set_lm(TypeSafe("jev-latest")) # reads TYPESAFE_API_KEY
result = rdm(text=handbook, question="What is the nightly hotel limit in London?")
print(result.observations["look"])
# ['The nightly limit for accommodation is 180 euros in most cities and 260 euros in London, ...']
# 15 requests, each a yes or a no
While this is a terrible way to actually solve this problem in the real world, it helped us demonstrate how we can use recursion to have decision models programmatically process a potentially very large context.
Self-similarity and external state
Readers familiar with RLM will probably be offended by the last example, so let’s now build out an analog that clearly demonstrates “the shape of RLM”. Specifically, that means we need to:
-
Have RDM call RDM all the way out to a leaf
-
Explicitly externalize the context
-
Keep it LLM-free to prove the point that we don’t need generation for this to work
The halving search sent the whole handbook to the model in its first request. This time the handbook stays in Python as state, and no request ever shows it whole. Plain code turns it into an outline: each heading, with the headings and passages directly under it.
def outline(markdown, title):
"""{heading: [the headings and passages directly under it]}, made by plain code."""
under, path = {title: []}, [title]
for block in markdown.strip().split("\n\n"):
depth = len(block) - len(block.lstrip("#"))
if depth:
block = block.lstrip("# ")
path = path[:depth] + [block]
under[block] = []
under[path[-2] if depth else path[-1]].append(block)
return under
state = outline(handbook, "Handbook") # stays in Python: no request ever shows it whole
def headings(entry):
"""Every heading below an entry: all that a request shows of what is under it."""
return [below for each in state.get(entry, []) if each in state for below in [each, *headings(each)]]
A request shows one entry and the headings below it, never the text under them. The same question is asked at every node, whether it is the whole handbook, a section, or a single passage.
class Navigate(dspy.Signature):
"""Search a handbook for the answer to a question, one entry at a time."""
entry: str = dspy.InputField(desc="A heading of the handbook, or a passage under one.")
contains: list[str] = dspy.InputField(desc="The headings under `inputs.entry`. Empty for a passage.")
question: str = dspy.InputField()
explore: Noul = dspy.OutputField(
desc="Does `inputs.entry` state the answer to `inputs.question`, or does `inputs.entry` or any "
"heading listed in `inputs.contains` name the topic that `inputs.question` asks about?"
)
def explore(entry, contains, question):
if entry not in state:
return [entry] # a leaf: a passage, and the model said yes to it
print("explored:", entry)
found = []
for each in state[entry]:
found += rdm(entry=each, contains=headings(each), question=question).observations.get("explore", [])
return found
A yes runs explore, which calls the same module with the same question on each entry under this one, all the way out to a passage. The RDM class is the one from the first section, unchanged. There is no LLM anywhere: Jev is the only model, and all it ever says is yes or no.
rdm = RDM(Navigate, tools=[explore], max_depth=10)
rdm.set_lm(TypeSafe("jev-latest"))
assert dspy.settings.lm is None # no LLM is configured anywhere
result = rdm(entry="Handbook", contains=headings("Handbook"), question="What is the nightly hotel limit in London?")
print(result.observations["explore"])
# explored: Handbook
# explored: Expenses
# explored: Travel
# explored: Hotels
# ['The nightly limit for accommodation is 180 euros in most cities and 260 euros in London, ...']
# 15 requests again, but the largest shows 272 characters of a 2,555-character handbook
Hierarchical search, without the beam
Let’s now refine the concept of Recursive Decision Models by remixing some of the Jev best practices. TypeSafe’s hierarchical classification cookbook walks a taxonomy with one Choice per node, and keeps the three best paths with a beam search that scores each path by the probabilities along it. The same walk falls directly out of recursion, so we can use the same recursive pattern to solve the problem without writing the beam search code explicitly. The external state here is the cookbook’s own Shopify product taxonomy: over twelve thousand categories that no request ever shows whole.
from urllib.request import urlopen
URL = "https://raw.githubusercontent.com/Shopify/product-taxonomy/v2026-02/dist/en/categories.txt"
taxonomy = {"All products": []} # {category: [the categories directly under it]}
for line in urlopen(URL).read().decode().splitlines()[3:]:
path = line.split(" : ")[1] # "Animals & Pet Supplies > Pet Supplies > Cat Supplies"
taxonomy[path.rpartition(" > ")[0] or "All products"].append(path)
taxonomy[path] = []
At each category the model is asked which of the categories under it fits best. The function follows every one that is at least half as likely as the likeliest, by calling itself, and when more than one leaf comes back it asks the same question again among those. There is no beam width and no path score: the model’s own odds decide where the search branches.
from dspy.experimental import Choice
predict = dspy.Predict("listing -> category")
predict.set_lm(TypeSafe("jev-latest"))
def choose(listing, options):
"""One request: how likely each category is to be the best match for the listing."""
if len(options) == 1:
return {options[0]: 1.0}
signature = predict.signature.with_updated_fields(
"category",
type_=Choice[tuple((option, None) for option in options)],
desc="Which category best matches the product in `inputs.listing`?",
)
return predict(signature=signature, listing=listing).category.probabilities
def classify(listing, category="All products"):
under = taxonomy[category]
if not under:
return category # a leaf
odds = choose(listing, under)
likely = [each for each in under if odds[each] >= max(odds.values()) / 2]
print(category.rpartition(" > ")[2], "->", [each.rpartition(" > ")[2] for each in likely])
found = [classify(listing, each) for each in likely] # the same function, one level down
odds = choose(listing, found) # more than one came back: ask again, among those
return max(found, key=odds.get)
listing = (
"Furniture listing: a wall-mounted window shelf bed. This padded floating shelf uses "
"suction cups and a washable cushion as a sunny perch for one cat."
)
print(classify(listing))
# All products -> ['Animals & Pet Supplies']
# Animals & Pet Supplies -> ['Pet Supplies']
# Pet Supplies -> ['Cat Supplies', 'Pet Beds']
# Cat Supplies -> ['Cat Furniture']
# Cat Furniture -> ['Cat Window Beds & Perches']
# Pet Beds -> ['Hammocks', 'Pet Cots', 'Pillow Beds', 'Radiator Beds']
# Animals & Pet Supplies > Pet Supplies > Cat Supplies > Cat Furniture > Cat Window Beds & Perches
That is the leaf the cookbook expects for its own listing, in 8 requests. The search is two functions and about twenty lines, where the cookbook’s search code runs to about 150 LOC. To be clear - I am not claiming that this is “better” in any way, I’m trying to show that we can achieve a similar outcome with a completely “model-native” approach. The potential advantage of a real RDM module in this situation is that it generalizes without writing bespoke “decision plumbing” for each new use case.
Signs of Life
It was very important to prove that RDMs work without any generative function at all, both as a sanity check on the implementation as well as to show how this approach is able to marry the concepts of RLM with these new System One models. The challenge with having “only decisions all the way down” is that you potentially need to enumerate the entire external state before running the RDM, because there is no easy way to generate new decisions to make.
My hunch is that for many applications such as browser use, video games, on-device decision models, and even computer use - this actually might be enough. You just enumerate every possible option in advance - then have the RDM sort through it all.
When you add LLMs to the mix, things start to come alive. I think this is part of why developers have found Jev so refreshing - these decision models are a powerful balancing force. The whole industry has been focused on expanding the divergent capabilities of models, TypeSafe has brought convergent capabilities into the spotlight.
The pair of generative model + decision model makes for a surprisingly lively automaton. One way of showing the difference here is by plotting the graph of calls to the decision model.
Here are the two handbook searches from earlier in this post. Every blue dot is one request to Jev. The grey tree behind it is everything the code laid out before the run started.

The shapes are regular because they are the shape of the data. Jev chooses a path, and nothing more.
TypeSafe’s own smart home assistant demo shows the general idea: let an LLM generate when the decisions aren’t enough. A yes/no question spots a request that asks for more than one thing, an LLM splits it into single commands, and each command goes back to the decision model. A request that is only conversation falls through to an LLM for a reply.
The graphs below came from our attempt to apply RDM to that smart home example, with a simulated house that pushes back: it refuses to lock a door that is open. Jev decides at every dot. The LLM only ever writes the next thing Jev is asked about, and each orange square is one of those.

Nothing was laid out in advance, and the same program made all four shapes. No line of it says to close the door before locking it. It also dies in ways nobody wrote: the last run is the same command, written and tried again at every level until the request limit.
As a final thought to wrap up this initial demonstration of combining language models and decision models, I want to make it clear that there are many possible permutations and combinations. I specifically am avoiding what I would consider the obvious one: giving an LLM agent a decision model as a tool. I’ve deliberately put the System One model at the center here, as I think that really highlights the potentially novel behavior.
Communicating with Maps
Now as soon as we add a second category of model, we need to think critically about the interface. The beauty of DSPy is that we can force the generative model to output a TypeSafe-compatible response, and this alone unlocks the “yin and yang” flow between the decision models and the generative models.
In practice, I found the direct connection between the System One models and large language models to be clunky. I went searching for a new primitive, something that would bridge the divide. Admittedly, this communication interface is the piece I am still wrestling with.
The working idea is a “map”. A decision model can only point, so something has to say what it is pointing at. A map is that: a set of names, each with a description for the model and a thing for the code.
point = dspy.Predict("request, house: dict -> next")
point.set_lm(TypeSafe("jev-latest"))
def run(lines, question, **shown):
"""`lines` is the map: {name: (description, thing)}. Returns the things that were done."""
done = []
while True:
options = tuple((name, description) for name, (description, _) in lines.items())
signature = point.signature.with_updated_fields("next", type_=Choice[options], desc=question)
name = point(signature=signature, **shown).next.value # the model points at a name
thing = lines[name][1] # the code looks up what the name stands for
if thing is None:
return done # nothing behind the name: stop
print("pointed at:", name)
thing()
done.append(thing)
That loop is the whole mechanism: show the names, the model points, the code does the thing behind the name, and the names are shown again. Here is a map for a house with a TV and a light.
house = {"tv": "off", "lights": "on"}
lines = {
"tv on": ("Turn the TV on.", lambda: house.update(tv="on")),
"tv off": ("Turn the TV off.", lambda: house.update(tv="off")),
"lights on": ("Turn the lights on.", lambda: house.update(lights="on")),
"lights off": ("Turn the lights off.", lambda: house.update(lights="off")),
"stop": ("Everything `inputs.request` asks for is already so in `inputs.house`.", None),
}
question = "What should happen next for `inputs.request`, with the house as `inputs.house` shows it?"
done = run(lines, question, request="Movie time.", house=house)
# pointed at: tv on
# pointed at: lights off
So far this is tool calling with a decision model. It gets more interesting when the map changes. What was just done can go back on the map as one more name. Its description says what it does, because the description is all the model sees of it.
lines["movie time"] = ("Do again what was done for 'Movie time.': TV on, lights off.", lambda: [thing() for thing in done])
house.update(tv="off", lights="on") # the next evening
run(lines, question, request="Let's watch a film.", house=house)
# pointed at: movie time
Two decisions became one, and a request in other words found it. Nothing matched strings: the decision model judged that “Let’s watch a film.” means the line kept for “Movie time.” This is the one map result that has held up so far. Over a few simulated days of requests to a toy house, a map that keeps what was pointed at got 87% of them right, against 66% for a map that keeps nothing, and it needs no LLM.
The part I am still working out is who writes that new line. Above, I wrote it in Python. The point of a map is that both kinds of model can change it, each in the way it is able to:
-
An LLM cannot write a function, but it can write a name, a description, and a list of names that are already on the map. That is a structured output.
-
A decision model cannot write at all, but it can point at a line that says “keep what was just done” or “forget that one”.
Both of these run today, as the same kind of edit to the map, and neither has yet done better than my one line of Python. Asked what to keep, the decision model kept everything and never chose to forget. Lines written by an LLM have not beaten plain pointing in any project we have tried. The one place an LLM has looked useful so far is smaller: writing the description of a line that was kept, so that a request in other words finds it.
A name can also stand for a group of more names, or for another map, so that pointing at it asks a question one level down. That is where the recursion comes back in: maps of maps of maps…
…but it is time to converge.
What is it good for?
1. “What is it good for?” 2. “Where are the benchmarks?” I can already hear the replies to this post.
1. I don’t know and 2. the benchmarks will come only after I nail down a v1 of the RDM module.
One of my favorite books about life is “Why Greatness Cannot Be Planned”, which happens to also be about machine learning (and specifically novelty search over objective optimization, which is highly relevant to the lineage of both RLMs and Jev). I am not sure where RDMs will be “better”, and I find Diogo’s distaste for benchmarks to be extremely refreshing. Right now I’m more interested in the novelty search: what might be possible with Recursive Decision Models that is impossible or impractical with other approaches? I also know that I need more “reality backpressure” to refine what the “essence of RDM” is.
What’s Next?
-
I don’t know exactly, but I will organize and share more code and examples soon
-
Definitely a dspy.RDM module that you can use
-
Maybe even a standalone RDM agent, if you ask me nicely
As usual, I can’t help but dive right into anything that comes out of Omar Khattab’s lab. “Machine Studying” is a very interesting blog post that attempts to define “expertise” and “intelligence” based on an agent’s ability to quickly learn new material from a corpus.
“In a way, Machine Learning asks how a system can improve from data when we have a precise objective to optimize. Machine Studying asks what an agent should do when it’s given a declarative corpus and no downstream task.”
I was curious to see what would happen with a modern harness (Codex CLI) and a top model (GPT 5.5 xhigh reasoning).
I was able to get a very high score right off the bat (76%, roughly 3x the highest reported in the original blog), which would suggest that either this combination of model and harness has high expertise on these tasks (already knows them well), or generally has fairly high intelligence for this category of task (ability to gain expertise).
Both of the examples in the current version of StudyBench involve testing an agent on its ability to use a specific version of a programming tool (DSPy & OpenClaw). The dates of the version in the test are from March and April of 2026, so I thought that it might be possible that these versions might already be in the training data of GPT 5.5, although the official page from OpenAI says the knowledge cutoff is Dec 1, 2025.
This led me to the idea that there are at least two different environments to study in: “the library” and “the lab”. In the library - you have books (i.e. documentation) and in the lab you have instruments (i.e. packages). My initial environments for DSPy and OpenClaw gave Codex access to both the documentation and the package. So I asked my buddy Codex (GPT 5.5 xhigh) to build out the remaining 3 environments. (“Lab only” turned out to be way trickier than anticipated as you have to block the package’s self-documentation.)
So now for each OpenClaw and DSPy StudyBench questions we end up with these four environments:
- Closed book: the agent only sees the question. It gets no docs, no source
tree, no package metadata, and no runtime to experiment with.
- Lab only: the agent gets no docs, no source tree, and no static study
corpus. It can only learn by running small probes against the pinned
tool/runtime through the allowed lab surface.
- Library only: the agent can read and search the official static material:
pinned source tree, docs, package metadata, and exposed corpus files. It cannot
run the package, import it, execute tests, or use a binary/runtime.
- Library + lab: the agent gets both surfaces. It can read the official
static material and also run experiments against the pinned environment.
| Treatment |
Library |
Lab |
DSPy |
OpenClaw |
Mean |
| Closed book |
no |
no |
51.70 |
6.35 |
33.56 |
| Lab only |
no |
yes |
78.07 |
9.70 |
50.72 |
| Library only |
yes |
no |
82.97 |
60.30 |
73.90 |
| Library + lab |
yes |
yes |
85.40 |
62.05 |
76.06 |
So these results suggest a few things:
- (Obviously) having access to both the documentation and the tool itself (library and lab) allows the agent to acquire the most expertise on DSPy and OpenClaw.
- GPT 5.5 appears to have much better built-in expertise (closed book score) about DSPy than OpenClaw. This makes sense, given that DSPy is many years old, and that OpenClaw is both very new and also has had three different names in its brief history.
- “Lab only” (tools but no documentation) led to modest improvements for both environments. “Library only” (docs and source code but no runnable package available) was slightly higher than “lab only” for DSPy, but dramatically higher for OpenClaw. I haven’t done an in-depth analysis of this, but again my suspicion is that it is related to the newness of OpenClaw, the quality of the documentation, and the structure of the StudyBench questions.
- Today’s SOTA coding agents may already be highly effective learners, when given access to documentation, source code, and runnable packages.
I did not make any attempts to control the budget (time, tokens, turns) or to vary the reasoning level (fixed xhigh) - so I cannot compute either the “expertise” or the “intelligence” in the framing of the Machine Studying post. These are obvious next steps, in addition to trying out some different models and harnesses.
My environments and results are available at: github.com/rawwerks/studybench-lab-and-library
Until next time…study hard!
Agent skills are a powerful and portable way to transform generally capable AI agents into specifically useful tools — even teammates. There are skills for almost anything you can imagine: doing your taxes, setting up Cloudflare, speaking like a caveman, generating algorithmic art…there’s even a skill dedicated to industrial brutalism.
Skills became so easy to make that the next challenge became finding them. This was quickly solved by various “skill registries” like Vercel’s skills.sh. OpenClaw was a huge inflection point for the skills boom — it made agent skills first-class in the product architecture and provided a high-permission substrate for millions of people to experiment with.
Of course, then safety became a problem. That’s a topic for a longer post, but in case you are curious, I developed the Skill Safety Data Sheet as an analogy to material safety data sheets — for evaluating the risks of specific agent skills.
So where do skills stand today?
If you are an information hoarder like me, your computer is also full of dozens or even hundreds of awesome agent skills — some that 10x devs shared on GitHub, and some that Claude Code made for you after you got tired of saying the same thing over and over and over again.
And if you are like me, your Claude Code and Codex have a terrible habit of finding random skills you don’t even remember installing. Worse, they never load the ones you just added — or at least not until you scold them.
There are two details of the skills implementation that make it very unreliable:
- LLMs are nondeterministic. They aren’t guaranteed to load the right skill at the right time.
- Progressive disclosure (implemented as a context window workaround) means the agent has to go looking for the information fresh each time. The clanker has to think to find the skill first, which is really inefficient.
In fact, Vercel recently showed that model-mediated skill activation lost to a simple index — a compressed 8KB AGENTS.md hit a 100% pass rate on their Next.js evals while a carefully crafted skill maxed out at 79%, and the skill was never invoked at all in 56% of cases. The winning approach still used a form of progressive disclosure — it just moved the routing layer into stable passive context. (In case you’re curious, I made dirpack as a general utility for creating indices of a fixed token budget for any directory.)
Engineer the skills!
So how can we take advantage of the power of agent skills without giving up our own agency to decide when and how they should be used? Engineer the skills!
Right now I’m having a lot of fun working on OpenProse — a “programming language” that is compiled inside of a coding agent. If this sounds sci-fi, it is…and OpenProse only really works with today’s top models.
The fun thing about OpenProse is that you can express very complex workflows in very simple markdown files. As an example, using the legacy v0 syntax purely for brevity:
input topic
loop until **editor approves** (max: 5):
session "research {{topic}}, address editor's prior notes"
session "draft from research, revise per prior notes"
session "review draft: approve as report or emit notes"
return report
This kind of logical statement is impossible to express in any other language. Prose is super fun!
But the problem I quickly ran into is that many of the prose programs I would want to run assume that my agents will use a specific skill. So I recently added the ability to deterministically declare agent skills inside prose programs — here’s the PR.
As a fun example of what’s possible with this new feature, I created auto-pocock: a headless prose program that incepts your favorite coding agent into running a deterministic sequence of Matt Pocock’s engineering skills (grill-with-docs → to-prd → to-issues → tdd → verify → commit), all from one input — a description of the feature you want built.
This combination of specific instructions (skills), deterministic processes (the prose contract), and nondeterministic magic (coding agents) is extremely versatile. By engineering skills with OpenProse, you can express complex multi-agent workflows, imbue each agent with detailed discipline, and hopefully get a much-needed break from the keyboard.
I personally do not care if my AI programs do their reasoning in latent space or code. I want results.
I am currently very intrigued by LongCoT, a new benchmark that is designed to push the limits of what is possible with today’s LLMs. Part of the original intention of the benchmark was to create something that would be both challenging for LLMs and less enmeshed with the details of the harness.
My recent results using DSPy.RLM have caused a bit of drama with some of the leaderboard owners, and the creation of a new tools-enhanced leaderboard. I understand the academic value in having “pure latent space” results without tools, but it just isn’t interesting to me…I want my agents to have tools.
So I set out to give the LLMs their desire path - the python tools that they tried to use in my prior benchmarking experiments.
This works surprisingly well with DSPy.RLM and Opus 4.7, which achieved a new SOTA on LongCoT-mini.
Opus 4.7 + DSPy.RLM → 377/500 (75.4%) on LongCoT-Mini — new top of the leaderboard, and a clear jump over the Sonnet 4.5 + DSPy.RLM 45.4% I posted in April.
| Mini |
Opus 4.7 + RLM |
Sonnet 4.5 + RLM |
Sonnet 4.5 vanilla |
| chess |
98/98 |
85/100 |
0/100 |
| logic |
101/101 |
106/110 |
0/110 |
| chemistry |
66/98 |
31/100 |
13/100 |
| cs |
71/97 |
4/100 |
0/100 |
| math |
41/77 |
6/95 |
0/95 |
| total (official /500) |
377/500 (75.4%) |
227/500 (45.4%) |
13/500 (2.6%) |
The new runs scored on the 471-question working set after the LongCoT team audited out 29 Mini questions as unsolvable. The official /500 totals here count those audited-out rows as wrong, matching the denominator used for the Sonnet baselines.
The paper splits the five domains into two classes: implicit (Logic, Chess, CS), where the dependency structure can be externalised to code, and explicit compositional (Math, Chemistry), where it can’t. The headline claim is that even with code execution enabled, RLM lifts the implicit class but leaves the compositional class near zero — direct quote: “explicit compositional domains (Math, Chemistry) remain at zero.” My April Sonnet run replicated that shape (math 6/95, hardest cs templates 0/75). Opus + RLM gets math up to 41/77 and cs to 71/97. Not zero.
Special thanks to Prime Intellect for sponsoring inference on this experiment — I promise I will publish my LongCoT environments soon. Even with that support I ran out of credits quickly, so I switched to OpenAI Codex CLI on the latest GPT-5.5 at “xhigh” reasoning, to put my $200/mo sub to work.
Codex CLI + gpt-5.5 “xhigh” → 398/500 (79.6%) on LongCoT-Mini — +21 rows over Opus.
| Mini |
Codex 5.5 xhigh |
Opus 4.7 + RLM |
| chess |
98/98 |
98/98 |
| logic |
100/101 |
101/101 |
| chemistry |
78/98 |
66/98 |
| cs |
67/97 |
71/97 |
| math |
55/77 |
41/77 |
| total (official /500) |
398/500 (79.6%) |
377/500 (75.4%) |
Two scaffolds, two models, similar bottom line. Codex’s persistent in-sandbox Python loop is grinding harder on chemistry and math, while DSPy.RLM holds a small cs edge. The LongCoT-Mini scoreboard is starting to feel more like a measure of how the agent’s tool loop is wired than of which frontier endpoint is behind it.
Mini is the easy slice. The full benchmark — medium + hard, ~2000 questions, where the dependency DAGs grow long enough to actually bite — is the real test of the paper’s compositional-walls claim. Over the weekend I let Codex loose on it; this time it took multiple Codex subscriptions to finish.
Codex CLI gpt-5.5 xhigh on full LongCoT: 1446/1995 (72.5%) — about 3× the top of the live longcot.ai Open Harness leaderboard, where the LongCoT team’s own GPT 5.2 + rlm run holds #1 at 25.12%. (My April Qwen 3.5 27B + DSPy.RLM run is #2 at 22.18%.)
| Full LongCoT · Codex 5.5 xhigh |
class |
medium |
hard |
total |
Open-Harness #1 (GPT 5.2 + rlm) |
| logic |
implicit |
187/195 (95.9%) |
165/199 (82.9%) |
89.3% |
68.3% |
| cs |
implicit |
150/150 (100.0%) |
210/250 (84.0%) |
90.0% |
26.7% |
| chess |
implicit |
92/150 (61.3%) |
200/250 (80.0%) |
73.0% |
30.6% |
| math |
compositional |
110/150 (73.3%) |
168/250 (67.2%) |
69.5% |
0.0% |
| chemistry |
compositional |
114/200 (57.0%) |
50/200 (25.0%) |
41.0% |
0.0% |
| total |
|
77.3% |
69.0% |
72.5% |
25.12% |
The compositional class — Math and Chemistry, the domains the paper said would “remain at zero” even with code execution — comes in at 69.5% and 41.0%. cs medium goes 150/150. The walls aren’t walls; with a stronger model and a more aggressive tool loop, they’re just programs.
With the full LongCoT now in hand, I think we can clearly state that the paper’s compositional walls don’t hold. Math and Chemistry — the domains the paper claimed would “remain at zero” even with code execution — come in at 69.5% and 41.0%. The wall isn’t compositional reasoning; it’s the harness used to measure it.
I think there is very clear evidence of something many people have been saying for the past ~18 months: it’s not (just) the model, it’s the harness. Additionally, it is very clear proof that “coding agents” are useful for long horizon tasks that don’t necessarily present themselves as coding problems.
Reasoning models were the first clear proof that language model capability can scale with test-time compute. Recursive language models (RLMs) ask what the correct abstraction for spending that compute is.
The insight behind RLMs is obvious in hindsight: it is the direct marriage of two important axes of model capability — reasoning and tool use. This is more radical than it first sounds. RLMs collapse reasoning and tool use into a single inference abstraction: the model treats its own prompt as an environment it can inspect, slice, and recursively query. Context itself becomes the object of computation.
This post is my attempt to explain why RLMs matter. I define what a RLM actually is, place it in the short history of reasoning and tool use, walk through the ~6 months of empirical results that have quietly turned “RLM” from a benchmark trick into the next reasoning paradigm, flag the honest limitations, and point at a few places to start building.
What is a RLM?
A Recursive Language Model, as introduced by Zhang, Kraska, and Khattab, is an inference paradigm in which a language model treats its input prompt as an environment rather than a fixed string. The root LM is given a REPL in which the prompt is bound to a variable it can inspect, slice, and partition programmatically. When it decides a region is worth a closer look, it issues a recursive subcall — to itself or another LM — over that slice and incorporates the result. Recursion bottoms out at the base model’s ordinary forward pass.
One consequence is that input size is no longer a hard ceiling on the computation. The paper reports RLMs processing inputs up to two orders of magnitude beyond the underlying model’s context window and outperforming vanilla frontier LLMs and common long-context scaffolds across four long-context tasks. Beyond long-context answering, recent results demonstrate that RLMs are a powerful paradigm for a wide variety of challenging tasks.
Reasoning & Tool Use — A Brief History
Reasoning and tool use are related, but they are not the same thing.
Reasoning is about how well a model can allocate inference-time compute to a problem: break it down, explore alternatives, verify intermediate steps, backtrack, and choose a better answer. Early reasoning gains came from methods like chain-of-thought, self-consistency, and later tree-search-style prompting. Those methods improve how the model thinks even when it never touches the outside world.
Tool use is about whether a model can decide to call an external function, search engine, calculator, browser, code runner, or UI action; pass the right arguments; interpret the result; and continue. That is partly a reasoning problem, but it is also an interface and reliability problem: schemas, argument formatting, retries, stop conditions, state tracking, and error recovery. Toolformer made this distinction especially clear by treating tool use as something a model could learn during generation.
Historically, the timeline looks roughly like this:
2022: reasoning first, mostly without tools.
Chain-of-thought prompting showed that asking models to generate intermediate reasoning steps could dramatically improve multi-step reasoning. Self-consistency pushed this further by sampling multiple reasoning paths and selecting the most consistent answer. The key lesson was that a large share of “reasoning” gains could come from spending more inference-time compute on the same prompt, not just from adding more knowledge.
Late 2022: the first real bridge between reasoning and acting.
ReAct was the key milestone. It framed the model as alternating between reasoning traces and external actions such as retrieval or environment interaction. This was the moment the field started to see tool use not as a one-off API call, but as a loop in which reasoning selects actions and tool outputs reshape the next reasoning step.
2023: tool use becomes an API discipline, not just a prompting trick.
Toolformer argued that models could learn when to call tools, which tools to call, and how to incorporate the results. Around the same time, vendors began standardizing function-calling interfaces. OpenAI’s June 2023 function calling release was a major product milestone because it made structured tool invocation reliable enough for developers to build on. This improved tool-use reliability faster than it improved deep reasoning.
2023 also deepened the separation between reasoning and tool use.
Tree of Thoughts made it even clearer that inference-time reasoning could improve through internal search alone. It let models explore multiple candidate thought branches, look ahead, and backtrack. That is search over reasoning traces. It can be paired with tools, but it does not require them.
2024: reasoning models become their own product category.
OpenAI’s o1 launch was the clearest signal. The company described o1 as a model family designed to “spend more time thinking before they respond,” and the initial API announcement explicitly noted that features like function calling were not yet included. That was strong evidence that, product-wise, reasoning and tool use were still separable.
2024 is also when agentic tool use got much more serious.
Anthropic’s Claude 3.5 Sonnet emphasized stronger tool use for coding and agentic tasks, and later in 2024 Anthropic introduced computer use: a model interacting with a real computer via screenshots, mouse, and keyboard. This is a good example of the two axes starting to merge into one agentic stack.
Late 2024 into 2025: vendors start presenting tool use as native, but still distinct from thinking.
Google’s Gemini 2.0 messaging explicitly framed the model family around the “agentic era” and native tool use, while keeping “thinking” as a distinct capability for harder multi-step planning. That split mirrors the real architecture: one layer governs deliberation, another governs interaction with external affordances.
RLMs are the abstraction where that split finally collapses. The past ~6 months of results are what make the case concrete.
Recent RLM Results
The arc of RLM results moves through three successive failure modes of the single forward pass: long context, then memory, then long reasoning. Each has been demonstrated by its own benchmark — Oolong, LongMemEval, and LongCoT respectively — and RLM-style systems have posted leading numbers on all three. Just as importantly, the follow-up work is already splitting into two camps: work that strengthens the original RLM implementation, and work that argues the deeper win is broader externalized program search rather than recursion alone.
Part of what makes RLMs challenging to appreciate is that frankly there aren’t very many benchmarks that really showcase the differences. In particular, I don’t view Oolong and LongMemEval as having much correlation to performance on real world agentic tasks. LongCoT is much more exciting to me, but it is brand new and only time will tell how it holds up.
2024: the memory target appears.
LongMemEval defines the benchmark for long-term interactive memory: 500 questions over sustained chat histories spanning extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. It matters here because it gives RLM-style systems a way to test whether recursive/tool-mediated processing can function as a memory system, not just a long-context hack.
October 2025: the original public RLM write-up lands.
In Recursive Language Models, Alex Zhang introduces the core idea: treat the prompt as an external environment, manipulate it through a REPL, and recursively subquery models over slices of context. The post reports an unusually strong early result profile: a GPT-5-mini RLM beats GPT-5 by more than 2× on an Oolong split while being cheaper per query on average, beats ReAct + test-time indexing/retrieval on a BrowseComp-Plus-derived long-context research task, and does not visibly degrade even at 10M+ input tokens.
November 2025: Oolong raises the bar for long-context reasoning.
Oolong is important because it measures something harder than needle-in-a-haystack retrieval: models have to analyze many local chunks and then aggregate them into a global answer. At release, GPT-5, Claude-Sonnet-4, and Gemini-2.5-Pro all score under 50% on both splits at 128K, making Oolong the clearest early benchmark for the kind of “context as workspace” reasoning RLM is trying to solve.
December 2025: the arXiv paper formalizes RLM.
The Recursive Language Models paper turns the blog’s intuition into a general inference paradigm: prompts are externalized, the LM programmatically inspects and partitions them, and recursive subcalls become part of test-time compute. The headline results are strong: RLMs process inputs up to two orders of magnitude beyond model context windows, outperform vanilla frontier LLMs and common long-context scaffolds across four long-context tasks at comparable cost, and a fine-tuned RLM-Qwen3-8B improves 28.3% on average over its base model.
February 2026: RLM starts posting real “memory” numbers.
In Recursive Language Models as Memory Systems, I reported early LongMemEval results with DSPy.RLM: 87.2% for a baseline Gemini 3 Flash setup, 89.2% with tools + a delegation prompt, and 89.8% with an observational-memory-style structured scaffold. That was a public Top-5-ish result at the time, below Mastra’s 94.87% but already strong evidence that RLM can act as a competitive memory system without a classical retrieval stack. In ypi: a recursive coding agent, I show an earlier tool-use REPL path scoring 77.6% on LongMemEval — a useful datapoint because it shows the gradient from “tool-using agent” to “true recursive scaffold” inside the same implementation lineage.
March 2026: follow-up papers clarify both the strengths and the limits.
Think, But Don’t Overthink reproduces RLM and finds that depth-1 recursion helps on Oolong, but deeper recursion can “overthink,” hurting accuracy and exploding runtime and token cost. Recursive Language Models Meet Uncertainty pushes a sharper critique: recursion itself is not the whole secret, and uncertainty-aware self-reflective program search can improve up to 22% over RLM under the same time budget. Then Coding Agents are Effective Long-Context Processors generalizes the broader thesis: off-the-shelf coding agents outperform published SOTA by 17.3% on average, and on Oolong-Synthetic / Oolong-Real their reported scores (71.75 / 33.73) exceed the paper’s RLM baselines (64.38 / 23.07). That does not really refute RLM; it suggests RLM was the first clearly articulated expression of a larger family of executable, tool-mediated long-context reasoning systems.
April 2026: the theory catches up to the results.
In The Mismanaged Geniuses Hypothesis, Zhang reframes the whole arc: RLM is not just a benchmark trick for long prompts, but a more expressive scaffold for plans written through code execution, recursive subcalls, and tools-as-functions. That is a useful conceptual update because it connects the empirical results back to the bigger claim: reasoning performance is starting to look less like a property of a single forward pass and more like a property of how well a model can manage executable external computation.
The empirical case moves just as quickly.
April 2026: the benchmark story shifts from long context to long reasoning.
LongCoT introduces 2,500 expert-designed problems for long-horizon chain-of-thought reasoning. At release, the best published models are still under 10% accuracy (GPT-5.2 at 9.8%, Gemini 3 Pro at 6.1%), which makes it an ideal test for whether recursive scaffolds are merely “good at reading long context” or whether they genuinely unlock reasoning depth.
April 2026: RLM immediately breaks LongCoT open.
In LongCoT — A benchmark worthy of a RLM’s attention, I showed Claude Sonnet 4.5 + DSPy.RLM reaching 45.4% on LongCoT-Mini versus 2.6% for the same model without recursion/tools. Then in RLMs are SOTA on LongCoT, I show the scaffold doing almost all of the lifting for small open models: Qwen3-8B jumps from 0/507 to 33/507 (6.5%) on LongCoT-Mini; Qwen3.5-9B + DSPy.RLM reaches 15.69% on full LongCoT, about 1.6× GPT-5.2 on the same slice; and Qwen3.5-27B + DSPy.RLM reaches 22.18%, more than 2× GPT-5.2. If these numbers hold up, they are some of the clearest evidence yet that recursive scaffolds can manufacture reasoning performance that is not visible in the base model alone.
The arc is now hard to ignore. Oolong gives the long-context failure mode. LongMemEval gives the memory version. LongCoT gives the long-reasoning version. Across all three, the recurring pattern is the same: when the task requires navigating, decomposing, and aggregating information over a structure that is too large or too entangled for one passive forward pass, recursive tool-mediated processing starts to look less like an implementation trick and more like the next reasoning paradigm.
Challenges with RLMs
A new paradigm is not a clean paradigm. Reasoning models were scoffed at for being too expensive. Early tool calling reliability was horrible. Even some leading reasoning models today are pretty bad at function calling.
RLMs have their challenges. My earliest contributions to the standalone RLM package and the DSPy.RLM implementation were purely practical: budgets, timeouts, managing the recursion depth.
Recursion sounds cool but isn’t always a good thing. Remember those viruses that would make your browser open a million popups until your computer crashed?
Recursion can be scary.

The most obvious limitations right now are cost and time. RLMs are expensive. They can take a long time. Worse, in the naive implementation that time is unpredictable and unbounded, because the model is deciding for itself how to decompose the problem.
Cost and time will be solved. Use smaller or faster models for each sub-call, and balance the agent-native “self-similar” decomposition with deterministic control of the graph topology and timeline.
The harder challenge, or at least the challenge that is more interesting to me personally, is how to get the language models to “act recursively”.
Obviously the concepts of recursion are in the pre-training data. Clearly reasoning and parallel tool calling are behaviors that the post-training incentivizes. Sub-agents are arguably a close behavioral analog to RLMs. And yet anyone who has worked with RLMs will tell you that the models generally suck at behaving recursively. It is not in their nature to decompose their prompt into sub-queries for many other instances of themselves to help solve them.
What’s next?
Well one obvious next step is to explicitly post-train the models in a RLM harness. Alex Zhang et al. are actively working in this area: MIT OASYS on HuggingFace (see e.g. mit-oasys/rlm-qwen3-8b-v0.1).
But what is the reward function for “optimal recursion”? I suspect this is a multi-billion-dollar question.
The most surprising result to me from my last few days of experimenting was how well very small models can do in RLM harnesses. These models are small enough to run on consumer devices, which potentially means that they offer an opportunity to upset the current “balance of power” between the GPU-rich and GPU-poor.
Yes, more money means you can run more. The best GPUs will always be faster. A RLM of Opus is smarter than a RLM of Llama 3. But I cannot help but feel excited and empowered to believe that an individual or consortium running many instances of small models on affordable/legacy/local compute infrastructure can now access model capabilities that are on par with or exceeding those of the most expensive LLMs from the frontier labs. If that is even directionally right, the frontier stops being a place only the largest labs can reach.
Getting Started with RLMs
Here are just a few of the many ways to get started with RLMs:
- alexzhang13/rlm — the reference implementation from Alex Zhang and the RLM paper authors; the cleanest place to read the core recursion loop.
- dspy.RLM — the DSPy integration, which exposes RLM as a composable module inside larger DSPy programs and is what I’ve been using for most of my own experiments.
- ax-llm/ax — a TypeScript DSPy-style framework with first-class RLM support via
AxAgent-driven recursive decomposition, bounded sub-queries, and a persistent JS runtime.
- rawwerks/rlm-cli — my CLI wrapper around
rlm with directory-as-context, JSON-first output, and self-documenting commands, for running RLMs against local repos and folders.
- rawwerks/ypi — my recursive coding agent built on Pi: one
rlm_query tool, one rlm_map fanout helper, and per-child jj workspaces for isolated recursive execution.
P.S.
I almost forgot: fractals.

A few days ago I showed Sonnet 4.5 + dspy.RLM hitting 45.4% on LongCoT-Mini. Exciting results, but a bit pricey for my taste.
So I set out to see what might be possible with some very small models.
First, I wanted to run a 3x2 comparison matrix of doing LLM vs. RLM vs. DSPy.RLM for both Qwen 3 8B and the MIT OASYS RLM finetune.
I will need to save the full analysis for another day (I wasn’t really able to get the finetuned model working), but the meaningful result is that on LongCoT mini, DSPy.RLM can take Qwen 3 8B Instruct from literally 0/507 correct to 33/507 (6.5%). This would be #7 on the leaderboard, from an 8B model!
So my immediate next thought was: “what about Qwen 3.5 9B”? This hit 17.2% on LongCoT mini (3rd place), and was so cheap that I decided to run the full benchmark (all 2500 questions)! (Now using Together AI via OpenRouter, I don’t think their endpoint is quantized but I’m not 100% sure.)
Surprisingly, DSPy.RLM with Qwen 3.5 9B is comfortably SOTA on the full LongCoT at 15.69%, outdoing GPT 5.2 by ~1.6x.
Now I was having too much fun, so I had to run Qwen 3.5 27B (this time via Alibaba Cloud via OpenRouter)…and unsurprisingly, a new LongCoT full king is crowned at 22.18%.
I’m really excited to finally have a meaningful benchmark that can clearly demonstrate the power of RLMs. This clearly deserves a much longer writeup, which I hope to post soon! (And I’m now running Qwen 3.6 35B-A3B at the suggestion of many folks on X.)
After the LongMemEval experiments in February, I’ve been hungry to find a better benchmark that will actually showcase the power of recursive language models (RLMs) on useful tasks. LongCoT is exactly that: a benchmark built to stress-test long-horizon reasoning.
As soon as I saw the benchmark, I aimed DSPy.RLM at it. (Even before reading the paper.)
The setup
- Model:
claude-sonnet-4-5 for both conditions. Same max_tokens=64000, same judge models, same prompts.
- RLM: stock
dspy.RLM 3.1.3, max_iterations=50, default Pyodide REPL, sub_lm=lm.
- Vanilla: raw Anthropic SDK, single user message, no tools. Leaderboard shape.
- Dataset: LongCoT-Mini, all 500 questions (easy slices across logic / cs / chemistry / chess / math).
The entire RLM surface area is one dspy Signature:
class LongCoTSolve(dspy.Signature):
"""Solve a LongCoT problem.
The `prompt` already contains the full problem statement and the answer
format requirement (always ends with `solution = ...`). Reason through
the problem with the available REPL, then return the final response —
which MUST contain the literal `solution = ...` line as instructed.
"""
prompt: str = dspy.InputField(desc="Full LongCoT problem prompt with answer-format instructions")
response: str = dspy.OutputField(desc="Full final response containing the required `solution = ...` line")
The headline
|
Vanilla |
RLM |
| Correct |
13 / 500 |
227 / 500 |
| Accuracy |
2.6% |
45.4% |
| Captured cost |
$31 |
$621 |
On the full 500-row overlap: 219 wrong→right flips, 5 right→wrong, 268 both-wrong, 8 both-right. The vanilla 2.6% lines up with the published Sonnet 4.5 Mini number, so the control is calibrated, not sandbagged.
Per-task
| Task |
RLM |
Vanilla |
| Dungeon · Packaging · Hanoi · Wizards · TrapezoidCounting · Sudoku |
15/15 each (💯) |
0/15 each |
| BlocksWorld |
9/10 |
0/10 |
| Sokoban |
7/10 |
0/10 |
| Chess |
85/100 |
0/100 |
| Chemistry |
31/100 |
13/100 |
| cs / DistMem |
4/25 |
0/25 |
| cs / MaxFlow-MinCut + Hindley-Milner |
0/75 |
0/75 |
| math |
6/95 |
0/95 |
The pattern is coherent: RLM crushes anything whose dependency structure externalises cleanly to code. The orchestrator writes a short Python program, the REPL runs it, the answer comes out. Logic puzzles, Hanoi, Sudoku, chess with Pyodide’s chess module — all 💯 or near it.
The walls are the opposite picture. Hindley-Milner and MaxFlow-MinCut go 0/75 because the orchestrator can’t find a decomposition where subproblems can be usefully farmed out — exactly the “graph-structured dependencies” failure the paper calls out.
And math? The paper’s wall holds at least for now, for Sonnet 4.5. 6/95 on Mini isn’t zero, but it’s terrible. Sonnet 4.5 × dspy.RLM replicates the paper’s math result on a different model and split.
What I think this means for the paper
The paper’s RLM discussion is genuinely thin — one paragraph, one figure, no dedicated table. With that as the bar, cross-model replication is useful:
- Logic, chess, CS wins: replicate and amplify. Same shape on a different frontier model.
- Math stays at zero: maybe? Model swap doesn’t rescue it. But it’s also from a baseline of 6 so you can argue it’s either modest or infinite improvement.
- Chemistry lifts modestly (13 → 31), which is the only spot where I’d push back on the paper’s phrasing — but I’m both a chemist and RLM addict.
P.S. - RLMs are expensive, and supposedly the 500-problem mini version is the easy subset of the full 2500-problem set. So…who wants to fund the Opus 4.7 run?
Over the last 2 days, we’ve stumbled upon a really powerful coding agent interaction pattern: git notes as an underground information network.
Git notes are both ubiquitous (part of git) and “invisible” (GitHub chose not to display them). This presents a very interesting communication channel for agents, who can now include rich details and discussions about the code without cluttering up the “visible” layer of the repo.
mycelium is my tool to make these interactions easier.
# agent arrives, reads what's known about a file
mycelium.sh context src/auth.ts
# agent works...
# agent leaves a note explaining what it did
mycelium.sh note HEAD -k context -m "Refactored retry logic. See warning on auth.ts."
Agents read notes on arrival. They leave notes on departure. The network grows.
The CLI makes it easy to link notes and git refs together — files, commits, directories, even edges between notes. Notes can have kinds (decision, warning, summary, context) and edges (depends-on, explains, warns-about) but the vocabulary is open. The tool tries to stay unopinionated about the actual workflow. From the SKILL.md: “That’s the whole contract. How you work, what you build, how you talk to your user — that’s your business. Mycelium just asks you to read the breadcrumbs and leave new ones.”
mycelium is meant to be agent-native — load the SKILL.md into your agent framework and it teaches the convention. But it’s just git & bash, so it works with any agent in any git repo.
curl -fsSL https://raw.githubusercontent.com/openprose/mycelium/main/install.sh | bash
Still wrapping my head around the consequences of this, and very curious to hear your thoughts.
P.S. — this is the foundation of some very cool tools I’m collaborating with OpenProse on.
I built ypi — a recursive coding agent. It’s Pi that can call itself.
The name comes from the Y combinator in lambda calculus — the fixed-point combinator that enables recursion. (“rpi” has other connotations.)
The idea was inspired by Recursive Language Models (RLM), which showed that an LLM with a code REPL and a llm_query() function can recursively decompose problems, analyze massive contexts, and write code — all through self-delegation.
The idea
Pi already has a bash REPL. I added one function — rlm_query — and a system prompt that teaches Pi to use it recursively. Each child gets its own jj workspace for file isolation. That’s the whole trick.
┌──────────────────────────────────────────┐
│ ypi (depth 0) │
│ Tools: bash, rlm_query │
│ Workspace: default │
│ │
│ > grep -n "bug" src/*.py │
│ > sed -n '50,80p' src/app.py \ │
│ | rlm_query "Fix this bug" │
│ │ │
│ ▼ │
│ ┌────────────────────────────┐ │
│ │ ypi (depth 1) │ │
│ │ Workspace: jj isolated │ │
│ │ Edits files safely │ │
│ │ Returns: patch on stdout │ │
│ └────────────────────────────┘ │
│ │
│ > jj squash --from <child-change> │
│ # absorb the fix into our working copy │
└──────────────────────────────────────────┘
The recursion works like this: rlm_query spawns a child Pi process with the same system prompt and tools. The child can call rlm_query too:
Depth 0 (root) → full Pi with bash + rlm_query
Depth 1 (child) → full Pi with bash + rlm_query, own jj workspace
Depth 2 (leaf) → full Pi with bash, but no rlm_query (max depth)
Each recursive child gets its own jj workspace, so the parent’s working copy stays untouched. You review child work with jj diff, absorb it with jj squash --from.
How it works
The architecture maps directly to the Python RLM library:
| Piece |
Python RLM |
ypi |
| System prompt |
RLM_SYSTEM_PROMPT |
SYSTEM_PROMPT.md |
| Context / REPL |
Python context variable |
$CONTEXT file + bash |
| Sub-call function |
llm_query("prompt") |
rlm_query "prompt" |
The key insight: Pi’s bash tool is the REPL. rlm_query is llm_query(). No bridge needed.
Guardrails
Recursive agents without guardrails will burn through your API budget. ypi has several:
| Feature |
Env var |
What it does |
| Budget |
RLM_BUDGET=0.50 |
Max dollar spend for entire recursive tree |
| Timeout |
RLM_TIMEOUT=60 |
Wall-clock limit for entire recursive tree |
| Call limit |
RLM_MAX_CALLS=20 |
Max total rlm_query invocations |
| Model routing |
RLM_CHILD_MODEL=haiku |
Use cheaper model for sub-calls |
| Depth limit |
RLM_MAX_DEPTH=3 |
How deep recursion can go |
| Tracing |
PI_TRACE_FILE=/tmp/trace.log |
Log all calls with timing + cost |
The agent can check its own spend at any time:
rlm_cost # "$0.042381"
rlm_cost --json # {"cost": 0.042381, "tokens": 12450, "calls": 3}
The path here
ypi went through four approaches before landing on the current design:
- Tool-use REPL — Pi’s
completeWithTools(), ReAct loop. Got 77.6% on LongMemEval.
- Python bridge — HTTP server between Pi and Python RLM. Too complex.
- Pi extension — Custom provider with search tools. Not true recursion.
- Bash RLM —
rlm_query + SYSTEM_PROMPT.md. True recursion via bash. This is the one that stuck.
Try it
curl -fsSL https://raw.githubusercontent.com/rawwerks/ypi/master/install.sh | bash
Or via npm/bun:
npm install -g ypi
ypi "What does this repo do?"
Or without installing:
bunx ypi "Refactor the error handling in this repo"
Code is at github.com/rawwerks/ypi. It’s built on Pi and inspired by RLM.
My morning’s notes from yesterday:

As I was waiting for Claude Code to help me with my goal of modifying DSPy to be able to “RLM everything”, I came across this result from Mastra.AI which describes a SOTA result on LongMemEval using an “observational memory” pre-processing approach.
As you can see, my thought was “maybe RLM will blow this out of the water?” I wasn’t able to find a public result of LongMemEval using Recursive Language Models, so I decided to explore it myself.
The initial results with Gemini 3 Flash Preview and the standalone RLM package weren’t great, but in the past I had noticed that Flash struggled to grok the RLM concept. Gemini 3 Pro fared much better.
Surprisingly - the additional structure enforced by DSPy.RLM was a huge boost, enabling Gemini 3 Flash to match Pro with the regular RLM package.
Most of the rest of the day’s experiments were less successful. I was able to eke out a few more points by attempting to re-create Mastra’s “Observational Memory” as a Pydantic type enforced by DSPy, but unfortunately a few hundred dollars worth of GEPA optimizations didn’t bear any additional fruit.
Surprisingly - with the full structure of DSPy.RLM and the structured observation, Gemini 3 Pro is not actually any better on this benchmark.
Here’s a summary of our experiments:
| # |
Experiment |
Model |
Score |
Cost/q |
Notes |
| 1 |
dspy.RLM baseline |
Gemini 3 Flash |
87.2% |
~$0.01 |
Huge boost over standalone RLM with Flash (58%) |
| 2 |
+ session tools (naive) |
Gemini 3 Flash |
87.3% |
$0.032 |
Context rot: +31 flips, -32 regressions = net zero |
| 3 |
+ tools + delegation prompt |
Gemini 3 Flash |
89.2% |
$0.031 |
“Don’t read sessions yourself, delegate” |
| 4 |
+ observational memory (Pydantic) |
Gemini 3 Flash |
89.8% |
$0.035 |
Our best. Typed observations force structured reasoning |
| 5 |
GEPA prompt optimization |
Gemini 3 Flash |
87.8% |
$0.042 |
Regressed. ~$400 spent. Overfits to small val sets |
| 6 |
Observational memory |
Gemini 3 Pro |
~89.6% |
~$0.20 |
Pro ≈ Flash with this scaffold |
And here’s how that stacks up on the LongMemEval leaderboard:
Not bad for a day’s work, we were able to demonstrate a “Top-5” LongMemEval result with very minimal modifications to dspy.RLM, just some helper functions to process the “multi-chat” sessions.
I think this demonstrates a few exciting things:
- RLMs can be very powerful memory systems without any pre-processing.
- The structured output enforced by the DSPy.RLM implementation is helpful for keeping (at least these Gemini models) “on the rails” vs. the more freeform standalone RLM package.
- Very fast and inexpensive models can achieve near-SOTA results inside the RLM scaffolding, and more speculatively…
- …perhaps RLM as a test-time scaling method is “orthogonal” to model size, in the same way that reasoning models with built-in CoT were able to eke out gains separately from model parameter count.
P.S. — Several improvements to DSPy.RLM were developed during this work and submitted upstream: stanfordnlp/dspy#9295