Recursive Language Models (RLMs) have proven a powerful harness architecture for Large Language Models, enabling dynamic agent workflows (Anthropic), extremely long context processing (MIT), SOTA results on “long chain-of-thought” benchmarks (RAW.works), and even best-in-class performance on benchmarks such as ARC-AGI 3 (Prime Intellect).
“Recursive Decision Models” (RDMs) is an exploration of how the principles of RLMs can be applied to “decision models” (aka “System One models” aka “classifiers”) such as Jev.
Briefly, a note on the naming: I don’t particularly care what term people choose for the “genre of Jev”. Diogo officially prefers “System One models”1, and has explicitly argued against the alternative: “Some people are trying to call them decision models… I wouldn’t do that, because I think there’s other types that are machine-native that are not decisions.” Nevertheless, the zeitgeist seems to be trending towards “decision models” and I’m not willing to wait any longer for people to make up their mind.
Team! What is the name of the "genre" of Jev?
— Raymond Weitekamp (@raw_works) September 21, 2026
Please let's decide now so we don't need to re-name all of our projects later...
In developing this concept of RDMs, I’ve been working from three “non-negotiables” of how the system must behave:
-
The RDM system must externalize context and be able to programmatically decompose slices of that context.
-
The decision model must have the choice to call another decision model, up to some maximum depth of recursion.
-
The RDM system must be able to work with and without the presence of any generative model or function. In other words, it must stand alone as “pure decision models” and also compose with LLMs and agents.
I initially started exploring some of these ideas in a repo that I dubbed “one-system”, which started as a “LiteLLM for Jev” and then I quickly realized could support recursion by composing requests through the gateway. Lucky for me, the DSPy team (in particular my esteemed peers at cmpnd.ai) has rapidly embraced Jev. Now that the TypeSafe System One API is supported in DSPy, my task of conveying the essence of Recursive Decision Models is greatly simplified.
A Simple RDM
Let’s start with a dspy.Module that recursively makes decisions up to its max depth.
import dspy
class RDM(dspy.Module):
def __init__(self, signature, tools, max_depth=3):
super().__init__()
self.predict = dspy.Predict(signature)
self.tools = {tool.__name__: tool for tool in tools}
self.max_depth, self.depth = max_depth, 0
def forward(self, **inputs):
answers = self.predict(**inputs) # one request to the decision model
observations = {}
if self.depth < self.max_depth:
self.depth += 1 # anything a tool calls is one level deeper
for name, tool in self.tools.items():
if answers[name].value: # a yes runs the tool of the same name
observations[name] = tool(**inputs)
self.depth -= 1
return dspy.Prediction(observations=observations, **answers)
As a simple example, we write a signature that asks if the answer to a question is in a chunk of text, and a function to divide and conquer by recursively halving the content.
from dspy.experimental import Noul, TypeSafe
class Search(dspy.Signature):
"""Search a text for the answer to a question."""
text: str = dspy.InputField()
question: str = dspy.InputField()
look: Noul = dspy.OutputField(desc="Is the answer to `inputs.question` stated in `inputs.text`?")
def look(text, question):
lines = text.splitlines()
if len(lines) <= 1:
return lines # down to one line, and the model said yes to it
middle = len(lines) // 2
found = []
for half in lines[:middle], lines[middle:]:
found += rdm(text="\n".join(half), question=question).observations.get("look", [])
return found
To finish the specific example, we can search a handbook recursively to find the specific line that answers a query.
handbook = open("handbook.md").read() # 71 lines
rdm = RDM(Search, tools=[look], max_depth=10)
rdm.set_lm(TypeSafe("jev-latest")) # reads TYPESAFE_API_KEY
result = rdm(text=handbook, question="What is the nightly hotel limit in London?")
print(result.observations["look"])
# ['The nightly limit for accommodation is 180 euros in most cities and 260 euros in London, ...']
# 15 requests, each a yes or a no
While this is a terrible way to actually solve this problem in the real world, it helped us demonstrate how we can use recursion to have decision models programmatically process a potentially very large context.
Self-similarity and external state
Readers familiar with RLM will probably be offended by the last example, so let’s now build out an analog that clearly demonstrates “the shape of RLM”. Specifically, that means we need to:
-
Have RDM call RDM all the way out to a leaf
-
Explicitly externalize the context
-
Keep it LLM-free to prove the point that we don’t need generation for this to work
The halving search sent the whole handbook to the model in its first request. This time the handbook stays in Python as state, and no request ever shows it whole. Plain code turns it into an outline: each heading, with the headings and passages directly under it.
def outline(markdown, title):
"""{heading: [the headings and passages directly under it]}, made by plain code."""
under, path = {title: []}, [title]
for block in markdown.strip().split("\n\n"):
depth = len(block) - len(block.lstrip("#"))
if depth:
block = block.lstrip("# ")
path = path[:depth] + [block]
under[block] = []
under[path[-2] if depth else path[-1]].append(block)
return under
state = outline(handbook, "Handbook") # stays in Python: no request ever shows it whole
def headings(entry):
"""Every heading below an entry: all that a request shows of what is under it."""
return [below for each in state.get(entry, []) if each in state for below in [each, *headings(each)]]
A request shows one entry and the headings below it, never the text under them. The same question is asked at every node, whether it is the whole handbook, a section, or a single passage.
class Navigate(dspy.Signature):
"""Search a handbook for the answer to a question, one entry at a time."""
entry: str = dspy.InputField(desc="A heading of the handbook, or a passage under one.")
contains: list[str] = dspy.InputField(desc="The headings under `inputs.entry`. Empty for a passage.")
question: str = dspy.InputField()
explore: Noul = dspy.OutputField(
desc="Does `inputs.entry` state the answer to `inputs.question`, or does `inputs.entry` or any "
"heading listed in `inputs.contains` name the topic that `inputs.question` asks about?"
)
def explore(entry, contains, question):
if entry not in state:
return [entry] # a leaf: a passage, and the model said yes to it
print("explored:", entry)
found = []
for each in state[entry]:
found += rdm(entry=each, contains=headings(each), question=question).observations.get("explore", [])
return found
A yes runs explore, which calls the same module with the same question on each entry under this one, all the way out to a passage. The RDM class is the one from the first section, unchanged. There is no LLM anywhere: Jev is the only model, and all it ever says is yes or no.
rdm = RDM(Navigate, tools=[explore], max_depth=10)
rdm.set_lm(TypeSafe("jev-latest"))
assert dspy.settings.lm is None # no LLM is configured anywhere
result = rdm(entry="Handbook", contains=headings("Handbook"), question="What is the nightly hotel limit in London?")
print(result.observations["explore"])
# explored: Handbook
# explored: Expenses
# explored: Travel
# explored: Hotels
# ['The nightly limit for accommodation is 180 euros in most cities and 260 euros in London, ...']
# 15 requests again, but the largest shows 272 characters of a 2,555-character handbook
Hierarchical search, without the beam
Let’s now refine the concept of Recursive Decision Models by remixing some of the Jev best practices. TypeSafe’s hierarchical classification cookbook walks a taxonomy with one Choice per node, and keeps the three best paths with a beam search that scores each path by the probabilities along it. The same walk falls directly out of recursion, so we can use the same recursive pattern to solve the problem without writing the beam search code explicitly. The external state here is the cookbook’s own Shopify product taxonomy: over twelve thousand categories that no request ever shows whole.
from urllib.request import urlopen
URL = "https://raw.githubusercontent.com/Shopify/product-taxonomy/v2026-02/dist/en/categories.txt"
taxonomy = {"All products": []} # {category: [the categories directly under it]}
for line in urlopen(URL).read().decode().splitlines()[3:]:
path = line.split(" : ")[1] # "Animals & Pet Supplies > Pet Supplies > Cat Supplies"
taxonomy[path.rpartition(" > ")[0] or "All products"].append(path)
taxonomy[path] = []
At each category the model is asked which of the categories under it fits best. The function follows every one that is at least half as likely as the likeliest, by calling itself, and when more than one leaf comes back it asks the same question again among those. There is no beam width and no path score: the model’s own odds decide where the search branches.
from dspy.experimental import Choice
predict = dspy.Predict("listing -> category")
predict.set_lm(TypeSafe("jev-latest"))
def choose(listing, options):
"""One request: how likely each category is to be the best match for the listing."""
if len(options) == 1:
return {options[0]: 1.0}
signature = predict.signature.with_updated_fields(
"category",
type_=Choice[tuple((option, None) for option in options)],
desc="Which category best matches the product in `inputs.listing`?",
)
return predict(signature=signature, listing=listing).category.probabilities
def classify(listing, category="All products"):
under = taxonomy[category]
if not under:
return category # a leaf
odds = choose(listing, under)
likely = [each for each in under if odds[each] >= max(odds.values()) / 2]
print(category.rpartition(" > ")[2], "->", [each.rpartition(" > ")[2] for each in likely])
found = [classify(listing, each) for each in likely] # the same function, one level down
odds = choose(listing, found) # more than one came back: ask again, among those
return max(found, key=odds.get)
listing = (
"Furniture listing: a wall-mounted window shelf bed. This padded floating shelf uses "
"suction cups and a washable cushion as a sunny perch for one cat."
)
print(classify(listing))
# All products -> ['Animals & Pet Supplies']
# Animals & Pet Supplies -> ['Pet Supplies']
# Pet Supplies -> ['Cat Supplies', 'Pet Beds']
# Cat Supplies -> ['Cat Furniture']
# Cat Furniture -> ['Cat Window Beds & Perches']
# Pet Beds -> ['Hammocks', 'Pet Cots', 'Pillow Beds', 'Radiator Beds']
# Animals & Pet Supplies > Pet Supplies > Cat Supplies > Cat Furniture > Cat Window Beds & Perches
That is the leaf the cookbook expects for its own listing, in 8 requests. The search is two functions and about twenty lines, where the cookbook’s search code runs to about 150 LOC. To be clear - I am not claiming that this is “better” in any way, I’m trying to show that we can achieve a similar outcome with a completely “model-native” approach. The potential advantage of a real RDM module in this situation is that it generalizes without writing bespoke “decision plumbing” for each new use case.
Signs of Life
It was very important to prove that RDMs work without any generative function at all, both as a sanity check on the implementation as well as to show how this approach is able to marry the concepts of RLM with these new System One models. The challenge with having “only decisions all the way down” is that you potentially need to enumerate the entire external state before running the RDM, because there is no easy way to generate new decisions to make.
My hunch is that for many applications such as browser use, video games, on-device decision models, and even computer use - this actually might be enough. You just enumerate every possible option in advance - then have the RDM sort through it all.
When you add LLMs to the mix, things start to come alive. I think this is part of why developers have found Jev so refreshing - these decision models are a powerful balancing force. The whole industry has been focused on expanding the divergent capabilities of models, TypeSafe has brought convergent capabilities into the spotlight.
everyone has been building divergence (generate)
— Raymond Weitekamp (@raw_works) September 16, 2026
i'm very excited for jev as convergence (decide)
yin and yang https://t.co/upbOC4ZsYb
The pair of generative model + decision model makes for a surprisingly lively automaton. One way of showing the difference here is by plotting the graph of calls to the decision model.
Here are the two handbook searches from earlier in this post. Every blue dot is one request to Jev. The grey tree behind it is everything the code laid out before the run started.
The shapes are regular because they are the shape of the data. Jev chooses a path, and nothing more.
TypeSafe’s own smart home assistant demo shows the general idea: let an LLM generate when the decisions aren’t enough. A yes/no question spots a request that asks for more than one thing, an LLM splits it into single commands, and each command goes back to the decision model. A request that is only conversation falls through to an LLM for a reply.
The graphs below came from our attempt to apply RDM to that smart home example, with a simulated house that pushes back: it refuses to lock a door that is open. Jev decides at every dot. The LLM only ever writes the next thing Jev is asked about, and each orange square is one of those.
Nothing was laid out in advance, and the same program made all four shapes. No line of it says to close the door before locking it. It also dies in ways nobody wrote: the last run is the same command, written and tried again at every level until the request limit.
As a final thought to wrap up this initial demonstration of combining language models and decision models, I want to make it clear that there are many possible permutations and combinations. I specifically am avoiding what I would consider the obvious one: giving an LLM agent a decision model as a tool. I’ve deliberately put the System One model at the center here, as I think that really highlights the potentially novel behavior.
Communicating with Maps
Now as soon as we add a second category of model, we need to think critically about the interface. The beauty of DSPy is that we can force the generative model to output a TypeSafe-compatible response, and this alone unlocks the “yin and yang” flow between the decision models and the generative models.
In practice, I found the direct connection between the System One models and large language models to be clunky. I went searching for a new primitive, something that would bridge the divide. Admittedly, this communication interface is the piece I am still wrestling with.
The working idea is a “map”. A decision model can only point, so something has to say what it is pointing at. A map is that: a set of names, each with a description for the model and a thing for the code.
point = dspy.Predict("request, house: dict -> next")
point.set_lm(TypeSafe("jev-latest"))
def run(lines, question, **shown):
"""`lines` is the map: {name: (description, thing)}. Returns the things that were done."""
done = []
while True:
options = tuple((name, description) for name, (description, _) in lines.items())
signature = point.signature.with_updated_fields("next", type_=Choice[options], desc=question)
name = point(signature=signature, **shown).next.value # the model points at a name
thing = lines[name][1] # the code looks up what the name stands for
if thing is None:
return done # nothing behind the name: stop
print("pointed at:", name)
thing()
done.append(thing)
That loop is the whole mechanism: show the names, the model points, the code does the thing behind the name, and the names are shown again. Here is a map for a house with a TV and a light.
house = {"tv": "off", "lights": "on"}
lines = {
"tv on": ("Turn the TV on.", lambda: house.update(tv="on")),
"tv off": ("Turn the TV off.", lambda: house.update(tv="off")),
"lights on": ("Turn the lights on.", lambda: house.update(lights="on")),
"lights off": ("Turn the lights off.", lambda: house.update(lights="off")),
"stop": ("Everything `inputs.request` asks for is already so in `inputs.house`.", None),
}
question = "What should happen next for `inputs.request`, with the house as `inputs.house` shows it?"
done = run(lines, question, request="Movie time.", house=house)
# pointed at: tv on
# pointed at: lights off
So far this is tool calling with a decision model. It gets more interesting when the map changes. What was just done can go back on the map as one more name. Its description says what it does, because the description is all the model sees of it.
lines["movie time"] = ("Do again what was done for 'Movie time.': TV on, lights off.", lambda: [thing() for thing in done])
house.update(tv="off", lights="on") # the next evening
run(lines, question, request="Let's watch a film.", house=house)
# pointed at: movie time
Two decisions became one, and a request in other words found it. Nothing matched strings: the decision model judged that “Let’s watch a film.” means the line kept for “Movie time.” This is the one map result that has held up so far. Over a few simulated days of requests to a toy house, a map that keeps what was pointed at got 87% of them right, against 66% for a map that keeps nothing, and it needs no LLM.
The part I am still working out is who writes that new line. Above, I wrote it in Python. The point of a map is that both kinds of model can change it, each in the way it is able to:
-
An LLM cannot write a function, but it can write a name, a description, and a list of names that are already on the map. That is a structured output.
-
A decision model cannot write at all, but it can point at a line that says “keep what was just done” or “forget that one”.
Both of these run today, as the same kind of edit to the map, and neither has yet done better than my one line of Python. Asked what to keep, the decision model kept everything and never chose to forget. Lines written by an LLM have not beaten plain pointing in any project we have tried. The one place an LLM has looked useful so far is smaller: writing the description of a line that was kept, so that a request in other words finds it.
A name can also stand for a group of more names, or for another map, so that pointing at it asks a question one level down. That is where the recursion comes back in: maps of maps of maps…
…but it is time to converge.
What is it good for?
1. “What is it good for?” 2. “Where are the benchmarks?” I can already hear the replies to this post.
1. I don’t know and 2. the benchmarks will come only after I nail down a v1 of the RDM module.
One of my favorite books about life is “Why Greatness Cannot Be Planned”, which happens to also be about machine learning (and specifically novelty search over objective optimization, which is highly relevant to the lineage of both RLMs and Jev). I am not sure where RDMs will be “better”, and I find Diogo’s distaste for benchmarks to be extremely refreshing. Right now I’m more interested in the novelty search: what might be possible with Recursive Decision Models that is impossible or impractical with other approaches? I also know that I need more “reality backpressure” to refine what the “essence of RDM” is.
What’s Next?
-
I don’t know exactly, but I will organize and share more code and examples soon
-
Definitely a
dspy.RDMmodule that you can use -
Maybe even a standalone RDM agent, if you ask me nicely
-
Diogo Almeida on Latent Space: “We need a new class of models. We’re not attached to naming that class of models. The most accurate name we’ve come up with is System One models… there’s a reason why we don’t call them decision models… System One is beyond that. That’s all I can say. We didn’t expect this to be our big launch, so we have stuff in the tank.” ↩︎