Agent orchestration in production
  1. Tools
  2. Memory
  3. State
  4. Decisions
Refik Anadol Studio
[email protected]
01 / 22</> press C

Agent Orchestration
in Production

Tools, Memory and State
Mert Cobanov
Senior AI Engineer, Refik Anadol Studio
DevNot Developer Summit 2026, 17 October

Dataland

The world's first Museum of AI Arts, in Downtown Los Angeles. I build its museum agent at Refik Anadol Studio.

  • Answers questions about the artworks and the museum
  • Inside, it reads your heart rate and room from wearables, and sometimes speaks first
  • Visitors are pseudonymous and every request is authenticated, but chat has no rate limit: assume unlimited attempts
FastAPIPydantic AIGeminiPostgreSQLQdrant, via a RAG service
The model picks a tool The visit opens a turn Code acts on a typed verdict
What do we leave to the model, and what stays in code?
Let the model decide what needs language.
Let code decide what needs guarantees.

An agent is a loop

The model picks the next move. The loop around it is ours: how long it may run, how often each tool may fire, what we keep.

model decides code decides

The harness is half the agent

Everything that is not the model is the harness. That is where the guarantees live.

What broke in ours

Five incidents from our agent's git log in 2026. Every fix landed in code, not in the prompt.

Draw the line on purpose

Two questions for every decision: does it need language? What does a wrong answer cost?

The model proposes. Code disposes.

Structured output guarantees the shape, not the truth. Our only typed output is a complaint judge: the model fills the verdict, code decides what that verdict may do.

01Tools
The model's hands. Every tool is a prompt, an API and a blast radius at the same time. Design them for the model.
Scope them to the surface.
Budget them in code.

Tools are prompts

The name, the docstring and the returned text are all read by the model. Each one steers the next step.

  • Docstrings are routing: one sentence decides which tool fires
  • Return what the model should say, not what the sensor sent
  • Errors are instructions: say what to do instead

All three cases from our museum agent, June 2026.

Scope tools by surface

A tool the model cannot see is a tool it cannot misuse. We decide before the model runs.

  • The app picks the agent, not the model
  • Before the visit: no live-state tools at all
  • A unit test fails if one is added
  • Phases or dozens of tools: filter per phase, or search

Selection degrades past 30 to 50 tools: Claude docs. Five MCP servers, about 55K tokens of definitions: Anthropic, Nov 2025.

Budgets live in code

Never ask the model to stop. Make it unable to continue.

  • One turn ran the same search three times
  • A repeat returns data plus a directive, not an error
  • Size timeouts to the real latency
  • We cap: 60 s per turn, calls per tool, 50 requests
  • Not yet: tokens, cost, a fallback answer

Our numbers: our logs, June 2026. Storm: OpenClaw issues #76293 and #78865, May 2026, user-reported.

Untrusted text gets no new powers

Private data, untrusted text and a way out: together they leak. You can't prompt your way out of that. Remove one leg in code.

  • Our model has zero write tools
  • Visitor identity comes from the session, never from model arguments
  • A regex gate runs before we pay for a model call; a managed scanner is the second layer
  • Images leave as structured cards, not as markdown the client fetches

Lethal trifecta: Simon Willison, Jun 2025. ForcedLeak: Noma, Sep 2025. Rule of Two: Meta, Oct 2025.

Delegate reads, keep one writer

A sub-agent buys a clean context, not speed. Let many read, let one write, and put the limits in config.

02Memory
The model is stateless. Memory is what the system sends again, and what code writes down. What must survive every turn goes into instructions.
What must survive the visit becomes a record.
Longer version: memory.cobanov.dev

What must survive every turn

We replay the whole conversation on every turn. Our system prompt still vanished after the first one.

  • No error, no warning: answers grew from 82 to 392 words
  • The fix: instructions, sent with every request
  • What must survive every turn goes into instructions, not into the transcript

Measured on our museum agent, June 2026. NoLiMa (Modarressi et al., 2025). Context rot (Chroma, 2025).

Code writes the record

The model only narrates it. Nothing the model says is stored as a fact about the visitor.

  • Persist first: a 2xx means durable
  • The model narrates a bounded fact sheet
  • Recall is a tool scoped to one visitor
  • Forgetting is an endpoint, not a prompt
03State
Memory is what we know about the visitor. State is where the visit is, right now. Owned by the systems that see the visit.
Moved by plain code, never by the model.
One writer per conversation at a time.

The visit is the state machine

Our agent has no phase field. The museum's systems move the visit, and plain code decides when the guide speaks first.

Two writers, one conversation

A visitor and our own rule engine write to the same conversation. What we run, where it stops, and the next step.

04Decisions
Some decisions are a probability with a threshold, not a paragraph.Typed questions in, probabilities out.
Code owns the threshold.
Fast enough to ask at every gate.

A decision is a distribution, not a string

Structured outputs guarantee shape. Typed decisions show doubt. Code decides what doubt costs.

Decisions at every gate

Rules first, one request with parallel questions, the threshold in code, every verdict in a ledger.

Run decision models locally

Ollaya runs open decision models on your own machine, the way Ollama runs LLMs. Same API as TypeSafe.

  • One forward pass: probabilities for your questions, no text
  • Change the base URL and the model name; agents get it as an MCP tool
  • Open models land near Jev on one public split. Measure on your own data.

I maintain Ollaya: ollaya.dev. Front page of Hacker News, 618 points, 25 Sept 2026. Community map: github.com/cobanov/awesome-jev.

One turn, measured

Today we measure a turn with one log line. Next we replay it, many times.

Takeaways

  1. 01Let the model decide what needs language. Let code decide what needs guarantees.
  2. 02The harness is half the agent. Version it, and test every change like a model upgrade.
  3. 03Permissions are architecture. A model with no write tools cannot be talked into writing.
  4. 04Code writes the record. The model narrates it.
  5. 05Some decisions are a probability and a threshold, not a paragraph.
  6. 06Measure k turns you can replay, not one demo.

Thank you!

Every number in this talk, with its source: orchestration.cobanov.dev/sources
Decision models, locally: ollaya.dev. The community map: github.com/cobanov/awesome-jev
More: memory.cobanov.dev, kvcache.cobanov.dev
snippet.py python
Chide code→steps move the highlight