Agent orchestration in production
  1. Tools
  2. Memory
  3. State
Refik Anadol Studio
[email protected]
01 / 22</> press C

Agent Orchestration
in Production

Tools, Memory and State
Mert Cobanov
Senior AI Engineer, Refik Anadol Studio
DevNot

Dataland

The world's first Museum of AI Arts, in Downtown Los Angeles. I build its museum agent at Refik Anadol Studio: visitors ask about artworks, the building, planning a visit. Every visitor is anonymous and can try as often as they like.

FastAPIPydantic AIQdrantRedisPostgreSQL
tools memory state
What do we leave to the model, and what stays in code?
Let the model decide what needs language.
Let code decide what needs guarantees.

An agent is a loop

The model picks the next move. The loop around it is ours: what may run, when to stop, what to keep.

model decides code decides

Where it breaks

  1. ToolsLoops that never endOne idle timeout, 761 API calls in 60 seconds (OpenClaw, 2026).
  2. ToolsConfident calls to the wrong toolWrong tool and wrong arguments are the most common tool failures (Anthropic, 2025).
  3. MemoryContext that grows until it forgetsGPT-4o: 99.3% at short context, 69.7% at 32K tokens (NoLiMa, 2025).
  4. StateState that lives in one processTwo writers orphaned a tool result; every resume failed (Claude Code issue, 2026).

Draw the line on purpose

Two questions for every decision: does it need language? What does a wrong answer cost?

The model proposes. Code disposes.

The model fills a typed proposal. Code checks it against reality and decides what actually happens. Rejections go back as instructions.

01Tools
The model's hands. Every tool is a prompt, an API and a blast radius at the same time. Design them for the model.
Scope them to the phase.
Budget them in code.

Tools are prompts

  • Few, coarse tools beat many thin ones
  • Names and docstrings are instructions
  • Return what the model needs to say next, not your table
  • Errors are instructions: say what to do instead

Not every tool, every turn

Every tool in the prompt is another wrong option and more tokens. Expose only what makes sense in the current phase.

Budgets live in code

Never ask the model to stop. Make it unable to continue.

  • Request, tool call, token and cost limits
  • A wall-clock timeout per turn
  • Same call, same arguments: refuse it
  • A deterministic fallback answer

Storm: OpenClaw issues #76293 and #78865, May 2026, user-reported.

Untrusted text gets no new powers

Private data, untrusted text and a way out: together they leak. You can't prompt your way out of that. Remove one leg in code.

  • Visitor identity comes from the session, never from model arguments
  • Read tools and write tools are separate
  • Hours, prices and refund rules come from data
  • The renderer loads no remote images and links only to museum domains

Lethal trifecta: Simon Willison, Jun 2025. ForcedLeak: Noma, Sep 2025. Rule of Two: Meta, Oct 2025.

02Memory
The model is stateless. Memory is everything the system does to carry context forward. The longer version, with interactive demos, lives at memory.cobanov.dev.

Context is not a database

Trimming history keeps you under the token limit. It also deletes facts: silently, by age, not by importance.

Read fast. Write carefully.

Recall is on the hot path, with a latency budget. Remembering happens after the reply, in a worker.

Updating is harder than remembering

The model proposes facts. Code decides what is stored, what is replaced, and what must never be written down.

03State
Memory is what we know about the visitor. State is where we are in the conversation, right now. Make it explicit.
Keep it out of the process.
Run one turn at a time.

Make the state explicit

The model can suggest the next phase. Code owns the transitions, and the phase decides which tools exist. A state machine for the process, not a script for the dialogue.

Same pattern in production: Rasa CALM, Salesforce Agent Script, Intercom Procedures, Decagon AOPs.

State lives outside the process

Any worker can serve any turn. The lock is for efficiency, the version check is for correctness, the key makes side effects happen once.

One turn, as a trace

You can only improve what you can replay. The trace is where the line between model and code becomes measurable.

Takeaways

  1. 01Let the model decide what needs language. Let code decide what needs guarantees.
  2. 02Tools are prompts. Permissions are code.
  3. 03Memory is a write problem, with history and a delete button.
  4. 04State lives outside the model and the process. Every write is versioned.
  5. 05Measure pass^k on turns you can replay.

Thank you!

Every number in this talk, with its source
orchestration.cobanov.dev/sources
More: memory.cobanov.dev, kvcache.cobanov.dev
snippet.py python
Chide code→steps move the highlight