Letting an agent act, but not without asking
RAION runs on my own machine and can touch my files, my shell and my inbox.
An assistant that can only talk is easy. One that can write files, run shell commands, send mail and hit external services is a different thing, because now every mistake it makes is a real mistake. That is the problem RAION is mostly about. The chat interface was the quick part.
A supervisor that spawns workers
The agent runtime is a LangGraph state machine. A supervisor node reads the conversation and decides what to do. If the task splits into independent pieces, it calls a spawn tool with all of them at once, and each becomes a worker agent with its own subgraph and its own restricted tool list.
Workers get an explicit tools_allowed array rather than the
full toolset. A worker sent to fetch a page and save a screenshot gets
browser tools and write_file, and nothing else. If it
decides halfway through that it wants the shell, it does not have it.
That is easier to reason about than a single agent holding every
capability for the whole run.
Four limits, because loops happen
An agent that can spawn agents can spawn agents. The run context carries four counters and every one of them is a hard stop:
- recursion depth, three layers of supervisor to worker
- total agents, ten including the supervisor
- total tool calls across everything, fifty
- wall clock, ten minutes
These are deliberately unsubtle. There is no adaptive budget or clever backoff. When a limit is hit the spawn tool returns a plain string saying so, and the model reads that as a tool result and has to work with it. A number I can find and change beats a heuristic I have to debug at 3am, and the failure mode of a runaway agent loop is a bill.
The permission layer is the actual point
Every tool call passes through a policy function that returns one of two
answers: auto or ask. Reads and directory
listings inside the workspace are auto. Writing a new file is auto.
Overwriting a file that already exists is ask. Deleting anything is
always ask, workspace or not.
Shell is where it gets specific. There is a regex list of patterns that
force an approval prompt, and it is long on purpose:
rm, git push --force,
git reset --hard, DROP TABLE,
sudo, kill, registry edits, and
curl piped into a shell. Nothing there is hypothetical.
Each one is a command that would ruin an afternoon if a model guessed
wrong about it.
When a policy returns ask, the graph calls
interrupt(). That pauses the run and checkpoints it,
mid-execution, to SQLite. An approval card appears in the browser. If I
approve, the graph resumes from exactly where it stopped with the tool
call intact. If I deny, the model gets that as the tool result and
continues without it. Every decision is written to an audit log.
The pause-and-resume is the part I would point at. It is not a confirmation dialog bolted on top; the agent's whole state is serialised at the moment of the question, so approval can happen a minute later or after a page refresh and the run picks up unchanged.
The rest of it
Around that core there is a fair amount of plumbing: WhatsApp through Green API polling groups every thirty seconds, a Telegram bot for notifications and questions from my phone, Gmail, Drive and Calendar over OAuth2, MCP connectors so any MCP-compatible server becomes usable tools, a local vector store for search over my own files, and automations described in plain English that compile down to cron jobs, Gmail triggers or filesystem watches.
Secrets that have to be stored, MCP environment variables and OAuth tokens, are encrypted with Fernet rather than sitting in the database as text.
What I would tell myself at the start
The capability list is the easy half. Adding a tool is an afternoon. Deciding which tool calls a person needs to see before they happen, and building a runtime that can stop cleanly in the middle to ask, is the part that took the time and is the only reason I am willing to run this against my own machine.
It is the same conclusion I reached building a WhatsApp bot for a college office: the interesting engineering in an AI feature is rarely the model call. It is what happens around it when the model is wrong.