note · October 2, 2026
A Flag Is Not a Guarantee
How nats-trail started as yet another NATS GUI and ended up as a place where an agent can look at production without being able to touch it.
In an event-driven system, the easy question is “what’s in this stream?”. Every NATS GUI answers it, and
answers it well. The hard question is a different one: “why did this flow fail?”. Answering it means
following a single request_id across four streams, three services and a dead-letter subject, and today
you do that by hand, with one terminal per stream and the patience of someone doing a jigsaw puzzle
without the picture on the box.
nats-trail answers that question directly: a web UI, a CLI and an MCP server over one bounded query engine. But it wasn’t born that way. It was born as one more GUI, and the story of how it stopped being one is all there in the git history, which is the only diary I keep without meaning to.
First I did what everybody does
The first commit is from May 29. Within three days I had everything you’d expect: connection contexts,
Core and JetStream panels, replay and live tail through an ephemeral consumer, a JSON viewer with a tree
and search, vendor-agnostic DLQ subject detection, and a guard that stops you before you connect to a
context marked prod or staging. On June 1 the commit message says, with all due solemnity,
“complete v1 inspection gaps”. The next one says “refactor design”, which is what you write when you
realize the thing you just finished isn’t the thing you wanted.
What I wanted wasn’t another panel. NUI already exists: it’s public domain, has more than 650 stars and ships desktop binaries for three platforms. Competing with it on GUI breadth, alone, is a race you lose gracefully or lose ungracefully, but you lose. I wrote that down in the roadmap, in my favorite section of any roadmap: Explicitly not doing.
The one who had to read the streams wasn’t me
On June 9 there’s a commit that changes everything: “feat(core): add query envelopes for agents”. That day and the next bring in the tool contracts for the CLI and MCP, the stdio server, timeouts, an audit entry for every call with its origin, and strict input validation: required fields, types, numeric ranges and unknown fields, all checked before anything runs.
The idea was simple. If the hard question is following a flow across four streams, the best one to ask it isn’t me with four terminals: it’s an agent with four tools. But an agent doesn’t read screens, it reads contracts. So every response became a small, boring envelope, on purpose:
{
"query": { "contextId": "dev", "limit": 50 },
"summary": { "returned": 12, "truncated": false },
"results": [],
"nextCursor": null,
"warnings": [],
"errors": []
}
The boring part is the interesting part. A human tolerates an ambiguous answer because they look at it and understand. A model fills it in with whatever seems most likely, which is not the same as what happened.
“Found nothing” without coverage is a lie with good spelling
On June 10 the bounded query engine landed, and it’s the decision I’m proudest of.
Scanning a whole stream to find a message is the most honest way to search and the most expensive. So every scan has a budget: 10,000 examined messages by default, 100,000 at most. Nothing unusual so far; every tool sets a limit. What matters is what happens when the limit is hit.
- If you pass no time window and no cursor, the query covers the most recent
maxScansequences and the envelope carries aquery.window_defaultwarning: I answered over a slice I picked myself. - If the scan stops on budget, you get a non-null
nextCursorand aquery.scan_truncatedwarning: I didn’t look at everything, and here’s where to continue.
Nobody:
Absolutely nobody:
An agent holding an empty result: “There are no errors in the payments stream.”
That’s the bug the contract prevents. The agent always knows whether coverage was complete and how to continue. “I found nothing in the last 10,000 messages” and “there is nothing” are two different sentences, and the second one can only be said once it’s been earned.
A flag is not a guarantee
None of this is worth much if the agent, besides looking, can break things. The first temptation is the
obvious one: a --read-only flag. I ruled it out for the very reason it exists: a flag can be flipped. A
badly copied environment variable, an inherited config, someone in a hurry on a Friday. A guarantee that
depends on nobody touching a knob is a suggestion.
So the rule became structural, not procedural:
| Where | Can it write? | Why |
|---|---|---|
MCP runtime (McpRuntimeData) |
No, never | The interface it’s handed has no write members. They aren’t disabled: they don’t exist. There’s nothing to call. |
| UI and CLI | Yes | Only with an authenticated human session or a write-scoped token, audited with its arguments. |
| Any flag, env var or config | Can’t change it | None of them can make a write reachable from executeMcpTool. |
When writes arrived in August (editing and purging KV keys with optimistic concurrency, creating streams and consumers, request/reply, storing objects), every one of them landed on the human side. The whole phase is called “Writes, on the human side only”, and it has a line at the top saying nothing in it may weaken that rule.
For teams that genuinely want an agent to write, the plan is a separate binary, natstrail-mcp-write,
with its own package and its own interface. Installing it is a deliberate act, with a different name in
the MCP client config. It will never be a flag on natstrail-mcp, because the whole value of saying
“it’s read-only” is that it can’t be switched off.
Two months of silence and a phase zero
After June 10 the history goes quiet until August 13. I won’t pretend it was a strategic pause. What I do know is what happened when I came back: I opened the roadmap and the first phase said “Nothing else matters until someone who is not the author can run this.”
I had a bounded query engine, auditing, a Sentry integration and fourteen MCP tools. I had no license, and
without a license nobody can legally use the project, however good it is. It wasn’t on npm. The binary was
called nats-ui, a name already taken on npm by an unrelated package, so the alias wasn’t deprecated: it
was dropped.
That day brought the Apache-2.0 license, package builds, a Docker image with a docker-compose and a
nats-server -js next to it, CI on Node 20 and 22 with a serve smoke test, and README screenshots seeded
with demo data. npx nats-trail serve started working from a clean machine. Between August 13 and 15 the
project went from 0.1.0 to 0.5.0, which is what happens when the work was already done and the missing part
was letting someone open it.
The tests found what I didn’t
On August 14 I wrote the query engine’s test suite: filters, envelope limits, truncation, and guards on the write boundary. 67 tests. One of them failed, and not because the test was wrong:
// The > wildcard matches one or more remaining tokens, never zero:
// the pattern orders.> must not match the bare subject orders.
if (token === ">") return s.length > i;
It used to say return true. In NATS, orders.> matches orders.created and orders.eu.paid, but not
bare orders. Mine did. It’s a tiny bug, the kind nobody sees in a demo and that in production shows you
messages you didn’t ask for, with complete conviction. The interesting part isn’t that it existed; it’s
that it had been there since the first commit in May, in a project I was using, and I never noticed. The
things you use every day are exactly the things you stop looking at.
What I don’t do, on purpose
The decisions that keep me most organized are the ones that say no:
- I don’t compete with NUI on GUI breadth. It already won, and it deserves to.
- I don’t do cluster administration or topology management. That’s Synadia Control Plane’s product. Cluster visibility is deferred, not refused, and the difference is written down.
- There are no agent writes behind a flag. See above, insistently.
- I don’t stream Object Store payloads through the bridge. Metadata only.
- I don’t decode protobuf field names. The wire format, yes, schema-less and with no dependencies; the names need a per-subject schema registry, and that’s another project.
- I’m not on Smithery. It now requires an
.mcpbbundle, which is a new distribution artifact, not a config file. I am on the MCP registry, republished on every tag.
The correlation index follows the same logic as writes: it’s opt-in per stream, it lives in node:sqlite
(hence the Node 22 floor), and building it is a human action. Querying it is automatic. The agent
benefits from the index; it doesn’t get to decide when to build one.
Milestones
| Date | What happened |
|---|---|
| May 29 | First commit: contexts, Core and JetStream panels |
| Jun 1 | “v1 complete” and, right after, refactor design |
| Jun 9 | Envelopes for agents, CLI and MCP contracts, auditing |
| Jun 10 | Bounded query engine, tokens, connection pooling |
| Aug 13 | Phase zero: license, npm, Docker, CI, renamed to nats-trail |
| Aug 14 | Tests, the > bug, flow tracing, human-side writes |
| Aug 15 | Schema-less protobuf and msgpack, PagerDuty, correlation index, Helm chart, v0.5.0 |
What an event bus taught me
I set out to build a tool for looking, and ended up building a tool about the limits of looking. Almost everything worthwhile in nats-trail is a way of saying this far: this far I scanned, this far I’ll let you reach, this far this binary goes. It surprised me how much that changes the relationship with an agent. A model doesn’t need you to trust it with production; it needs you to tell it precisely which part of production it saw, and for there to be no version of the conversation in which it can touch it.
The other lesson is quieter. I spent two months with a project that worked and that nobody else could use, and in my head it was “almost ready”. It wasn’t. A license, a free name on npm and a command that runs on a clean machine are worth more than the next feature, because they’re the difference between something that exists and something I know exists. What comes next (multi-user, cluster visibility) are open design questions, and I’d rather leave them written down as questions than pretend they already have answers.
Sol Soletti
- #NATS
- #MCP
- #agents
- #nats-trail