Annie, Anyshift's AI site reliability engineering (SRE) agent, started as one big, heavily instrumented Claude Code SDK service that reached into half of our product. Every large feature came with a real risk of regressions, so we rebuilt it around clear boundaries to iterate faster without breaking things.
Annie v3 in three points
- Why we rebuilt it. We had a dual mandate: shipping technical features like streaming, or fix bugs, without degrading answer quality; improving answer quality without introducing bugs.
- What changed. We split the agent in three. The harness is a business-neutral execution layer, the brain holds operational knowledge and skills, and side effects handle everything that talks to the outside world: delivery, integrations, notifications.
- What we gained. Features and answer quality now evolve independently, and the system is much easier to maintain.
The architecture in one diagram
Each layer has one job and one reason to change.
- The brain changes whenever we learn a better way to investigate incidents, which happens often.
- Side effects change every time we ship a product feature, so constantly, but they're now decoupled from everything else.
- The harness only changes when the model runtime or some deep execution logic evolves; once it works, we rarely touch it.
Their release cycles no longer drag one another along.
1. The brain moves at the speed of operational learning
Annie is an agent built on LLMs. To keep it reliable and trustworthy, we constantly improve its operational knowledge by updating its prompts and skills. Keeping that knowledge inside the execution service made improvements expensive to ship and regressions hard to diagnose. Pieces of prompts lived deep inside the code, so even small updates were risky and slow, and outdated prompts could linger without anyone noticing.
We moved every system instruction, skill, piece of runtime guidance, and Model Context Protocol (MCP) template into a versioned annie-brain repository. At session start, the harness writes only the enabled skills into the workspace: disabled means unavailable, not merely ignored. In theory, you could point any harness and model at this directory and get results close to what production would give you.
Everything Annie knows about operations now lives in one place, which we can update and review without touching the execution layer. That's what lets us iterate quickly and fold user feedback straight into Annie's behavior.
The brain also has its own Git history. We can compare, revert, or fork it without redeploying the harness, building on our earlier work encoding on-call judgment as loadable SRE skills.
2. The harness is a black box
Say we want to support a new model vendor. Without a model-neutral harness, vendor-specific branches would spread across the product, every feature would have to speak several protocols, and maintenance would get worse with each vendor. Our rule is simple: a single if vendor == "new_model" outside the black box is a failure.
We looked at the current capabilities of Claude Code and Codex and designed a minimal shared interface that fits both, with vendor-specific logic isolated in adapters. The interface covers every interaction pattern we need: input, result, token streaming, events, and observability. Keeping the contract small follows Erik Schluntz and Barry Zhang's guidance at Anthropic: add complexity only when it earns its cost.
Then we rewrote our existing Claude Code SDK integration as the first adapter behind that interface. With model-specific logic out of the harness, we can swap or upgrade models with minimal changes to the rest of the system.
To test the idea, we shipped OpenAI Codex support. The pull request changed one production file (not counting tests), and it worked the first time.
3. Skills replaced complex pipeline handling code
A lot of business logic used to live in the code wrapping the Claude Code SDK. The biggest piece was a routing pipeline: an LLM call classified each incoming request, then sent it to one of two fully separate pipelines that drove Claude Code differently, one for chat and one for root cause analysis (RCA). Each pipeline had ramifications in other services of our product, so maintaining them meant either duplicated work or painful refactors.
We went back to a plain approach: let a simple harness run and pick the right skill. That let us delete the routing mechanisms entirely. Every message now goes through the same path, Annie's system prompt decides which skill to load (chat or RCA), and each skill carries its own follow-up actions and logic. Skills can also chain, one triggering the next (rca-alert-intake → rca-methodology → rca-formatting), which gives us the pipeline behavior we used to have without the code around it.
The trade-off is behavioral routing: Annie can still pick the wrong skill. That's why we unit-test routing with input and expected-skill pairs. In exchange, the agent got more flexible. Before, once a conversation was flagged as an RCA, there was no going back to chatting about hypotheses. Now every message is just a harness turn with different instructions, so a conversation can start as chat, turn into an RCA when someone asks for an alert investigation, and go back to chat when the context changes.
4. Side effects subscribe to events, so we ship faster
For a long time, long-running investigations had essentially no streaming. Users waited for the final answer with no intermediate feedback and very few ways to steer the investigation while it ran.
We drew a hard line between the runtime and its side effects. The runtime is a black box that broadcasts events (tool calls, skill loads, text deltas, thinking) without knowing who listens. Each subscriber picks up those events and builds its own state to track the investigation's progress and outcome.
That's how we show live progress of the same investigation on several surfaces at once, with minimal coupling:
- the web app
- the Slack bot
- the CLI
- the MCP server
New features like better delivery or self-improvement are now smaller and much less risky. As long as we don't touch the runtime, the investigation still completes. If a side effect has a bug, we fix it on its own without touching the investigation.
5. Traces tell us which part of the run failed
"The model was bad" is not a diagnosis. A weak answer can come from loading the wrong skill, missing tool data, an instruction buried in a long prompt, or reasoning the model abandoned halfway.
Langfuse puts instructions, skill loads, tool calls, timing, token use, and the final answer in one trace. The harness exports that structure through the standard OpenTelemetry Protocol (OTLP), so we can replace Langfuse if we need to. As Jina Yoon describes at PostHog, traces become really useful once reviewed failures turn into evaluation cases.
The cost is exposure: more visibility means more sensitive data stored in more places, so traces need explicit retention, access, redaction, and sampling rules.
6. New testing strategies unlocked
We test Annie in two ways, unit tests and end-to-end evaluations, and they answer different questions.
Unit tests rest on a simple theory:
- if Annie reliably picks the right skill for a given context (input, available tools, memory state),
- and, once a skill is loaded, Annie makes the right tool calls and puts their results together well,
- then Annie will probably do well on a full investigation.
That theory lets us write small, scoped unit tests, split between decision-making and skill execution. They run on synthetic data and cost few tokens, and they catch regressions immediately: when Datadog removes a tool from its remote MCP server, the Datadog skill's unit test fails right away.
Their blind spot is that the real world is messy and rarely binary. That's where end-to-end evaluations come in.
End-to-end evaluations run the whole system on past real-world cases, with all their complexity and randomness. They're slower, costlier, and probabilistic, but they're worth it: every production failure we add to the corpus becomes a permanent regression case, so the same failure can't come back unnoticed.
This follows Hamel Husain's evaluation loop: measure quality, inspect failures, and change the system.
What's still imperfect
Two things will be hard to solve cleanly:
- This design bets that harnesses stay the state-of-the-art way to build agents. If that stops being true, we'll have to rebuild (and we won't be the only ones).
- Prompts and skills aren't just static Markdown files. Some content is generated per user and per context, and we haven't found a clean way to fit that into the brain repository yet. We're working on it.
Did it work?
The Codex pull request was the real test of our abstractions. If adding a runtime had leaked into the brain or the side effects, we'd have known exactly which boundary failed. It stayed inside one adapter, and since it shipped, not a single bug has been traced back to it.
Continue reading:
