← All projects

Case study

Hive Mind: an AI chief of staff with a team behind it

A multi-agent system I built to run my own workload. I message it like a colleague, and it sorts, delegates and remembers, with firm rules about what it's allowed to do on its own.

Role
Design, build and run
Built with
Hermes Agent, Claude & OpenAI models, Discord, Signal, Git
Status
In daily use, still evolving

The problem

I juggle IT support, marketing and a growing pile of my own projects, mostly on my own. AI assistants helped, but every new chat started from zero: I'd re-explain the project, the decisions already made and what I'd tried, every single time. And one general-purpose assistant is a poor fit for work that jumps between research, writing code and looking after machines.

All that re-explaining, context-switching and hand-holding was costing me two to three hours a day. I didn't want a smarter chatbot; I wanted a small team that remembers, knows who does what, and only comes to me when it genuinely needs a decision.

What I built

I talk to one agent, Argus, from Signal on my phone or a Discord channel on the desktop. Argus works out what the request needs and hands it to the right specialist, each running as its own agent with its own tools and limits:

  • Argus · Chief of staff

    Takes every request, decides who should handle it, and is the only one allowed to update the shared memory.

  • Dex · Developer

    Writes and tests code. Can edit, run tests and commit locally, but pushing, deleting and secrets stay off-limits.

  • Scout · Researcher

    Digs into questions, compares options and reports back with sources.

  • Talos · Sysadmin

    Looks after the machines and runs scheduled checks, like watching for a software release every six hours.

  • Athena · Reviewer

    Called in sparingly, for architecture, security and major releases.

  • Iris · Artist

    Handles images and visual assets.

The piece that makes it work is shared memory: a plain Markdown knowledge base kept in Git. Every agent reads from it, so context and past decisions carry over between sessions. Only Argus writes to it; specialists propose updates and Argus files them, so it stays tidy instead of turning into a pile of contradicting notes.

A live "office" dashboard shows each agent at its desk when it's working and back in the huddle when it's idle, so I can see what's happening at a glance.

Guardrails

An agent team is only useful if you can trust it while you're not watching. The rules are simple and hard-coded:

  • It asks me, by rule, not by feel. Anything touching money, credentials, security, production systems or deletion, a task that fails twice, or sources that disagree goes back to me. The model's own confidence never decides.
  • Least privilege per agent. The developer can edit, test and commit locally; pushing, deleting, secrets and configuration are blocked outright. The reviewer can only read and run tests.
  • A spending ceiling. Routine API spend has a small monthly cap, and anything above it, or any billing change, needs my approval.
  • No single point of failure. Each agent has a primary model and an automatic fallback to a different provider, so a rate limit or outage doesn't stop the work.

What went wrong, and what I changed

The first version got too clever. The agents spent about five days building verification systems around their own work (receipts, ledgers, checks on the checks) instead of shipping anything. It looked diligent and produced very little.

So I reset it. "Done" now means three plain things: the agent finished cleanly, the tests pass, and the change actually exists. Each task goes to a fresh run as one whole job rather than being sliced into dozens of steps, and the reviewer only comes in when the stakes justify it. It's cheaper, faster and far easier to reason about.

That's the main lesson I bring to client work: the value is in small, well-guarded loops that do one job reliably, not in the most elaborate agent setup.

Results

So far its biggest project has been itself. Most of what's in the office dashboard, from the walking animations to the live status and job labels, was built by the team from requests I sent in plain language, each change delivered with its tests passing. It also runs standing jobs, like checking every six hours for a software release I'm waiting on. Next, it takes over the day-to-day on my apps and websites.

A real example, from one evening:

  1. From my phone: “Get Dex to make Argus show Working in the office whenever he’s handling a request, on any platform. Store only a state and a timestamp, never message text. Add tests, deploy it and commit.”
  2. Five minutes later: Argus had read the existing code, written a full brief and started Dex on it in the background.
  3. An hour in: Argus reviewed Dex’s work, found tests failing and a rule broken, and sent it back once with corrections. I didn’t have to spot either problem.
  4. Twenty minutes after that: “Built, tested, deployed and committed,” with every test passing, then it stopped and asked me before restarting anything live. A couple of quick “restart now” approvals from me later, it confirmed the feature working end to end.

Total effort on my side: a handful of short messages, mostly saying yes.

Could something like this help your team?

Most businesses don't need six agents. They need one or two doing a specific, repetitive job well, with sensible limits. If there's work eating up your week, I'm happy to tell you honestly whether an agent is worth it.

Start the conversation