Claude Code is excellent at one task. It reads the context, makes the change, runs the tests, and stops. The problem is everything that happens between tasks. What comes next? What was the state of the branch when I walked away? Did the last thing actually land? The agent does not remember, because there is nothing to remember into.

I kept losing the thread. So I built Sugar - an open-source task queue for coding agents. This note is the short version of why.

The gap it fills

The Claude Agent SDK and Claude Code both assume a human is in the loop, watching one task finish before starting the next. That works for a single focused change. It breaks down when you have a backlog of ten things and you want the agent to work through them while you sleep. You end up babysitting the terminal, which is exactly the thing you were trying to automate.

Sugar sits one layer above the agent. It holds the queue, tracks state per task, and re-invokes the agent for the next item when the current one finishes. You add tasks the way you would add issues to a tracker - by type, priority, and a short spec - and the queue drives the agent through them.

Sugar - one agent, a queue, and a memory
Task queue
Sugarstate + dispatch

Why a queue, not just a script

The naive version of this is a shell script that runs the agent on a list of files. I tried that. It falls over for three reasons:

  1. State leaks between tasks. Without per-task isolation the agent carries context from task A into task B and starts making decisions based on the wrong premise.
  2. Failure is ambiguous. A script exits non-zero. Did the task fail, or did the tests flap? You cannot tell without reading the log, and the log is 4,000 lines.
  3. Nothing remembers. Next session, you start from scratch. The agent has no idea what was already attempted.

A queue fixes all three. Each task is its own unit of work with its own result. Failures are categorized. State persists across sessions. The agent gets exactly the context it needs and no more.

What it is not

Sugar is not an orchestrator that fans out to many agents in parallel, and it is not trying to be a CI system. It is a thin layer that makes one coding agent usable for real backlogs. If you want multi-agent swarms, this is the wrong tool. If you want to stop babysitting a terminal, it is the right one.

Where it is now

Open source, on GitHub: roboticforce/sugar. We use it internally for the boring part of client work - the long tail of small, well-specified changes - so the human hours go to the parts that need judgment. If you try it, tell me what breaks.