The company metaphor has arrived.
Paperclip describes itself as open-source orchestration for teams of agents. Its pitch is direct: bring your own agents, place them in an org chart, assign goals, schedule heartbeats, track costs, and use approvals or pause controls when work needs intervention. The repository names task checkout, budget enforcement, persistent state, scoped secrets, and audit records as product concerns.[1]
This is a more honest shape than twenty unlabeled terminals. Work has an owner. Goals have a route. Spend has a number. A person can see that something is waiting. Those are real upgrades.
The dashboard tells you who should be in charge. The runtime decides who actually is.
Claims need a running artifact.
The public repository is active, MIT-licensed, and ships a CLI package. On this pass, the registry resolved paperclipai version 2026.916.1. The CLI returned that version under Node.js 24.11.0. The checked repository revision was d3e0f0a. Its current README requires Node.js 24.11 or newer and offers an isolated test-drive route as well as a self-hosted install.[2]
That proves there is software behind the screenshots. It does not prove every governance claim in every deployment. A version response checks packaging and launch wiring. It does not test secret boundaries, budget races, cancellation latency, review bypasses, or recovery after a worker dies.
Control words are not control tests.
OWASP names excessive agency as a risk when an LLM receives unchecked autonomy to act. Its prompt-injection guidance also warns that crafted input can steer decisions and reach downstream systems. An orchestration layer can organize those actions, but organization does not remove the risk. It becomes another place where permissions, inputs, and stop behavior need tests.[3]
Terms such as "budget," "approval," and "terminate" need operational definitions. Does a budget stop new work or kill work already running? Does approval protect one tool call or the whole task? Does terminate close child processes and revoke temporary credentials? Can a heartbeat wake stale work after a human paused it?
A fleet needs a cheap roster and deep records.
VS Code 1.139 moved lightweight session and chat metadata into a central catalog while keeping full conversation content in separate databases. Microsoft reports faster listing on a development machine with about 645 sessions. The same release makes pending input and approval visible when compact rows expand.[4]
That is not evidence about Paperclip. It is a useful design signal. A fleet view should stay fast enough to scan, while the evidence for one task remains detailed enough to inspect. A green roster row cannot replace the command, diff, cost ledger, approval decision, or stop receipt behind it.
Test the four paths.
- Scope. Give one agent one disposable workspace, one network policy, and one credential set. Attempt a write, call, and read beyond each boundary.
- Spend. Set a small ceiling. Test concurrent starts, delayed usage reports, retries, and a task already running when the limit lands.
- Review. Mark one action for approval. Try alternate tools, delegated children, resumption after reboot, and stale approvals.
- Stop. Pause the parent. Confirm queued wakes, child processes, network calls, leases, and temporary credentials all reach a named final state.
Keep the receipts boring. Record the revision, policy, task, expected denial, actual output, cost at stop, child-process state, and reviewer. If the system cannot print that packet, the company metaphor is ahead of the control system.