Chief of Staff
Route a multi-step task across the workers that are actually good at each step.
-
1 Decompose Claude Opus 4.8 940ms
Two jobs, not one: groom the backlog, then summarise the delta. The second depends on the first.
-
2 Route — 12ms
backlog-grooming scores 86.5% — above the 80% floor. No worker in this workspace covers "brief me", so that step comes back to the user.
-
3 Hand off Claude Haiku 4.5 380ms
Passed the grooming output only — 3 issues touched, 1 escalation. The next step does not need the reasoning that produced it.
Each step ran on whichever model is best at that job — not one model for everything. The worker is the workflow.
Permissions & guardrails- Read: the worker capability index, current eval scores and recent run traces
- Write: nothing directly — it only invokes already-deployed workers and records each hand-off
- Escalate: the user when no worker clears the score threshold for a step, or when a step has no owner
- Never: deploy, modify or configure a worker — it routes to existing ones only
- Never: route a step to a worker scoring below 80% on its own eval
- Routed step 1 to backlog-grooming (86.5%). Completed.
- Refused step 2 — no deployed worker covers briefing. Escalated instead of guessing.
- Handed off 180 tokens of output, not 4,200 tokens of transcript.
79.2%
a general model gets 66.4% · +12.8pp
66.3 · 70.2 · 73.6 · 76.3 · 78.1 · 79.2
private rules were taught in one workspace and never leave it. general rules are true about the work, so every copy of this worker gets them.