You can spin up a dozen agents before your coffee’s cold. You’ll also find out shortly that you can only really run about four at once. Because every agent needs your supervision.
An agent produces a draft, a PR, a plan. None of it is value yet. It becomes value the moment a human trusts it enough to ship it, and trust is a human act. When making a mistake is cheap, you check one sample and let the rest run untrusted. But usually the work worth pointing agents at is the work where a mistake is expensive - and there you have to trust each output on its own, one at a time. So before the output counts, someone reviews it, redirects the agent when it drifts, or reloads the context when the agent loses the thread. Every running agent therefore has a human supervisor attached, and that supervisor is you.

Supervision is attention
Review, redirect, reload: on all three of those, you spend one resource - attention. The reload is the expensive part. Watching an agent play well costs little, but the moment it goes sideways you have to rebuild the whole task in your head to steer it back, and that reconstruction is what drains you.
So the real cost of running an agent is the attention that is demanded by the supervision. You can partially offload the routine of it onto gates and reviewer-agents - but the last call, ship or don’t, lands on one human, and that call is attention-heavy. Which raises the only question that matters next: how much attention do you have?

Human attention is finite, and it’s been measured
We don’t have much of it, and the numbers converge on how little. In 2007, researchers measured how many robots one operator could actively control at once: 4 to 6; effectiveness thinned above that (Crandall & Cummings, 2007). UAV work found the same shape from the other side: operators could monitor about 15 drones but actively control only ~3 (Porat et al., 2016). Watching is cheap; driving is capped - and every agent you steer is mostly driving.
That 4-6 isn’t hardwired in your brain. It’s what active control costs you at a short feedback loop - hands-on, checking often. Lengthen the loop checking interval and the number of agents climbs; that’s lever one, below. But it climbs against something that doesn’t stretch: the attention each review cycle spends.
Nearby limits do cluster in the same numeric neighborhood - working memory holds about four chunks (Cowan, 2001); sustained intense cognitive work tops out at 2 to 4 hours a day (Ericsson et al., 1993); monitoring vigilance decays within 30 minutes (Frontiers, 2025); and switching off an unfinished task leaves residue that leaks into the next (Leroy, 2009), so every paused agent is a half-open loop. Suggestive of a shared ceiling - not proof of one constant; these are different instruments landing near the same number, which is worth noticing.
You could argue even 4-6 is just economics - the seventh agent isn’t worth its coordination cost, no biology needed. But the tell is physiological: people don’t stop at four agents feeling fresh, they stop feeling fried - some have started calling it AI brain fry. Missing out on value doesn’t exhaust you; attention running out does. So the finite input is biological - and that’s the part that doesn’t move, even though we can buy more intelligence every day.

Your attention holds still, even though the models evolve
Total value from your agents is plausibly bounded by your attention budget - not by model quality, not by how many agents you can afford to run. And total value is just that budget in action: attention spent per agent x agents run.
Even though the best models now hold a coherent task for over five hours unattended, they finish the task only about half of the time (METR, Jan 2026, as of July 2026). Five hours is the coin-flip frontier - half of those runs fail. The task length you can actually trust is far shorter. Minutes, not hours. Keeping us more or less in the same place.
We are still bounded by the feedback loop interval. But we all somehow feel that smarter models should be able to sort more things on their own, so the feedback interval will grow. Not sure that’s the case now. But it definitely will be in the future.
Then what changes would be the scope of our goal. We will just provide higher level goals to the models. Our attention will still be required. And on the top of that, we have to come up with those higher level goals. And that’s truly homework - you have to precisely know what you want.
Two main levers move the productivity number
If value is attention spent per agent x agents run, then for one person, mainly two levers move it: spend less attention per agent, or have more attention to spend. Hiring more supervisors is a third path, let’s skip that for now.

Lever one: spend less attention per agent
Two moves stretch how long an agent runs unattended without the output going sideways.
Your playbook: ways of working plus the skills built with the agent over time - accumulated prompts, project conventions, reusable setups - so you’re not re-explaining context every session. The compounding lives here.
Verification gates: a way for the agent to check its own output before it reaches you. In coding that can be all sorts of tests - unit, integration, and the rest; in writing it could be a pass that asks ‘is this AI slop, yes or no’ and won’t release the draft until it clears. Both push work off your attention onto the loop.
The rule that falls out: the longer the feedback loop you can trust, the more agents you can lead. The math is simple. Say one agent works 15 minutes on its own, then needs 5 minutes of you - review, redirect, done. (That unattended stretch is more like 5 to 15 minutes in practice - take 15 as the good case.) While one agent works alone, you have 15 free minutes - enough to serve three more agents at 5 minutes each. Three waiting plus the one in your hands: four agents.

Now stretch the same agent to 30 unattended minutes: 30 ÷ 5 (serving time per agent) = 6 waiting, plus one in your hands, seven agents. Nothing got smarter - the loop got longer. And that’s exactly what the playbook and the gates do: makes the agent run longer without you.
This is how the 4-6 moves: every hour of trusted autonomy you add, every touch you remove, buys room for another agent - until a reload hits, and rebuilding that context in your head pins you down again. Better tooling - cleaner diffs, a review UI that surfaces what actually changed, one dashboard instead of ten terminals - shrinks the touch-time too, but not to zero: every decision still passes through one human judgment, and that floor is the part tooling can’t refactor away. Smarter models help - better results per run, often longer runs too - but autonomy duration is the term that grows your agent capacity most.
Of the two moves, gates matter more than they look. Your own judgment is the one instrument you can’t audit from the inside: the more an agent’s output looks right, the less you check it - exactly when a plausible-looking error on expensive work is most costly. A gate doesn’t get tired and doesn’t get impressed. It catches what your confidence won’t.

Lever two: have more attention to spend
The budget you’re spending is biological, so you raise it the way you raise anything biological - through the body. Two inputs, in order.
Sleep first - and nothing else here is even close. Short yourself on sleep and the first thing to go is sustained attention, the exact function agent supervision runs on. It fails before memory, before mood, before you notice. Total sleep deprivation drives attention lapses at g = −0.76 (Lim & Dinges, 2010), the largest hit in the whole test battery (−0.76 is large: clearly and reliably worse, not a rounding error); everyday partial restriction still costs around g = −0.41 (Lowe et al., 2017). And the deficit hides from you.
Two weeks of 6-hour nights leaves you as impaired as two full nights without sleep, and the subjects were largely unaware of it (Van Dongen et al., 2003); at 17-19 hours awake - a late night stacked on a normal day - you perform at a 0.05% blood-alcohol level (Dawson & Reid, 1997). Sustained attention fades first, and it’s the exact function you use to judge whether you’re still sharp - so the instrument reading “fine” is the broken one.
Slow cardio second: walking, rucking, biking, rowing, swimming. The measured effect of chronic aerobic training is modest - attention up about g = +0.16, executive function g = +0.12 (Smith et al., 2010). The study pinning zone-2 to workday attention span doesn’t exist yet; I run it on myself anyway.
Sleep → attention is proven; attention → supervision is the capacity math above; the full chain, health → daily agent-driving throughput, is a well-grounded inference, labeled as one.

The consequence: find a balance you can actually hold
Follow the chain to its end and the moat is obvious. Intelligence you can rent by the hour. Attention oversees everything you rent, and cannot be bought. And the return on it just went vertical: a rested hour once bought about an hour of your own work; now it gates the output of an agent that ran five hours while you were away - the same hour, worth several times more, and the multiplier compounds while your attention stays flat. So the game was never exhausting all of your attention in a two-week sprint - it’s holding a level you can sustainably keep.
That leaves two honest ways to work. Conductor: scale the number of agents, and fund the recovery that many agents demand. Craftsman: run one or two deep streams and make them better. Both work.
Key takeaways
- Drive 4-6 agents on a fast feedback loop - grow the number by lengthening the loop, not by muscling through. Active control caps around 4-6 today; a playbook, gates, and more autonomy raise the feedback loop interval.
- Your playbook is how you make a difference - the longer the feedback loop you can trust, the more agents you can lead.
- Autonomy duration grows your agent capacity more than raw model IQ - twenty more unattended minutes beats a marginally smarter model that interrupts you just as often. Smarter models usually buy more autonomy, so the two move together - but it’s the autonomy doing the work, not the IQ.
- Sleep is the biggest biological lever - it’s the g = −0.76 hit on sustained attention, the exact function every rented hour draws from, and the one most people spend first.
Where the numbers come from
All sources accessed 2026-07-08, links re-verified 2026-07-10.
- Fan-out / supervisory control (4-6 robots active): Crandall & Cummings, IEEE Transactions on Robotics, 2007
- Active-vs-passive supervision (~3 active / up to 15 monitored): Porat et al., 2016
- Working memory ~4 chunks: Cowan, Behavioral and Brain Sciences, 2001
- Attention residue from unfinished tasks: Leroy, OBHDP, 2009
- Sustainable intense cognitive work 2-4 h/day: Ericsson et al., 1993
- Vigilance decrement within ~30 min: Frontiers in Cognition, 2025
- Agent autonomy time horizon (320 min at 50% success; doubling ~89 days since 2024) - as of July 2026: METR Time Horizon 1.1, Jan 2026
- Sleep deprivation → sustained attention (g = −0.76): Lim & Dinges, 2010
- Partial sleep restriction → attention (g = −0.41): Lowe et al., 2017
- 14 days at 6h ≈ 2 nights total deprivation, unnoticed: Van Dongen et al., SLEEP, 2003
- 17-19h awake ≈ 0.05% BAC: Dawson & Reid, Nature, 1997
- Chronic aerobic training → attention/EF (g ≈ +0.16 / +0.12): Smith et al., 2010






Leave a Reply