For the entire history of knowledge work, the expensive step was making the thing. Writing the memo, drafting the contract, cutting the analysis, building the deck — that is where the hours went, and judging the result rode along almost for free. You produced something slowly and deliberately, and the deliberation was baked into the production. Care and output were the same motion. Every process we have — every review cycle, every approval chain, every inbox — was designed around that fact.

That fact is gone. Generation is now free and effectively instant. And the moment production stops being the constraint, judgment becomes the only thing that is. This is not a small adjustment. It is an inversion of the very thing our workflows were built to conserve, which means the processes we trust are now optimized for a scarcity that no longer exists. They protect production effort — the cheap thing — and spend judgment recklessly, because they were never designed to protect it.

The old world was short on production and long on judgment. The new one is exactly the reverse — and our processes never got the memo.

I

The slop grenade

You can watch the mistuning happen in real time. The most visible symptom got a name this year. Shopify’s CEO Tobi Lütke — who a year earlier had made AI use a “baseline expectation,” telling staff to prove they couldn’t do a task with a model before asking for more people — sat down on a podcast and named the thing his own mandate had produced. He called it a slop grenade: AI-generated material, unread and unvetted, lobbed at a colleague who now has to read, verify, and repair what the sender never took responsibility for.1

His example is almost too perfect. An employee uses a model to inflate a short thought into a long email; the recipient runs it through a different model to compress it back down. Volume up, comprehension flat, net work higher. The machine ran twice and moved nothing.

What makes Lütke a specimen rather than an anecdote is the direction of the reversal. This is not a skeptic warning from the sidelines. Shopify’s internal agent handles something like half of its production code changes; the man is not anti-AI in any sense that matters. He is the loudest adopter in the room, documenting his own backfire on the record. When the person who pushed hardest is the first to describe the failure mode, the failure is structural, not cultural.

And it is measurable. Researchers at BetterUp Labs and Stanford’s Social Media Lab gave the phenomenon a name — workslop — and a number.

The cost of crossed streams

40% of desk workers received AI “workslop” in the last month — polished-looking output that looks finished but isn’t, displacing effort rather than doing it.2
1h 56m average time to sort out each incident — roughly $186 per employee per month, an invisible tax on self-reported salaries.
$9M a year in lost productivity for a 10,000-person company — before the downstream hit to trust and morale.

Source: BetterUp Labs & Stanford Social Media Lab, survey of ~1,150 U.S. desk workers, Harvard Business Review, Sept 2025.

Duolingo’s Luis von Ahn made the same turn in the same window — from AI-first mandate to a demand for human oversight. Two of the loudest voices, pivoting to the same correction at the same time. That is a pattern, not a coincidence. And Stanford’s Jeff Hancock, who ran the workslop study, names the driver directly: the trouble starts with “indiscriminate AI mandates, when leaders encourage people to ‘use AI everywhere’ without clear guidance.”3

Be precise about the failure, because the obvious reading is the wrong one. The problem is not that a machine wrote the thing. The problem is that generation and consumption got wired directly together, human to human, with no layer between them to catch that nobody had taken responsibility for it. The output moved at machine speed; the judgment it needed still moved at human speed; and the gap between those two speeds landed on someone’s desk as work.

II

A faster horse, or a car

So the interesting question is not “how do we get people to be more careful.” Carefulness is a moral answer to a structural problem, and moral answers don’t scale — they are a tax on individual willpower that erodes the moment anyone is busy. A mandate to use AI with no redesigned target is not leadership either; it is ambition without direction, and it manufactures the very slop it later laments. The real question is structural: what does an operation look like when it is designed for the new scarcity instead of the old one?

Two heavy adopters, the same year, opposite answers — and the gap between them is the whole argument. The first pushed AI onto the organization he already had. Everyone got a copilot; the structure stayed exactly as it was, and only the speed demanded of the people inside it went up. That is a faster horse, and the slop grenades were the direct result.

The second asked the harder question: what does AI let an organization fundamentally stop doing? In “From Hierarchy to Intelligence” (March 2026), Block’s Jack Dorsey and Sequoia’s Roelof Botha argued that the corporate hierarchy is a two-thousand-year-old information-routing protocol — built to move context up and decisions down under the limits of human cognition — and that AI can now do that coordination directly. Block recut itself around the answer rather than layering AI on the old chart.4 Whatever one makes of the specifics, the move is the point: he treated AI as a reason to redesign the system, not a tool to bolt onto it.

Using AI inside the existing structure makes the horse faster. Redesigning the structure around what AI makes possible builds the car.

Most organizations cannot cross that gap — not because the technology isn’t ready, but because the org chart won’t let them. And it is worth being clear-eyed: Block’s particular prescription is contested. Cutting deeply and pushing people to the edge of an AI-centered core assumes the main thing people contributed was routing — when people also generate and interpret the signal an organization runs on. By that reading the bottleneck was the structure, not the people in it. Take the transferable lesson, not the blueprint.

This also explains why the mandate is the default and the redesign the exception: most chief executives are neither trained for the shift nor structurally free to make it. The training gap is real — today’s leaders came up under the old scarcity, where the job was to allocate scarce production capacity, not to redesign an operation around machine-speed generation. The authority gap is just as binding — a hired CEO rarely holds the decision rights or the risk tolerance to recut deeply.

The leadership gap, measured

72% of CEOs now say they are the primary decision-maker on AI — double last year. Ownership has moved to the top.5
15% are BCG’s “Trailblazers” — decisive, investing at scale, upskilling broadly. The other 85% are waiting for proof or moving with the market.
15% the share generating meaningful value from AI — the same slice. Owning the decision is not the same as knowing how to make it.

Source: BCG AI Radar 2026 (survey of 2,360 executives incl. 640 CEOs); “AI for CEOs,” BCG, 2026.

Dorsey is an outlier on both axes at once: a founder with the control to act and the temperament to think past the existing chart, moving from financial strength rather than distress. Treating him as the template misleads, because most leaders have neither of his freedoms. Which is the whole case for a method — a way to reorient incrementally and reversibly, for the great majority who cannot cut forty percent of the company on a Tuesday.

III

The operating model

An operation designed for the new scarcity protects judgment the way the old one protected production effort. Three principles define it.

1 · Separate the streams
There are two economies inside any operation now, running at different speeds. An agent-to-agent stream carries output machine to machine — one system generating, another critiquing, a third revising, none of it needing a human to move it along. A human-to-human stream is slower and more valuable, and it should carry judgment: intent, framing, reviewed decisions. The slop grenade is what crossing the streams looks like. Keep output in the machine lane; reserve the human lane for what only humans can carry.
2 · Share the judgment, not the output
The highest-leverage move available, and most organizations have it backwards. When a team encodes what “good” means — a spec, a set of guardrails, a context file that tells the agents how this organization thinks — and shares that, it makes judgment portable. Encoded once, it steers every downstream generation and propagates to other people without anyone re-deriving it. Circulating the files that instruct your agents is not shortcut-taking; it is the reviewed, deliberate, human part of the work made reusable. The unit of exchange between humans should be the judgment. Let the agents exchange the output.
3 · Gate by consequence
Between the agent stream and any human sits a decision about whether a human is needed at all — and most of the time the answer is no. A gate that routes by blast radius sends low-consequence output onward automatically under monitoring, samples the middle tier, and pulls a person in only where stakes or ambiguity require a call. Reviewing everything is not the safe alternative; it is the same failure in a responsible costume. Real assurance means a human owns the envelope — the standard, the monitoring, the authority to gate up — and answers for whether it is real. Verification moves from a pre-release snapshot to a continuous property of the running system.

Put the three together and the shape resolves. Humans live at the two ends of the chain — framing at the front, where deciding what to point the agents at is a creative act in an engineering coat, and ownership at the back, where someone answers for whether the result is fit. Agents fill the middle, fast and tireless and indifferent to whether any of it is any good. The gate is the load-bearing wall. It is the gate that dissolves applied one level in from the ship-to-production edge: the same tiering by blast radius, now guarding the hand-off to a colleague.

IV

Target, backcast, pilot, redesign

Reorientation is a discipline, not an improvisation — and it is precisely the part a mandate skips. Any leader who genuinely understands vision and strategy arrives at the same sequence:

Set the target
Start from a clear picture of what the organization is for and where it is going — the strategic aim, not “use more AI.” The target is exactly what a mandate lacks; without it, speed has no direction and every generated artifact is motion mistaken for progress.
Backcast
Work backward from that target to the present, mapping the capabilities, decisions, and structures the future state requires — rather than forecasting forward from today’s org chart, which only ever yields a faster version of what already exists.
Pilot
Prove the redesigned way of working on a bounded slice before betting the operation on it. The pilot is where the streams actually get separated, the gate gets tuned, and judgment utilization gets measured for real: cheap to be wrong, fast to learn.
Redesign, systematically
Only then propagate the change deliberately — role by role, workflow by workflow — encoding the judgment as you go, so the redesign compounds instead of scattering.

The contrast with the mandate is total. A mandate starts at the last step and skips the first three: it orders a change in behavior with no target, no backcast, no pilot — which is why it produces slop instead of transformation. The sequence is what lets the great majority who are not founders reorient incrementally and reversibly, without a single irreversible bet.

V

Judgment utilization

There is a single number that tells you whether an operation has made the turn: the share of its people’s time spent at genuine decision points, versus the share spent triaging and repairing machine output. Call it judgment utilization.

Where the human hours actually go Tuned for the old scarcity 25% decisions 75% triaging & repairing output Tuned for the new scarcity 70% at genuine decision points 30% triage The gain is not more hours worked — it is a smaller triage share. judgment utilization = time at decision points ÷ total human time
Judgment — genuine decision points Triage — reconstituting machine output
Judgment utilization is the share of human time spent deciding rather than reconstituting. Volume up with the triage share up is not progress — it is the bottleneck moved from production, where machines can help, to judgment, where they can’t.

In most organizations right now that number is falling, and the fall is invisible, because output volume is up and up looks like progress. It isn’t. An operation generating ten times the material while its people spend their days deciding whether that material is true has not gotten faster. It has moved the bottleneck from production, where machines could have helped, to judgment, where they can’t — and then flooded it. Judgment utilization is the metric that makes the invisible loss legible, and the target the operating model is built to raise: not by working people harder, but by keeping them out of the stream everywhere a judgment call isn’t actually required.

VI

You can’t move humans at machine speed

This is the deepest version of the error, and the one worth carrying out of all of it. If you are trying to move humans at agentic speed, it will never work. A person cannot be accelerated to match a machine, and every attempt to do it just relocates the bottleneck onto the person and calls the resulting pile-up productivity. The answer is not to speed the human up. It is to reorient the workflow so the human is not in the fast lane at all — present where judgment is required, absent everywhere it isn’t.

None of this is a case against AI, and it should not be read as one. The organizations making the mistake are, almost without exception, the ones that adopted hardest — the believers, not the skeptics. That is the tell that the problem is architectural rather than attitudinal. You cannot mandate your way out of it, because the problem was never the technology — and the mandate to generate more is the thing producing the flood. You can only design your way out: build the operation so that agentic output is owned before a human ever consumes it, so the scarce resource is spent only where it is scarce, so the machines carry the volume and the people carry the calls.

This is the same shape the rest of the canon keeps arriving at. The loop moved upstream, to spec and judgment. Capability is discovered, not specified, so you bound authority rather than what a system can do. And the human factor — creativity and judgment as the constant — is what those arguments look like when you stop talking about systems and start talking about people. Here it becomes a design instruction: protect the judgment, because it is the only thing left that is scarce.

The era we’re entering does not reward the operation that generates the most. It rewards the operation that spends its judgment best.

That was always going to be the test. Production was the cheap thing even when it felt expensive; what an organization was really made of was the quality of the calls it made and who was willing to answer for them. AI just removed the disguise. That is a design problem — and it is the one worth solving.

Related on this site: The Human Factor · The Loop Moved · The Gate Dissolves · Capability Is Discovered · Nobody Owns the Decision · The Seam. The framework in full: Human · Org · Tech.

  • Tobi Lütke (Shopify CEO), on The Knowledge Project podcast, reported in Fortune and Business Insider, Sept 2026 — “slop grenades,” the round-trip email example, “humans take responsibility,” and the note that Shopify’s internal agent handles a large share of production code. AI use as a “baseline expectation” is from his 2025 all-hands memo. fortune.com
  • BetterUp Labs & Stanford Social Media Lab, “AI-Generated ‘Workslop’ Is Destroying Productivity,” Harvard Business Review, Sept 2025. Survey of ~1,150 full-time U.S. desk workers: ~40% received workslop in the prior month; ~1h56m to resolve each incident; ~$186/employee/month; >$9M/year for a 10,000-person org. hbr.org · betterup.com/workslop
  • Jeff Hancock (Director, Stanford Social Media Lab), interview with UNLEASH, 2025 — on “indiscriminate AI mandates” that tell people to “use AI everywhere” without clear guidance as the driver of workslop. The Duolingo / Luis von Ahn oversight pivot is reported alongside Lütke’s in coverage of the same period. unleash.ai
  • Jack Dorsey & Roelof Botha, “From Hierarchy to Intelligence,” Block / Sequoia, March 31, 2026 — the thesis behind Block’s reorganization toward “a company organized as intelligence rather than hierarchy.” Cited here for the altitude of the move — a thesis formed and owned from the top — not as an endorsement of the ~40% workforce cut or the specific role design. block.xyz
  • BCG AI Radar 2026, “As AI Investments Surge, CEOs Take the Lead” (survey of 2,360 executives across 16 markets, incl. 640 CEOs): 72% of CEOs the primary AI decision-maker (double 2025); ~15% “Trailblazers.” The “only ~15% generating meaningful value” figure is from BCG’s companion piece “AI for CEOs,” 2026. bcg.com
On the numbers The workslop figures are BetterUp/Stanford’s, first reported in HBR (Sept 2025); across write-ups the “received last month” share is stated as ~40–41% and the per-incident time as 1h56m. The CEO archetype and value figures are BCG’s. The reading that ties all of this to a single cause — production made free, judgment left scarce, processes still tuned for the old scarcity — is the argument of this site, not a claim in any of the sources. Lütke and Dorsey are used as opposite specimens of one lesson; neither is endorsed wholesale, and Block’s workforce decisions in particular remain contested. Two independent adopters reaching the same correction is corroboration, not coordination.