Research note 12 min read

Delegation, lock-in, and the shrinking verifier class

When rational delegation meets measured pipeline collapse — and the literature refuses a single verdict.

  • Labor & institutions
  • Human–AI interaction
  • Learning & cognition

A familiar story goes like this: you delegate a little more each week — drafting, checking, coding, summarizing — because the marginal cost of doing so falls toward zero. Individually, that looks like rational adaptation. Collectively, it may be something else: a path toward sociotechnical lock-in, a shrinking pool of people who can verify what machines produce, and institutions that keep cheaper automation even as the long-run costs compound.

We did not set out to write a doom note. The literature on AI and work is better read as an empirical debate with live tensions than as a settled consensus. Some mechanisms are now formally modeled; some labor-market patterns are measured; some cognitive claims are actively disputed. The point of this note is to lay out those ingredients — and leave the verdict to the reader.

The delegation–verification dilemma, formalized

The complementarity story — humans and AI as collaborators, each doing what they do best — dominates public framing. A 2026 paper by Angjelin Hila, The Human-AI Delegation-Verification Dilemma, argues that this framing is incomplete.

Hila models how agents adapt delegation strategies over time, then asks what happens when individually stable choices aggregate. In the absence of communicative safeguards or institutional norms, mass delegation can produce sociotechnical lock-in: a macro-level equilibrium modeled as a collective-action problem (a prisoner's dilemma) in which shared epistemic standards degrade even as individual payoffs from delegating remain attractive in the short run.

That is close to a formal version of a claim many practitioners feel in their bones: once you dip, you AI — not because anyone is lazy, but because verification is costly, delegation is rewarded, and the feedback loop that would correct over-delegation is weak. Hila's analysis suggests mitigation requires more than better models: local signaling, communicative commitments, or institutional norms that make verification socially obligatory, not merely personally optional.

A related formal line — Delegation and Verification Under AI — shows that even fully rational workers can over-delegate when institutions evaluate outcomes rather than process. Small differences in verification ability can produce sharp phase transitions: some workers become institutionally more valuable with AI access; others degrade through rational reduction of oversight. No behavioral bias required.

Task displacement, reinstatement, and the junior pipeline

Daron Acemoglu and Pascual Restrepo's task-based framework (Automation and New Tasks, 2019) gives language for the economic mechanism. Production is a bundle of tasks allocated to labor or capital. Automation creates a displacement effect — capital substitutes for labor in tasks humans used to perform. New tasks can create a reinstatement effect — labor returns in activities where humans retain comparative advantage.

The framework's policy implication is uncomfortable: automation can raise productivity while reducing labor's share and, under some conditions, net employment. Whether that happens depends not on whether AI is "good" but on whether new labor-intensive tasks appear fast enough to offset displacement. That is an empirical question, not a philosophical one.

Recent data suggest displacement is concentrating at the entry point of careers, not uniformly across experience levels. The 2026 Stanford HAI AI Index reports that employment for software developers aged 22–25 fell nearly 20% since 2024, while older developers in the same firms saw headcount hold or grow. Employer surveys point to further change ahead — including anticipated reductions in software engineering and service operations even where aggregate employment has not yet collapsed.

Erik Brynjolfsson, Rishi Chandar, and Ryan Chen (Stanford Digital Economy Lab, 2025) analyze high-frequency payroll data and find a relative employment decline for workers aged 22–25 in the most AI-exposed occupations — on the order of 13%, holding after firm-level controls — while less-exposed occupations and more experienced workers in the same fields remained stable or grew. Adjustment appears to run through fewer jobs, not just lower wages.

Christian Catalini, Xiang Hui, and Jane Wu (Some Simple Economics of AGI, 2026) name the structural consequence: the Missing Junior Loop. Human expertise is a stock built through the friction of routine execution. Automate entry-level cognitive work and you do not merely save cost — you foreclose the pipeline through which future verifiers, managers, and domain experts would have formed. The professional not hired in 2026 is not in the running for senior leadership in 2038.

That directly supports a claim we hear often from institutions: you cannot simply hire or train a replacement workforce on demand if the junior cohort that would have become that workforce is no longer entering the pipeline. The pool is not just shifting skills. In exposed domains, it may be shrinking at the base.

Moral crumple zones: accountability without control

The "crossing guard for agents" metaphor has an academic cousin. Madeleine Clare Elish introduced moral crumple zones (2019): when automated systems fail, responsibility tends to crumple onto the nearest human operator — even when that person lacked the information, authority, or time to prevent the failure.

Like a car's crumple zone, the design protects the integrity of the technical system at the expense of the human in the loop. The operator becomes a liability sponge. In aviation, nuclear control, and early autonomous systems, Elish shows how accountability can track symbolic human oversight rather than effective control.

As agents enter drafting, hiring, clinical support, and compliance workflows, the same pattern is easy to imagine: a human "in the loop" who cannot realistically verify model output at scale, but who absorbs blame when output fails. Ben Green (2022) documents a related failure mode in government algorithmic oversight: mandated human review that provides legitimacy without capacity — presence is not control.

If delegation scales faster than verification authority, moral crumple zones are not a edge case. They are an institutional default.

Cognitive offloading: atrophy, reorganization, or both?

Automation complacency — reduced vigilance as systems prove reliable — has a long literature in human factors (Parasuraman & Manzey, 2010). Newer work extends this to human–AI collaboration. A 2025 review in AI & Society on automation bias emphasizes that explainability alone often does not restore appropriate skepticism; verification effort and engagement matter more than transparency theater.

Neuroscience adds texture without settling the normative question. A 2026 NeuroImage study using the N2pc component — a marker of selective attention — found that reliable AI reduced monitoring (smaller N2pc amplitudes) compared with unreliable AI, tracking implicit trust calibration in a visual search task. Brains offload attention when automation looks competent. That is not proof of permanent cognitive decline. It is evidence that trust and attention move together — which matters when competence is misestimated.

The deeper dispute is whether offloading atrophies human capability or reorganizes it — cognitive extension rather than cognitive loss. That dispute is not closed. Human intelligence has long used tools that store and structure knowledge externally. The open question for high-stakes verification is whether current delegation patterns leave enough repeated, consequential practice for evaluative skill to form — especially among people who never received the junior loop in the first place.

We think the honest posture is to hold both possibilities: offloading can be adaptive and the institutional loss of verification capacity can be real.

Social substitution: a sadder, more defensible claim

The narrative that "AI is eating our social fabric" is simpler than the best current longitudinal evidence.

Dunigan Folk and Elizabeth Dunn (Psychological Science, 2026) followed more than 2,000 adults across four Western countries for 12 months, modeling bidirectional links between social chatbot use and loneliness.

The results depend on measurement — which itself is instructive:

  • With a narrow single-item measure of emotional isolation, increased chatbot use predicted later increases in loneliness (consistent with a "potato chip" pattern: comforting in the moment, costly over time).
  • With a broader 20-item social-connection scale, the cleaner finding ran the other way: feeling less connected predicted subsequent increases in chatbot use, while chatbot use did not significantly predict decreases in social connection on that broader measure.

The authors urge caution — the analyses are partly exploratory — but the shape of the evidence is not "AI replaces friends, full stop." It is closer to the lonely turn to AI, and narrow measures suggest it may not help much — and might make isolation feel worse. That is a harder story to sloganize. It is a better hook for reasoning: which direction do you think dominates, and for whom?

What we think is solid — and what we will not claim

Well supported today:

  • Formal models of delegation-driven lock-in and verification failure (Hila, 2026; related verification-under-AI models).
  • Measured concentration of employment loss among young workers in AI-exposed roles, with multiple independent data sources converging.
  • Explicit recognition of pipeline / apprenticeship collapse as a structural risk (Catalini, Hui, & Wu, 2026; Brynjolfsson et al., 2025).
  • Accountability misallocation toward humans with limited control (Elish, 2019).
  • A contested but active literature on automation bias, attentional offloading, and verification effort.

Complicated — keep the tension visible:

  • Whether cognitive change constitutes decline or reorganization.
  • Whether social chatbot use causes loneliness broadly, or primarily selects for and fails to repair pre-existing disconnection — with measure-dependent answers.
  • Whether reinstatement of new labor-intensive tasks will arrive fast enough to offset displacement in the current wave.

We are not asking readers to pick a team. We are asking what happens when all of these feedback loops run at once.

A prediction question, not a verdict

Institutions face four interacting pressures:

  1. Individual rationality toward delegation (lock-in equilibria).
  2. Labor-market displacement at the junior tier (shrinking verifier stock).
  3. Accountability architectures that protect systems more than humans (crumple zones).
  4. Social substitution among people already disconnected (loneliness-driven adoption).

Acemoglu and Restrepo remind us that the long-run outcome depends on new tasks — including tasks that rebuild verification, apprenticeship, and judgment. Catalini and coauthors ask whether synthetic practice and accelerated mastery can replace what entry-level repetition used to provide. Folk and Dunn remind us that psychological harm may track who was already isolated, not only who adopted first.

So we end with a question we are actively working on, and welcome collaborators to stress-test:

Which of these feedback loops closes first — delegation lock-in, pipeline collapse, accountability misallocation, or social retreat — and does the order matter for whether institutions can still tell who knows what?

If that question fits your mandate — measurement, labor economics, learning systems, or institutional design — start a conversation.


Selected references