The Apprenticeship Crisis: AI and Entry-Level Jobs
The bottom rung was never only work. It was the training ground. Automate it and you remove the mechanism that manufactures the seniors you will be trying to hire in 2032.
Companies are automating the exact work their future senior people were supposed to learn on.
That's the apprenticeship crisis. AI is genuinely good at the bottom rung of white-collar work: the research pass, the first draft, the reconciliation, the boilerplate function, the deck that summarises the other decks. And the bottom rung was never only work. It was the training ground. It was where someone did a task badly two hundred times, got corrected, and slowly turned correction into judgment. Automate it and you don't just remove a salary line. You remove the mechanism that manufactures the seniors you'll be trying to hire in 2032. And the bill doesn't arrive for years, which is exactly why nobody's paying it.
This is not a hiring-market story. It's an org-design story. Every company making this trade is making it rationally, one company at a time, and the cost lands on all of them together.
The bottom rung was a training ground disguised as a cost centre
Think about what a first-year analyst, a junior developer, or a graduate marketer actually produced. Some of it was valuable. A lot of it was economically marginal. Work a senior person could have done faster and better, handed down anyway.
Handed down because it was pedagogically dense. The junior version of a task is where the reps live. You build the model wrong, someone shows you why the assumption was load-bearing. You ship the function, it breaks in a way you didn't predict, and now you understand the system's edges in a way no documentation transfers. Judgment is not knowledge. It's compressed correction... a stack of times you were wrong about something specific.
Firms carried that inefficiency because it converted into capability. The junior's salary was partly tuition, paid by the employer, in exchange for a claim on the senior they'd become.
AI breaks the exchange, cleanly. The tool now does the marginal work faster than the junior, at a fraction of the cost, without the two years of variable output. Judged as a task, it's not close. The task was never the point... and the accounting was never set up to notice.
Will AI really eliminate entry-level jobs, or does it just feel that way?
The honest answer: the entry level is contracting sharply, the evidence is strongest in the most AI-exposed occupations, and "AI did it" is a reasonable read of the shape rather than a proven cause.
Stanford's Digital Economy Lab, working with ADP payroll data covering millions of US workers, tracks this directly. In the August 2026 revision of Canaries in the Coal Mine?, Erik Brynjolfsson, Bharat Chandar and Ruyu Chen report that employment for workers aged 22–25 in the most AI-exposed occupations sits 19% below where it would be had it kept pace with their less-exposed peers. A gap that has widened steadily since August 2025. Two details matter more than the headline. First, experienced workers in those same occupations don't show the pattern; it is specific to the young. Second, the gap runs through reduced hiring, not increased separations. Nobody is firing juniors. They're quietly not opening the role.
The authors are careful to call these descriptive indicators, canaries, not causal estimates. Take that caveat seriously. But the sector data has the same shape. SignalFire's State of Tech Talent Report 2026, built on its Beacon platform tracking 650M+ individuals, finds new-grad hiring down roughly 65% at the tech majors and around 76% at early-stage startups versus 2019.
And the outcome data agrees. The New York Fed's tracker for recent college graduates put unemployment at 5.6% in Q2 2026 with underemployment at 42%. That's 42 out of every 100 recent graduates working a job that doesn't require the degree they just bought.
Interest rates and the post-2021 over-hiring correction explain part of this. They don't explain why the damage concentrates in the young, in the exposed occupations, and in hiring rather than firing. That combination is what a company looks like when it decides the first rung isn't worth funding any more.
The cost that never appears on the budget
Here is the asymmetry that makes this crisis structural rather than a bad quarter.
A junior salary is a line item. Judgment transfer is not. One of them shows up in a spreadsheet the CFO reviews monthly; the other shows up as an absence, five to ten years out, in someone else's tenure. So the trade looks free at the moment you make it, and it stays looking free for the entire period during which you could still reverse it.
What you're actually doing is spending a stock you've stopped replenishing. Your senior people are that stock. Every one of them was produced by an apprenticeship system somebody paid for, often a previous employer. Turn off the intake and you don't feel it while the stock holds. You feel it when the stock leaves: retirement, a competitor, a founder itch. Then you go to the market for senior judgment and discover every other company defunded their intake at the same time, for the same good reasons, and the price of the thing nobody is producing has gone where prices go.
There's no cartel here and no villain. Just a lot of individually correct decisions producing a shortage nobody chose. It's the most predictable failure mode in the future of work, and it's already priced into exactly zero of the AI business cases I've reviewed.
This is the Hourglass Collapse, one floor down
I've written before about the Hourglass Collapse. The argument: AI automates the narrow middle of the org chart faster than the wide top or wide bottom, because middle management is fundamentally an information-routing technology and AI routes information better. The hourglass becomes a lens: wide top, wide bottom, nothing in between.
The apprenticeship crisis is the same event, seen from the bottom.
Because the middle layer wasn't only routing information (the half of management that goes, per The Third Great Restructure). It was also the ladder. Those coordination roles were where an operator became a manager, where a manager saw the whole system for the first time, where the person who was good at the task learned what the task was for. Collapse the middle and you don't just delete a layer of cost. You delete the staircase between the two layers you kept.
Which produces a specific, under-discussed problem with the lens-shaped org: it has no internal path from bottom to top. A lens runs on senior judgment. It has more need of senior judgment than the hourglass did, because there's no middle layer absorbing ambiguity on the executives' behalf. And it cannot manufacture that judgment internally, because the rungs that used to do it are the exact rungs it removed.
That's a stable shape only if senior judgment is something you can reliably buy. Right now it is, and it's cheap relative to the value, because we're all still living off a stock produced by the pre-AI apprenticeship system. Ask what that market looks like in 2033. This is the part of AI org design that gets skipped: everyone models the headcount they're removing, nobody models the pipeline they're removing with it.
Are junior developers actually getting worse?
Their output got better. Their judgment is being formed less. Those aren't in tension. Confusing them is why this argument goes in circles.
Start with what changed in the artifact. GitClear's Maintainability Gap research (January 2026), which analysed 623 million code changes across 2023–2026, found block duplication climbing from 40.3 to 73.0 duplicated lines per thousand changes, moved lines, the signature of refactoring, falling from 13% of changed lines in 2023 to 3.8% year-to-date in 2026, and cross-file function calls, a reuse indicator, down 35%. Code is being added faster and integrated less. That's not a claim about who wrote it. It's a claim about what the writing process now optimises for.
Then what changed in the review. Stack Overflow's 2025 Developer Survey (49,000+ developers, 177 countries) found 84% using or planning to use AI tools, and simultaneously 46% actively distrusting the accuracy of the output, up from 31% the year before. The single most-cited frustration, at 66%, was "AI solutions that are almost right, but not quite." Not wrong. Almost right. Wrong is cheap to catch; almost-right is the expensive kind.
And what changed in the workload. Upwork's Research Institute survey of 2,500 executives, employees and freelancers found 77% of employees using AI say the tools have added to their workload, with 39% reporting more time spent checking AI-generated content. Even seniors misjudge this: METR's randomised trial of 16 experienced open-source developers across 246 real tasks found they were 19% slower with AI tools while believing they'd been 20% faster.
Now the counter-evidence, because it's real and it's strong. Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,179 customer-support agents and found AI assistance raised productivity 14% on average, and 34% for novice and low-skilled workers, with almost no effect on experts. AI compresses the output gap between a beginner and a veteran faster than anything we've measured.
Both findings hold, and the reconciliation is the whole essay: AI transfers the answer, not the judgment that produced it. A junior with a good model ships senior-looking work on day one. What they don't get is the part where they were wrong, noticed, and adjusted. They weren't wrong. The model was right, and nothing happened to them.
There's a controlled result on exactly this. Fabrizio Dell'Acqua's field experiment with 181 professional recruiters ("Falling Asleep at the Wheel") randomly varied the quality of the AI recommendations they received. Recruiters given the higher-quality AI were less accurate than those given the worse one. They exerted less effort, spent less time, and deferred to the recommendation. The group with the mediocre AI stayed engaged, learned to work with it, and improved. That learning effect was driven by the experienced recruiters, the ones who had something to check the machine against.
That is the mechanism, stated as cleanly as I can state it. A good tool handed to someone with no internal model to check it against produces good output and no learning. Hand the same tool to someone with twenty years of pattern-matching and it produces good output they can audit. The audit is the part that compounds.
Which is also the honest answer to the senior engineer's complaint. You're not getting older. You're reviewing plausible work from people who cannot yet tell you why it's right. Reviewing confident, plausible, unexplained work is genuinely harder than reviewing the obviously-junior work you used to get.
So how does anyone earn judgment now?
We have a precedent, and it's not encouraging by default.
Matt Beane, now at UC Santa Barbara, spent two years observing hundreds of surgical procedures at five hospitals and interviewing trainees at thirteen more. His finding, published in Administrative Science Quarterly in 2018: robotic surgery let the attending operate essentially alone. Open procedures needed four hands, so the resident was structurally required. The robot made them optional. Residents got dramatically less hands-on practice, and the few who became competent did it through what Beane called shadow learning: norm-bending, off-the-books practice, out of the limelight, at real cost to themselves and their profession.
Read that as the general case. The technology didn't make trainees worse. It removed the trainee from the loop, and the loop was the training.
In The Skill Code, Beane names the three ingredients that survive across every apprenticeship he studied: challenge (work near, but not past, the edge of current capability), complexity (seeing the system around the task, not just the task), and connection (a relationship with someone more expert who will actually correct you). AI removes all three by default. It removes challenge by doing the hard part. It removes complexity by handing over an answer with the reasoning collapsed. And it removes connection by making the junior unnecessary in the room where the work happens.
Default is the operative word. None of those removals is a law of physics. Each is a design decision that most companies are currently making by not making it.
What to do about it: four moves
This is org design, not a training programme. L&D can't fix it alone, because the thing that broke is the work, not the curriculum. A lens-shaped org needs more senior judgment than the hourglass did, not less. That makes the pipeline a design constraint, not an HR initiative.
That's the sequence ORBIT inverts. People before Process before Tool. It's why "we'll run an AI upskilling module" fails: it changes what people know without changing what they do.
ORBIT names five functions an AI-augmented team has to cover: Orchestrate (direct the agents toward an outcome), Run (manage the workflows the agents are inside), Build (create the systems, tools and prompts they operate within), Influence (drive adoption and culture change) and Translate (bridge agent outputs and human decisions... interpret, validate, act on what comes back).
Read that list with the apprenticeship problem in mind and one function stands out. Translate is the judgment function: the person who looks at an agent's output and says "plausible but wrong," or "right, but reframe it before it goes to the customer." That call can't be delegated to a model, and it's the function teams leave unassigned more often than any other. It is also the closest thing the lens-shaped org has to the rung it removed. So the four moves below are one move seen four ways: put juniors on Translate work, deliberately, and make someone senior own their correction.
1. Sort junior work into zones before you automate it
The 3-Zone Co-Pilot map exists for this: Replace, Augment, Refuse. The common failure is dumping all junior work into Replace because AI can do it. Replace is correct for coordination and routine generation. But Zone 2, Augment, is where apprenticeship actually lives: work where a human decides under ambiguity and AI supplies context. Staff Augment work juniors-first, deliberately, and accept the drag. That drag is the tuition line you deleted from the budget.
2. Make review the teaching surface
If AI writes the artifact, the artifact stops being evidence of judgment. So stop reviewing artifacts. Review reasoning. Require the junior to state what they decided, what they rejected, and what would make this wrong. Someone who can't answer the third question hasn't learned anything, no matter how good the output looked. This costs senior time and nothing else, which makes it the highest-return move on this list.
3. Put connection in the plan or it won't happen
Senior attention is the scarcest input in an AI-augmented org and the first thing squeezed when a lean team gets leaner. If coaching time isn't scheduled, defended and counted as delivery, it evaporates. And connection was one of the three ingredients.
4. Capture the shadow learning instead of banning it
Your juniors are already using personal AI accounts on your work. That's Beane's residents on the simulator at 2am. Punish it and you lose both the learning and the visibility. Bring it inside, look at what they're actually asking the model, and you get the single best real-time map of where your training system has holes. One structure for keeping what you capture is in the Second Brain guide.
The binary
Two versions of your company in 2033.
In the first, your senior bench is the people you have right now: older, more expensive, and smaller, because some of them left. You go to the market for replacements and find that the market defunded its intake at the same moment you did. You pay whatever the number is.
In the second, you decided which rungs you were keeping while you still had juniors to put on them. The work was slower for two years. You have people.
One of those is your company. The choice is being made this year, whether or not anyone is making it deliberately.
Read this next: The Third Great Restructure... the same decision one floor up: which half of the manager layer you keep.
FAQ
Will AI really eliminate entry-level jobs?
Not eliminate. Contract, sharply and unevenly, concentrated in the most exposed occupations. Stanford's Digital Economy Lab, using ADP payroll data through June 2026, finds employment for 22–25-year-olds in the most AI-exposed occupations running 19% below where it would be had it tracked their less-exposed peers, with no equivalent effect for experienced workers in the same roles. Crucially, the mechanism is reduced hiring rather than layoffs. Roles quietly not opened. SignalFire's 2026 data shows new-grad hiring down roughly 65% at the tech majors versus 2019. The researchers describe their findings as descriptive indicators rather than proof of causation, and rates and over-hiring corrections explain some of it. But the concentration in young workers, in exposed occupations, through the hiring channel, is what a defunded first rung looks like.
Are junior devs getting worse because of AI?
Their output is better and their judgment is forming more slowly, which is a different problem and a worse one. GitClear's analysis of 623 million code changes found duplicated blocks rising 81% since 2023 while refactoring moves fell from 13% of changed lines to 3.8%. Code added faster, integrated less. Meanwhile 66% of developers in Stack Overflow's 2025 survey name "almost right, but not quite" as their top AI frustration. The mechanism isn't laziness. Dell'Acqua's experiment with 181 recruiters showed that people given higher-quality AI performed worse than people given a mediocre one, because good recommendations remove the incentive to stay engaged. The learning effect went to the experienced participants who had something to check the AI against. A junior with a good model gets correct answers and no correction. Correction is what judgment is made of.
Am I getting old, or is working with "AI juniors" becoming a nightmare?
You're not getting old. The review job genuinely changed shape. Junior work used to be visibly junior, which made it cheap to correct. AI-assisted junior work is confident, plausible, well-formatted and occasionally wrong in ways that take a senior person to spot. That's a harder review, not an easier one, and the data reflects it: 45% of developers say debugging AI-generated code is more time-consuming, 46% actively distrust AI accuracy (up from 31% a year earlier), and 77% of employees using AI in Upwork's survey report it has added to their workload. Even experts misread their own experience here. METR found developers were 19% slower with AI while believing they were 20% faster. The fix is to change what you review: ask for the reasoning and the rejected options, not the artifact.
I'm a junior CS major and AI is doing the work. Do I have any hope of getting hired?
Yes, but the thing you're selling has moved. The market has stopped paying for the ability to produce the artifact, because that's what got automated. SignalFire puts new-grad hiring down around 65% at major tech firms and 76% at early-stage startups versus 2019, and the New York Fed has recent-grad underemployment at 42%. What's now scarce is verification: being the person who can tell whether AI output is right, why it's wrong, and what it's about to break downstream. That's built by doing hard things without the model and being corrected, which nobody will schedule for you. Optimise your first job for correction density, proximity to someone more expert who will tell you you're wrong, over title and over salary within reason. And go where the system is visible: small teams and startups still let you see the whole machine, which is the complexity ingredient that big-company junior roles have mostly stopped supplying.
If AI automates the bottom rung, how do people earn judgment?
By design, now, rather than by default. The default path is gone. Judgment is compressed correction: a stack of specific occasions when you were wrong and found out. Matt Beane's research on robotic surgery is the clearest precedent. The technology let attendings operate alone, so residents got far less practice, and the ones who became competent did it through unsanctioned "shadow learning." Beane's three ingredients, challenge, complexity and connection, are what any replacement has to reproduce: work at the edge of current capability, visibility into the system around the task, and a relationship with someone who will correct you. In practice that means deliberately staffing juniors onto ambiguous Augment-zone work where a human still decides, reviewing their reasoning rather than their output, and treating senior coaching time as delivery rather than overhead. The companies that do this will be buying senior judgment at a discount in 2033. The ones that don't will be renting it at whatever the market asks.
Sources
- Erik Brynjolfsson, Bharat Chandar & Ruyu Chen: Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab, revised 12 August 2026; high-frequency ADP administrative payroll data covering millions of US workers through June 2026. Source of the 19% relative employment gap for workers aged 22–25 in AI-exposed occupations, the hiring-not-separations mechanism, and the absence of the effect for experienced workers. The authors explicitly frame these as descriptive indicators rather than causal estimates. https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/
- SignalFire: State of Tech Talent Report 2026 (22 June 2026). Built on SignalFire's Beacon platform tracking 650M+ individuals and 80M+ organisations. Source of new-grad hiring down ~65% at the tech majors and ~76% at early-stage startups versus 2019. https://www.signalfire.com/blog/signalfire-state-of-talent-report-2026
- Federal Reserve Bank of New York: The Labor Market for Recent College Graduates (Q2 2026 update). Recent-graduate unemployment 5.6%; underemployment 42%. https://www.newyorkfed.org/research/college-labor-market
- GitClear: The Maintainability Gap: 2026 AI Code Quality Research (January 2026). 623 million analysed code changes, 2023–2026. Block duplication 40.3 → 73.0 per thousand changes (+81%); moved lines 13% (2023) → 3.8% (YTD 2026); cross-file function calls −35%. https://www.gitclear.com/the_ai_code_quality_maintainability_gap
- Stack Overflow: 2025 Developer Survey (published 29 July 2025). 49,000+ developers across 177 countries. 84% using or planning to use AI tools (up from 76%); 46% distrust AI output accuracy (up from 31%); 66% frustrated by "almost right, but not quite"; 45% cite time-consuming debugging of AI-generated code. https://survey.stackoverflow.co/2025/ · https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/
- Upwork Research Institute: From Burnout to Balance / AI-Enhanced Work Models (2024). Survey of 2,500 C-suite executives, full-time employees and freelancers across the US, UK, Australia and Canada. 77% of employees using AI say it has added to their workload; 39% report more time checking AI-generated content. https://www.upwork.com/research/ai-enhanced-work-models
- METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (10 July 2025). RCT with 16 experienced developers on 246 real issues in their own large repositories: 19% slower with AI tools, against a self-estimate of 20% faster. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Erik Brynjolfsson, Danielle Li & Lindsey R. Raymond: Generative AI at Work. NBER Working Paper 31161; published in the Quarterly Journal of Economics (2025). 5,179 customer-support agents, ~3 million chats. +14% productivity on average, +34% for novice and low-skilled workers, minimal effect on the most experienced. https://www.nber.org/papers/w31161
- Fabrizio Dell'Acqua: Falling Asleep at the Wheel: Human/AI Collaboration in a Field Experiment on HR Recruiters. Pre-registered field experiment, 181 professional recruiters reviewing 44 résumés each with randomly varied AI quality. Recruiters with higher-quality AI were less accurate and exerted less effort; the low-quality-AI group learned and improved, an effect driven by more experienced recruiters.
- Matt Beane: Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail. Administrative Science Quarterly, 2018. Two years of fieldwork; hundreds of procedures observed at five hospitals plus interviews with surgeons and trainees at thirteen more. Robotic systems let attendings operate without a resident, sharply reducing hands-on practice; competent trainees relied on norm-bending "shadow learning." https://theconversation.com/young-doctors-struggle-to-learn-robotic-surgery-so-they-are-practicing-in-the-shadows-89646
- Matt Beane: The Skill Code: How to Save Human Ability in an Age of Intelligent Machines (2024). Source of the challenge / complexity / connection triad. https://ide.mit.edu/publication/the-skill-code-how-to-save-human-ability-in-an-age-of-intelligent-machines/
Frameworks referenced are my own: the Hourglass Collapse, the ORBIT Framework, and the 3-Zone Co-Pilot (Replace / Augment / Refuse).
The rungs worth keeping
The junior-work zone sort... which entry-level work to automate, which to protect as training, and how to tell the difference. It goes out to the list.
If you're about to make this trade and want the diagnosis first: Book a CAIO scoping call →