Every argument for AI on a program starts the same way: it takes the low-value work off your people so they can do the high-value work. I've made that argument myself, and I stand behind the economics. But there's a second-order effect nobody puts in the business case, and it doesn't show up for two or three years — which is exactly why it keeps getting missed.
The low-value work is where the high-value judgment came from.
I did not develop a feel for schedule risk by attending a training. I developed it by building schedules that were wrong, watching how they were wrong, and slowly assembling an internal library of the ways a plan lies to you. The tedium was the curriculum. Reconciling a broken dependency map at eleven at night is how you learn what a broken dependency map feels like from across the room — and that recognition, years later, is what lets you look at a green status report and say the thing nobody else in the room can articulate: something here is off. Hand that work to an agent and the output improves immediately. The apprenticeship quietly ends.
Aviation ran this experiment first
We don't have to speculate about where automation dependency leads, because commercial aviation has been living it for thirty years, with better instrumentation and higher stakes than any PMO. The pattern is documented and it is not comforting.
The DOT Inspector General reported that senior FAA officials estimate airline pilots use automated systems roughly 90 percent of the time in flight, and that as automated procedures expanded, opportunities to practice manual flying diminished. The same report cites a Flight Safety Foundation study of thirty experienced U.S. commercial pilots whose manual flying scores came in below FAA standards — while 80 percent of those pilots reported that they typically hand-fly the aircraft. They were not lying. They believed they still had it. That's the part that should make every program leader sit up: the degradation is invisible from the inside, and confidence is the last thing to go.
The mechanism transfers cleanly. Automation is reliable, so it handles the routine, so humans get their reps only on the exceptions — and you cannot build exception-handling judgment out of exceptions alone, because the exception is precisely the moment you need judgment you already have. Aviation's answer wasn't to remove the automation. It was to write manual practice into training policy, because reps that don't happen naturally have to be scheduled.
Automation doesn't just move work off your team — it moves the training ground off your team. If the reps that built judgment are gone, judgment has to be trained deliberately, or you're running on a skill base with no replacement plan.
The people most affected haven't started yet
If you have twenty years in, your instincts are already paid for. AI is pure leverage for you — you can smell a wrong number in an AI-drafted weekly because you spent a decade producing wrong numbers by hand. The exposure is entirely on the people two years into the job, who will spend their formative years reviewing outputs instead of producing them, and who will therefore learn what a good report looks like without ever learning what makes it true.
Reviewing is not a substitute for doing. It feels adjacent, which is why the risk is easy to wave off, but the cognitive work is different in kind: producing forces you to confront every gap in the source material, while reviewing lets you assess surface plausibility and move on. And there's a nasty circularity here. Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers across 936 real AI-assisted tasks and found that higher confidence in the AI predicted less critical thinking, while higher confidence in one's own expertise predicted more — and that workers tend to skip critical evaluation precisely when they lack the skill to inspect and improve the output. The people least equipped to catch the machine's errors are the people least likely to look for them. Deskilling doesn't just create a weaker reviewer; it creates a reviewer who doesn't know they're weak, which is how you end up with a rubber-stamp review that everyone still counts as a control.
Training the bench on purpose
The Army solved a version of this a long time ago, and the solution is boring: you train the way you'll have to fight, including the parts the equipment normally does for you. We navigated with map and compass while carrying working GPS — not out of nostalgia, but because the day the batteries die is not the day to start learning terrain association. Same logic, different domain.
On my programs that translates into a handful of unglamorous practices. Junior staff produce some artifacts cold — no agent — not because it's efficient, but because it's how they earn the pattern recognition that makes their later reviews worth anything. Before anyone reads the AI-drafted forecast, they write down their own call; comparing the two builds calibration, and a reviewer with a prior is a genuine check while a reviewer without one is a spellchecker. When an agent's output gets corrected, we capture why — the correction log becomes the fastest teaching material on the program, because it's a curated stream of exactly the ways this machine goes wrong on this work. And we run periodic manual drills on the critical path: rebuild the schedule by hand, reconstruct the risk picture from source, once a quarter. It costs real hours. It is cheaper than discovering during a crisis that nobody on the team can do it.
- Name the skills you're not willing to lose — schedule construction, estimation, risk reasoning — and protect reps in each one, on the calendar, not on good intentions.
- Have reviewers form their own answer before reading the AI's. No prior, no real review.
- Give junior staff production work, not just review work. Reviewing teaches you what good looks like; producing teaches you what makes it true.
- Log every correction to an agent's output with the reason. That log is your apprenticeship curriculum, already written.
- Run manual drills on the critical path quarterly, while the automation still works — the same reason you test a generator before the outage.
The bottom line
This isn't an argument against AI on programs; I've spent the last year building exactly these capabilities and I'd build them again. It's an argument that the efficiency case and the capability case run on different clocks. The savings land this quarter. The bill arrives in three years, when the person who should be running your recovery has spent their whole career approving drafts. Aviation didn't fix automation dependency by flying less — it decided proficiency was something you schedule rather than something you assume. Program leadership needs the same decision now, while the people who learned the hard way are still around to teach it. Delegate the work. Keep the reps.