The Human Was in the Loop. The Airplane Still Fell.
- Allen Westley

- Jul 24
- 8 min read
Updated: 36 minutes ago

At 2:10 in the morning over the equatorial Atlantic, an Airbus A330 handed control back to the people who were supposed to have it all along.
Ice had formed over three small sensors on the aircraft's skin. For less than a minute, the flight computers lost a reliable airspeed reading, and the autopilot did exactly what it was designed to do when it can no longer trust its inputs: it disconnected and gave the airplane back to its crew. Three qualified pilots. A functioning aircraft. A recoverable situation.
Four minutes and twenty-three seconds later, Air France 447 struck the ocean. All 228 people aboard were lost.
The airspeed problem corrected itself in under a minute. The airplane was flyable the entire way down. At one point the stall warning sounded continuously for fifty-four seconds, and the crew flew on as though it hadn't. They never understood they were stalling. The machine had been so capable, for so long, that when it finally asked the humans to take over, the humans no longer had the practiced judgment the moment required. They were in the loop. Their hands were on the controls. The authority was entirely theirs.
The judgment wasn't there.
The official finding, years later, refused to pin it on three individuals. The problem was systemic - an entire generation of pilots trained to manage automation rather than to fly, whose competence had quietly eroded in exact proportion to how reliable the automation became. Presence had been mistaken for control. On paper, oversight existed. In the cockpit, at 2:10 in the morning, it did not.
I keep this case close because it is the clearest picture I know of a failure now migrating, quietly, into the Defense Industrial Base - and it does not require an airplane.
One Degree
Our compass points to a single idea: a signal to follow. So let me use the instrument itself to make the point.
A navigator holding a course one degree off true notices nothing. The error is invisible on any timescale that matters in the moment. But one degree compounds - roughly ninety-two feet of drift for every mile traveled. Over an ocean crossing, one unexamined degree is the difference between the runway and the water.
Automation bias is one degree off true. No single approval feels like surrender. Each one is reasonable - the tool has been right before, the queue is long, the recommendation is clean. The drift is never in the individual decision. It is in the accumulation, in the slow migration of judgment from the human to the system, one defensible concession at a time. By the time the consequence arrives, the course was set thousands of decisions ago, and no one can point to the moment it went wrong, because no single moment did.
AI did not remove the human from the loop. It moved the loop. And the most dangerous drift is the kind that shows green the whole way down.
6:42 on a Monday
Move the cockpit into a DIB security operations center.
It is 6:42 on a Monday. Overnight, an AI-enabled workflow correlated an endpoint alert with identity activity, asset criticality, vulnerability data, threat intelligence, and a machine-readable runbook. Before the first analyst sets down her coffee, the system has assembled the evidence, assigned severity, drafted the ticket, recommended containment, and named the person authorized to approve it.
The package is clean. Every field is populated. The reasoning reads as coherent.
Underneath it, one relationship is wrong. A stale ownership record points to the wrong program environment. A threat-intelligence tag carried forward from an earlier report inflates the confidence of the attribution. The model did not hallucinate - it produced a plausible answer from structured context, and structure reads as authority. The analyst has seven more alerts waiting. The recommendation matches what the dashboard says. The button is ready.
She is in the loop. Her name will be on the record.
Is the judgment in the room?
This is the human-in-the-loop fallacy - the belief that human presence guarantees human control. But if the system selected the evidence, ranked the risks, framed the alternatives, drafted the rationale, and compressed the clock, then the decision was substantially formed upstream of the person approving it. Approval is not authorship. Approval is not validation. Approval is not ownership.
In my Cognitive Security work I call the failure mode Cognitive Surrender: an organization keeps the appearance of human approval while removing the conditions that make approval mean anything. It needs no rogue model and no careless employee. It grows from incentives that each look entirely rational - faster response, leaner teams, shorter cycle times, confidence in a tool that has usually been right, leadership expectations built around machine speed.
Every concession makes sense. Together they produce reciprocating consequences: the more the organization leans on machine judgment, the less its people practice their own; the less practiced that judgment becomes, the more indispensable the machine appears. That is not science fiction. It is AF447's training gap, rewritten for the SOC. It is a readiness problem.
Why the DIB Is Not an Ordinary Enterprise
In a commercial setting, this drift produces audit friction or financial loss. In the Defense Industrial Base, the same drift crosses different terrain - Controlled Unclassified Information, export-controlled technical data, classified program boundaries, supplier trust, mission systems, authorization artifacts, contractual obligation. A response can be technically efficient and operationally wrong: right procedure, wrong boundary; clean output, weak provenance; complete package, invisible author.
Picture an AI-enabled compliance agent assembling an authorization package. It searches policy, prior findings, system inventories, POA&Ms, control statements, vulnerability results, inherited controls - and drafts a persuasive narrative explaining why each control is satisfied. The reviewer sees a finished document. What the reviewer cannot see is which sentences came from authoritative evidence, which were inferred from adjacent artifacts, which were carried forward from a stale assessment, and which were generated by another AI system entirely.
The risk is not one wrong paragraph. It is interpretive displacement - the machine's explanation becomes the organization's understanding of its own system, while the humans drift steadily further from the evidence beneath it. Fix the stale source months later and the interpretation can still be living in agent memory, cached summaries, copied reports, workflow templates, and human assumptions.
That is the Persistent Knowledge Dilemma: removing the error does not remove its operational effect. In the DIB, that residue doesn't stay in a database. It travels into mission assurance, supply-chain trust, and the credibility of your security posture.
The New Attack Surface Is the Decision Path
We have spent a generation hardening systems, identities, networks, applications, and data. Those controls still matter. But agentic workflows expose a surface we have not been defending: the path by which facts become conclusions and conclusions become action.
An agent draws on endpoint telemetry, identity graphs, threat reporting, asset catalogs, engineering repositories, ticket histories, security plans, and runbooks - structured context I describe as Knowledge Graph Infrastructure. The benefit is extraordinary speed. The hazard is structured confidence. A stale document can be corrected. A poisoned relationship is harder to catch, because it quietly shapes many summaries, rankings, and tickets before anyone questions it - and, per the PKD, persists after the source is fixed.
Which is why the model is not the only thing that needs scrutiny. A defender has to be able to reconstruct what the agent could read, what it was permitted to infer, what it could update, which tools it could invoke, which other agents contributed, which human owned the decision, and what evidence let that human actually challenge it.
That reconstruction is the purpose of Agentic Role Mapping - a working map of identities, permissions, tools, memory, knowledge sources, decision rights, escalation paths, and evidence requirements. Not a list of approved models. A wiring diagram for authority.
A Harder Test for Human Oversight
NIST's AI RMF asks organizations to define roles, responsibilities, and oversight mechanisms across the manual-to-autonomous spectrum. OWASP's agentic work names memory, tools, identity, excessive autonomy, and human oversight as linked concerns. Good foundations. But the operational test is harder than "is a human assigned?" It is whether that human keeps enough understanding and authority to intervene when the system is fast, persuasive, and usually right.
I call that line the Cognitive Integrity Threshold. Before you count a human approval as a control, put it through five questions:
1. Can the reviewer explain the basis of the recommendation without repeating the AI's summary back to you?
2. Can the reviewer separate observed fact from machine-generated inference?
3. Does the reviewer have the time, expertise, evidence, and authority to actually challenge it?
4. Can the reviewer stop, redirect, or reverse the action without an unacceptable operational penalty?
5. Can the organization reconstruct which knowledge, permissions, rules, agents, and human actions shaped the outcome?
Several "no" answers mean you have a human in the workflow, not meaningful oversight. The approval is ceremonial - a signature standing in for a judgment that was never actually exercised. It is the pilot monitoring the automation instead of flying the airplane, right up until the automation hands it back.
Designing Authority at Machine Speed
The answer is not to slow every workflow to human speed. That surrenders the defensive advantage AI provides and adds risk through delay. The goal is decision authority at machine speed - and that takes a constellation of accountability rather than one tired person as the last stop before execution.
System owners, cyber operators, engineers, data stewards, human-factors specialists, mission leaders, acquisition professionals, and governance teams each hold part of the picture. Push informed challenge close to the point of action; keep accountability clear enough that no one can vanish behind the algorithm.
Five moves to start:
Map authority before autonomy. Document what the system may recommend, decide, initiate, and execute - and name the human owner for each consequential decision.
Separate evidence from inference. Decision packages should visibly distinguish source fact, machine-derived relationship, confidence level, assumption, and agent-written narrative. If the reviewer can't tell them apart, neither can the auditor.
Test the override, not just the output. Run tabletops that measure whether reviewers can detect stale context, challenge a persuasive recommendation, halt an action, and recover from a wrong one. AF447 was not a knowledge failure. It was an override failure at 2 a.m. Train for the 2 a.m. version.
Preserve provenance and portability. Logs, source references, graph changes, prompts, agent actions, and approvals must be reconstructable. If the decision record can't leave the vendor's platform, you don't really own it.
Watch for Cognitive Indicators of Compromise. Repeated rubber-stamping, declining manual proficiency, disappearing dissent, unexplained confidence, stale facts resurfacing, conclusions that can't be traced to evidence - these are the drift showing on the instruments before the technical incident ever appears.
The Question Leaders Cannot Delegate
The DIB is under real pressure to move faster. Threats don't keep committee hours, and no team clears machine-scale signal by hand. Used well, AI lets analysts put themselves in the path of opportunity - surfacing relationships, compressing research, sharpening response. It can genuinely strengthen human judgment.
But speed without decision integrity is not readiness. It is latency disguised as certainty - the gap between the moment authority quietly moved and the moment the organization finally feels the consequence. On AF447, that gap was four minutes and twenty-three seconds. In a DIB governance failure, it may be a quarter you don't discover until the audit, the incident, or the mission tells you where your judgment actually was.
The next major AI governance failure in the DIB probably won't announce itself with a rogue agent or a machine defying a human command. It will look like a professional recommendation, backed by a complete evidence package, approved by a qualified person who no longer had a practical way to know whether it was right.
The human will be in the loop. The signature will be on the record. The accountability will still be human.
The only open question is whether the judgment will be there too - or whether we'll find out, fifty-four seconds into a warning nobody heard, that the airplane has been falling for a while.
Cyber Explorer Field Question
Take one AI-assisted workflow in your organization and trace a single decision from source evidence to final action. At each step, ask: who shaped the judgment, who held authority, and what proof would let us reconstruct the path? If you cannot map it, you cannot govern it.
Follow Allen Westley and Cyber Explorer for continuing research on Cognitive Security, AI decision authority, Agentic Role Mapping, and the protection of human judgment in high-consequence environments.
Boundary note: This article reflects Allen Westley's independent Cyber Explorer research lens and is intended for educational purposes. It does not represent the position of any employer, customer, government agency, standards body, or partner, and should not be treated as legal, contractual, classification, export-control, or compliance guidance.




Comments