Automation Bias
People can over-rely on automated recommendations, sometimes following an incorrect recommendation or failing to notice a problem the automation missed, even when contradictory information is available elsewhere.
What Is It?
Automation bias describes the tendency to give automated or algorithmic outputs more weight than they’ve actually earned, including a documented tendency to both over-rely on automation when it’s wrong (an error of commission) and under-monitor a task closely enough to catch an automated error at all (an error of omission). It was studied and named within human factors and aviation psychology research in the 1990s, notably by Kathleen Mosier and Linda Skitka, whose research on pilots and automated cockpit systems found that people sometimes failed to notice or act on contradictory information available to them, or actively followed an incorrect automated recommendation. In some of these studies, unaided participants working without the automated recommendation actually caught errors that aided participants missed, even though automation improved performance on average across the full set of cases; the notable pattern isn’t that automation makes people worse overall, it’s the distinctive errors it can introduce in the specific instances where it fails. The core finding isn’t that automation is unreliable, most automated systems in these studies were highly accurate; it’s that people’s monitoring and independent judgment can weaken specifically because the presence of an automated recommendation invites a kind of trust that isn’t always warranted case by case.
The underlying phenomenon predates modern AI and is well established in decades of human-factors and decision-support research, particularly in aviation and clinical settings. Evidence of analogous overreliance in human-AI collaboration is growing quickly, but research on modern AI-assisted work is a newer and less mature literature than the older automation research, and factors like system accuracy, user expertise, task difficulty, and how easy a recommendation is to verify all appear to strongly affect when and how much overreliance shows up.
Why Does It Matter?
The organizational relevance has sharpened as more decisions run through dashboards, scoring models, and AI-assisted tools. A number produced by a system carries a kind of unearned authority simply by virtue of looking precise and system-generated, and a team can defer to a flawed model’s output over a domain expert’s contradicting judgment specifically because the model’s output feels more objective, not because it’s actually more accurate in that instance. This is a genuinely different failure mode from simply having a bad model: the model can be good on average and still be wrong in a specific case, and automation bias is what keeps a person from catching that specific case when their independent judgment would otherwise have caught it.
What Changes Once You See It?
You start treating an automated recommendation as one input to weigh, not a conclusion to defer to. Disagreement between an automated recommendation and an informed independent judgment should trigger verification, not automatic deference to either side, since experts can be wrong and algorithms can legitimately outperform expert intuition; the goal is calibrated reliance, not a default toward either humans or machines. You get more deliberate about building in a real, exercised habit of checking automated outputs against independent judgment on the decisions where being wrong is expensive, rather than treating that check as optional once a system has proven generally reliable. You also become more attentive to whether people are still actively monitoring an automated process for errors, or have quietly stopped watching closely because the system has been reliable so far.
Common Misunderstandings
- It isn’t a claim that automated systems are generally untrustworthy or that people should default to distrusting them. Reflexive distrust of automation just trades one bias for its mirror image; most automated systems studied in this research were quite accurate overall, the problem is specifically the erosion of independent monitoring and judgment in the individual cases where the system is wrong.
- It isn’t only about actively following a bad recommendation. The same research identifies a distinct failure mode, failing to notice or act on an error at all because attention has shifted away from independent monitoring, which can be just as consequential as actively trusting a wrong answer.
- It doesn’t mean more human oversight automatically fixes the problem. Oversight that consists of passively watching a normally reliable system doesn’t reliably catch the specific cases where it’s wrong; the corrective has to involve genuinely independent judgment being exercised, not just a human nominally present in the loop.
- It isn’t unique to advanced AI systems specifically. The effect was identified and named in the context of comparatively simple automated cockpit systems decades before modern AI tools existed, and it applies to any output that looks system-generated, a calculator, a decision-support system, a scoring algorithm, a dashboard metric, an AI system, not only to sophisticated algorithmic models. Worth keeping the category technological rather than stretching it to cover something like a scoring rubric a person fills out by hand, which risks the term swallowing other, distinct measurement effects.
Diagnostic Question
What independent evidence would make us override this automated recommendation, and are we actually checking for it? A useful companion question in the other direction: are we giving an automated recommendation less scrutiny than we’d give the same conclusion from a person, or more?
Explore Further
Field Notes
- None yet.
Related Field Guide
Origin
Studied and named within human factors and aviation psychology research in the 1990s, notably by Kathleen L. Mosier and Linda J. Skitka; see Mosier, Skitka, Heers, and Burdick, “Automation Bias: Decision Making and Performance in High-Tech Cockpits,” The International Journal of Aviation Psychology (1998).