Writing · Applied AI
Human Baseline Blindness: Why We Forgive Human Errors We Condemn in AI
How to compare AI systems with measured human work, then set a higher standard where scale or consequence demands it.
- Human baseline blindness
The evaluation error of comparing an AI system with an idealized employee instead of measuring how people perform the same bounded work under comparable information, authority and operating constraints.
An AI system makes a bad recommendation and the output is preserved perfectly. We can replay it, screenshot it and examine every unsupported sentence. A colleague makes a similar error and the organization instinctively adds context: the policy was unclear, the queue was long, the case was unusual, or the relevant information lived in another system.
That asymmetry is understandable. A person has intent, history and relationships. A system output arrives as an artifact. But the asymmetry can distort evaluation. The human side of the comparison becomes a story about competent people on a normal day; the AI side becomes a dataset of isolated failures.
Both sides need an honest standard. Measure the human workflow that exists, define the workflow the organization wants and state where an AI system must meet a deliberately higher standard.
The fictional perfect employee
Many evaluation conversations begin with sensible questions—Did the system follow policy? Did it identify the risk? Did it escalate correctly?—and then quietly assume that a trained person always would.
The assumed employee has complete recall, consistent attention, unlimited time, uniform access to information and no disagreement with peers. This person is not a baseline. It is a specification.
Specifications matter. Some outcomes are unacceptable even if people make the same mistake. But using a specification while calling it “human performance” creates two problems. First, it hides the real opportunity: the system may improve consistency in a workflow whose human variation was never measured. Second, it hides the real risk: an automated mistake may propagate faster and further than a human one.
Separate three things from the start:
- Observed human performance: what people do in representative work.
- Minimum system threshold: the lowest acceptable performance for deployment.
- Normative target: the outcome the organization ultimately wants, regardless of who or what performs the task.
They may produce different thresholds and may be judged by different rules. The important thing is that the difference is deliberate.
Why the asymmetry feels reasonable
Human errors arrive with explanations. We know the person was interrupted, the input was ambiguous, or two policies were in tension. We also know they may notice novelty, ask a colleague, or recover from an error through judgment that was never written into the process.
AI errors have different characteristics. They can be highly repeatable. A polished response can create unwarranted confidence. The same failure can be reproduced across every case at machine speed. When a system is integrated with tools, its output can become an action before someone has time to notice.
Those differences justify scrutiny. They do not justify an undefined baseline.
A fair comparison can use different standards. What matters is an explicit baseline and a documented reason for asking the software to exceed it.
Measure a bounded task
“Support agent,” “engineer” or “analyst” is too broad a unit for comparison. Start with a bounded task and its decision context.
A useful baseline protocol records:
- the task and intended outcome;
- the information available at the decision point;
- the tools and permissions available;
- the experience band of the person doing the work;
- case complexity and time pressure;
- the decision, explanation and escalation path;
- later evidence that makes review possible.
Sample routine work as well as difficult work. If the dataset contains only exemplary expert cases, it measures a ceiling. If it contains only costly failures, it measures a tail. Neither describes normal operating performance.
Disagreement should be preserved rather than forced into false ground truth. Two competent people may interpret incomplete evidence differently. That can reveal an ambiguous policy, missing instrumentation or a task that needs a distribution of acceptable answers rather than one canonical response.
Where practical, review human and system work against the same rubric and without revealing its source. Blind review does not remove every bias, but it prevents the evaluator from quietly grading polish, confidence or presumed intent differently because the author is known. Preserve reviewer disagreement as data; it often identifies a weak task definition more accurately than a forced consensus.
Build an evaluation contract
Before comparing systems, write down the rules that will govern the comparison. This avoids changing the standard after seeing a result.
| Part | Question | Decision to record |
|---|---|---|
| Task boundary | What exact work is being evaluated? | Inputs, outputs, tools, permissions and stopping conditions. |
| Human baseline | Who performs comparable work today? | Sampling method, experience bands and information available. |
| System threshold | What is the minimum acceptable behavior? | Severity-specific pass criteria alongside average accuracy. |
| Higher standard | Where must the system exceed people? | The specific scale, consequence or oversight reason. |
| Review trigger | When does the contract expire? | Model, workflow, policy, data or operating changes that require re-evaluation. |
This contract also makes leadership decisions legible. A system can outperform average human work and still be unsuitable for deployment because its rare failures are severe or hard to detect. Another system can be less capable overall but useful within a narrow, supervised boundary.
I treat observed human performance as diagnostic evidence. It is not a deployment threshold and certainly not permission to automate an existing failure.
Count consequence as well as errors
An undifferentiated error rate treats a harmless formatting issue and unsafe operational advice as equivalent events. A useful evaluation decomposes error along several dimensions.
| Dimension | Question | Why it changes the decision |
|---|---|---|
| Severity | What harm can follow? | High-consequence errors can dominate an otherwise strong average. |
| Detectability | Will a user or control notice? | A visible error and a plausible hidden error require different safeguards. |
| Recoverability | Can the effect be reversed? | Reversible work permits a different operating boundary from permanent change. |
| Propagation | Can the error influence later work? | Automated reuse can turn one mistake into a system pattern. |
| Authority | What may the system do without approval? | The same recommendation has different risk when it is also an executable action. |
Harmful advice
Incorrect, irrelevant and unsafe are not synonyms. A response can be factually wrong but easy to detect. It can be factually plausible yet unsafe because it omits a precondition. It can be technically correct but outside the user’s authorization.
Grade confidence, actionability and consequence together. The highest-risk output is often not an obvious fabrication; it is a specific, credible instruction that should not have been given in that context.
Unnecessary escalation
Escalation is usually scored as caution, but excessive escalation transfers cost and delay to the user and receiving team. A system that avoids every difficult decision may look safe in a simple metric while providing little value.
Review whether the escalation occurred at the right boundary, whether the reason was clear, and whether the handoff reduced or increased the next person’s work.
Missed escalation
Missed escalation deserves separate treatment, especially when consequences are high. Define the triggers before running the evaluation: ambiguous authorization, safety-sensitive conditions, contradictory evidence, novel cases or low confidence where action would be hard to reverse.
An average accuracy score can hide a system that is excellent on routine work and unsafe at the boundary. Tail cases are not statistical noise when the operating model routes important decisions through them.
Consistency cuts both ways
Humans are variable. That can create mistakes, but it can also create diversity of approach. One person may notice that a case does not fit the usual pattern. Another may remember a similar incident that documentation failed to capture.
AI systems can be more consistent. That helps when the desired process is well defined. It is dangerous when the system repeats the same blind spot with confidence. Compare the full distribution as well as the mean:
- How often does performance fall below a critical threshold?
- Are failures clustered around a class of inputs?
- Does the same unsupported assumption recur?
- How does behavior change under missing or contradictory information?
- Can monitoring detect the pattern before it propagates?
Consistency is not automatically quality. It is an amplifier of whatever behavior the system has learned.
Compare the whole operating system
Generation time is not task time, and a model is not a workflow. A credible comparison includes context preparation, tool use, review, correction, escalation, monitoring, incident handling and the work required to maintain the evaluation itself.
The same rule applies to human work. Salary divided by ticket count is not a human baseline. Training, tooling, coordination and quality review are part of the system. The comparison should be made at the workflow level, using a defined period and workload shape, without inventing a universal cost number.
Speed should also be measured at the point where the result becomes usable. A fast draft that requires slow expert reconstruction may help, but it is not equivalent to a completed task. A slower system that prepares an inspectable decision and a clean handoff may create more operating leverage than one optimized for first response.
A baseline is not a ceiling or a floor
Observed human performance answers a descriptive question: how does this work perform today? It does not answer the normative question: what should the organization accept?
If people routinely work around an ambiguous policy, the system should not be trained to reproduce the workaround unquestioningly. If a safety boundary is missed by people under pressure, matching that rate is not success. If expert judgment recovers from poor instrumentation, automating the task without improving the instrumentation may remove the very mechanism that keeps the workflow safe.
Use the baseline to locate variation, friction and unwritten judgment. Then decide which parts should be preserved, improved, constrained or redesigned before automation.
When AI should face a higher standard
There are strong reasons to require an AI system to exceed observed human performance:
- it acts at greater speed or scale;
- users may treat its language as institutional authority;
- failures are difficult to detect before action;
- one error can propagate automatically;
- the organization is reducing human oversight;
- the affected people have limited ability to appeal;
- the task has safety, legal or other high consequences.
The higher standard should be tied to the reason. “AI must be perfect” is not an operational requirement. “No autonomous state change when authorization is ambiguous” is. “Escalate every difficult case” is not a service boundary. “Escalate when evidence conflicts and the next action is difficult to reverse” can be designed and tested.
NIST’s AI RMF emphasizes that risk depends on context and that human roles, feedback and oversight should be defined. The human–AI interaction guidelines developed by Microsoft Research similarly treat correction, explanation and the ability to act when the system is wrong as design concerns rather than afterthoughts.
The rule I use is simple: measure human work to understand the opportunity; set the deployment standard from consequence and scale.
Replace the argument with a contract
Teams often debate AI quality in moral language: either the system is held to an impossible standard, or human imperfection is used to excuse machine risk. A measured baseline resolves neither question automatically. It replaces rhetoric with an operating decision.
It reveals where the existing workflow already fails, where a system improves consistency, where automation introduces new failure modes and where humans remain essential. Most importantly, it forces the organization to say what “good enough” means before deployment.
I would not approve deployment from a comparison with a fictional perfect employee. I want measured human work, including its uncertainty and recovery, followed by a deployment standard tied to consequence. The difference between those two numbers should be deliberate, documented and reviewed when the workflow changes.
Sources and further reading
- NIST, Artificial Intelligence Risk Management Framework, frames AI risk management around the specific context in which a system is designed, deployed and evaluated.
- NIST, AI RMF Core, includes outcomes for defined human oversight roles, contextual validation and feedback processes.
- Amershi et al., “Guidelines for Human-AI Interaction”, reports a synthesis and validation of 18 interaction guidelines.
- OpenAI, “Measuring the performance of our models on real-world tasks”, describes work-like tasks, expert-authored rubrics and blinded expert comparison, while noting that one-shot evaluation does not capture interactive workflows.