Executive Summary:
Automation is an efficiency tool; accountability is a human function no model can own. The danger isn’t the obvious error; it’s the confident one that goes unflagged. A model that’s only 92% accurate is a 100% liability: in high-stakes, high-variance domains, the missed 8% clusters at your edge cases and surfaces only at audit, with months of compounding damage already booked. Benchmark accuracy is not production reliability. The fix isn’t less automation; it’s full automation where errors are reversible, and a human gate where they aren’t. The question that matters: who owns the consequence, and how fast will they know?
Why an automation model that’s 92% accurate is a 100% liability, and you won’t know which 8% until the audit. In a high-stakes domain, that 8% isn’t randomly distributed; it clusters exactly where your edge cases live. You won’t know which transactions, which claims, which decisions landed in that 8% until an auditor tells you.
Picture a pipeline that has been confidently wrong for an entire quarter. It threw no exceptions. It posted no alerts. Every dashboard showed green. Then a regulator, an auditor, or a customer finds the error you didn’t, and hands it back to you with three months of compounding damage already baked into your ledgers, your CRM, and your regulatory filings.
That is the failure mode worth losing sleep over. Not the error. The lag before anyone knows it happened.
The problem was never automation. The problem is unsupervised autonomy pointed at high-stakes, high-variance, unstructured data, where benchmark accuracy gets mistaken for production reliability, and no human is positioned to own the consequence. The macro picture is well documented: McKinsey’s 2025 State of AI survey found that 51%[1] of AI-using organizations have already hit at least one negative consequence, nearly a third of it traced to inaccuracy. But the most expensive subset of failure is narrower and quieter than a botched rollout. It is the moment a team aims an autonomous system at messy data and assumes the model has it handled.
As a strategic technology and data partner, Damco’s position is clear: automation is an efficiency tool, accountability is a human function, and it cannot be delegated to a mathematical model.
When Model Failure Stays Silent
Deterministic software fails loudly. A query breaks, an integration drops, a value comes back null, and the system throws a stack trace and stops. You know within seconds, and the blast radius is contained at the point of failure.
Probabilistic AI fails differently. In rule-free, high-variance environments, it doesn’t halt. It smooths the anomaly over with a confident answer and keeps going. The error doesn’t stop at the point of failure; it travels.
This is the accuracy ceiling, and it is where most leaders miscalibrate. An 85% to 95% accuracy rate looks like a win in a pilot. In production, across hundreds of thousands of daily decisions, it is a standing guarantee that 5% to 15% of outputs are wrong and unflagged. Peer-reviewed benchmarking against real enterprise data found LLMs dropping 14-19[2] percentage points in precision and recall versus their public benchmark scores and for domain-specific tail entities, accuracy collapses to 6-20%.
“These systems that we call autonomous… they’re not autonomous.”
– Cassie Kozyrkov, Google’s First Chief Decision Scientist, Founder & CEO, Kozyr.
Stanford’s RegLab tested 200,000+[3] queries against GPT-3.5, Llama 2, and PaLM 2 on specific, verifiable legal questions and found hallucination rates of 58-88%. Even GPT-4, the strongest model in the study, hallucinated 58% of the time. A follow-up Stanford study testing 2024 legal AI tools found hallucination rates still running 17-43% depending on the product.
And the cost of that base rate doesn’t surface at the moment of failure. IBM’s Institute for Business Value found in 2025 that over a quarter of organizations lose more than $5 million[4] annually to poor data quality and 7% report losses of $25 million or more. Those aren’t line items that appear on the day of the error. They accumulate in rework cycles, failed AI outputs, and audit findings that surface months after the root cause.
The corrupted output gets ingested, and the loss surfaces somewhere else entirely. Underwriting misprices risk, and you learn it at default, not at decision. Compliance files an inaccurate report, and you learn it at examination, not at filing.
The single bad line item was never the cost. The cost is the retroactive audit, the months of database history someone now reconciles by hand.
Turn Ambiguous Transaction Strings Into Evidence-Backed, Audit-Ready Tags
The Financial Data Cryptogram: Why “The Model Said So” Won’t Hold Up
Let’s start with the financial sector, because it exposes the limit most cleanly. Here, unstructured input has to map onto a rigid, auditable ledger and the input arrives as a cryptogram.
A card transaction rarely shows up as a clean merchant name. It comes through as a layout-dependent string of branch numbers, processor acronyms, and codes like GIGADAT, FLEXEPIN, or BR. 0072. Point a general-purpose LLM at these at scale and three failures surface fast:
- Output is non-deterministic. The same string classifies as a micro-loan on Monday and a miscellaneous expense on Wednesday.
- Meaning is layout-dependent. A shift in spacing flips a merchant name into a routing code.
- The model has no localized knowledge. Ask it about BR. 0072 and it confidently calls it a branch transaction when it is, in fact, a specific non-sufficient-funds fee code used by major Canadian banks. Strings like GIGADAT and ILIXIUM get filed as generic fintechs, when they are payment gateways for online gambling, a distinction that drives the entire risk score.
For credit underwriting, AML, and fraud work, non-deterministic guesses are not acceptable inputs. So, a financial data intelligence firm, tracking anonymized consumer transaction panels for hedge funds and institutions, ran research-backed human validation per tag with Damco. Ambiguous strings get resolved by analysts against commercial registries and a curated knowledge base, and every classification links to traceable evidence.
“AI hallucinates. It should be the tech that you don’t trust that much.”
– Sam Altman, CEO, OpenAI.
That layer does more than lift accuracy. It builds the audit trail because “the model said so” is not a legal defense.
In Moffatt v. Air Canada[5] (2024 BCCRT 149), the airline argued its chatbot was a separate entity responsible for its own bad advice. The tribunal rejected that outright: an automated tool is an extension of the company, and the company owns its errors in full. The courts are not softening on this; in Q1 2026 alone, US courts imposed a record $145,000[6] in sanctions on attorneys who filed AI-hallucinated citations.
The principle scales straight into regulated financial operations. Misclassify an AML tag and your institution owns the BSA violation; not the model, and not the vendor. If your tool makes a commitment or misprices a risk, the liability is yours.
The Scratch Test: Where Pure Automation Hits Its Ceiling
The capability limit isn’t unique to finance. Clinical video annotation exposes it just as cleanly and the dollars are larger.
In sleep medicine trials for conditions like atopic dermatitis, researchers track unconscious nocturnal scratching to measure sleep disturbance and drug efficacy. That means frame-level inspection of eight to ten hours of infrared video per patient, per night. It is a brutal manual workload, and the obvious instinct is to hand it to computer vision.
The instinct breaks on the details. Near-infrared footage is dark, low-contrast, and noisy. The patient is half-covered by a blanket, so pose tracking has little to lock onto. And here is the part no pixel-motion model solves on its own: to a bounding-box algorithm, a patient vigorously scratching an arm is mathematically indistinguishable from one repositioning, rolling over, or tugging a sheet. The model sees movement. It cannot see meaning. Left to call those events autonomously, it floods the dataset with false positives and false negatives.
A corrupted dataset doesn’t slow a clinical trial: it invalidates it. A failed pivotal trial runs $50M-$300M+ in sunk cost and lost market exclusivity.
This is why a digital health firm generating clinically validated endpoints for pharma trials, runs a human-in-the-loop model with Damco rather than a fully automated pipeline. Trained annotators supply the situational nuance, reading the rhythm of a scratch against the arc of a roll-over, accounting for a body under covers. Automation still earns its place as a pre-filter, surfacing high-motion windows so humans aren’t scrubbing dead footage. But the final call sits with a person.
In clinical research, automation without context isn’t merely inefficient. It is a scientific and regulatory liability.
What It Costs, and When You Find Out
Treat the financial toll as consequence, not as proof, and the picture sharpens.
The primary cost driver is time-to-detection. Architectural failures in AI systems can sit undetected for weeks before anyone notices, and the longer the lag, the more the rework compounds because every downstream system that touched the bad data now needs cleaning too.
The regulatory exposure is now sized and dated. Under the EU AI Act, penalties for the most serious violations, prohibited AI practices, reach the higher of 7%[7] of worldwide annual turnover or €35 million. Violations of high-risk AI system obligations carry up to 3% of worldwide turnover or €15 million.
For a $10B business, that is up to $300M in violations of high-risk AI system obligations you assumed the system was handling itself.
The directional truth is enough: catch the error at a human gate and the lag is near zero, so the cost stays at correction. Let it run autonomously and the compounding term is what bankrupts the ROI case. You end up funding a hidden data factory; a retrospective cleanup operation that costs far more than the checkpoint you skipped.
Deciding What to Automate and What to Own
The move here is not “automate less.” It is to stop treating the decision as binary and start spending human attention only where an error is unrecoverable. Three tiers do most of the work.
Human-Led—for high stakes. Irreversible, regulated, or value-laden decisions run on an interrupt-and-resume pattern: the workflow pauses for human verification before it proceeds. This is the model that Damco followed for the digital health firm and financial data intelligence firm, where a wrong call earlier invalidated a trial or corrupted a risk score, and structured human validation is the product, not an overhead line.
Human-Governed—for medium stakes. The model runs independently, but outputs route to review queues with a real-time override, and anything below a confidence threshold escalates automatically. This fits billing verification, service routing, and CRM enrichment where errors are costly but reversible, with throughput that still matters.
Machine-Led—for routine. Low-risk, easily reversed tasks run fully autonomous inside locked parameters. Humans set the boundaries at design time and audit periodically. Phone-number formatting, duplicate detection, low-priority routing are some examples of the tasks. The accepted trade is exposure to minor silent errors, which is why even this tier needs drift observability.
Get Clinically Defensible Data, Not Pixel-Motion Guesses
Be honest about what oversight costs, because it is not free. A synchronous human gate turns milliseconds into minutes, which makes it impractical on a high-throughput line. Reviewers staring at hundreds of consecutive model outputs start to rubber-stamp them, and confirmation bias lets subtle errors through. Different annotators interpret guidelines differently over time, and that labeling drift pollutes your retraining data. Human capacity does not scale linearly with a volume spike, either.
The taxonomy exists precisely so you absorb these costs only in the lane where the alternative is an unowned, unrecoverable failure and nowhere else.
Build AI You’re Willing to Stand Behind
The narrative that AI can run fully autonomous across high-variance enterprise data is not a productivity story. It is a deferral. You don’t eliminate the manual work; you postpone it, and you pay a premium when it comes back as forensic cleanup. MIT’s NANDA initiative found that 95%[8] of enterprise GenAI pilots deliver no measurable P&L impact, and miscalibrated autonomy is a recurring driver in the post-mortems.
So change the question. Not “Can we automate this?”—that one almost always answers yes. Ask instead: Who owns the consequence when this model is wrong, and how fast will they know?
The goal is not less automation. It is full automation everywhere the error is reversible, and a human gate only where it isn’t. Human-in-the-loop in the high-stakes lane isn’t a step backward—it is how you keep the efficiency without inheriting a liability nobody signed up to own.
References:







