AI-powered detection tools have become one of the most closely watched additions to the endoscopy suite, promising to catch precancerous polyps that even experienced endoscopists might miss.
But a wave of new research is complicating that promise, raising questions about whether those detection gains are translating into better patient outcomes — and whether reliance on the technology is quietly eroding the skills of the clinicians using it.
Ignasi Puig, MD, PhD, and Maria Pellisé, MD, PhD, describe these concerns in a commentary published in Medscape Sept. 18. They write that computer-aided detection technology, known as CADe, increases the number of adenomas endoscopists detect during colonoscopies performed as part of fecal immunochemical test-positive screening. But in high-performing screening programs, that boost hasn’t extended to advanced adenomas, the lesions most closely tied to cancer risk.
The authors pointed to an observational study showing a reduced adenoma detection rate when endoscopists worked without CADe after becoming accustomed to using it — a pattern they described as a potential deskilling signal. A similar dynamic emerged in the CADILLAC trial, where CADe improved detection of small, proximal and nonpolypoid lesions but showed no improvement in detecting advanced colorectal neoplasia.
“Artificial intelligence is not a universal solution, but a context-sensitive tool whose value depends on baseline performance, operator expertise, and clinically meaningful outcomes,” Dr. Puig and Dr. Pellisé said.
The European Society of Gastrointestinal Endoscopy has taken a similarly cautious stance, according to the authors, issuing a weak recommendation for CADe use that cites uncertain effects on colorectal cancer incidence and mortality alongside the added detection and surveillance burden the technology creates.
How much of that burden gets absorbed by AI, and how much stays with the endoscopist, is itself an open question, according to a commentary in Gastroenterology led by Alan Barkun, MD. The benefits CADe demonstrates in randomized trials don’t always hold up in daily practice, Dr. Barkun and his co-authors wrote, because how much detection work an endoscopist delegates to the technology depends on the tool’s performance and how much the endoscopist trusts it.
That trust is often undermined by false alarms. Some studies have found CADe systems generate roughly 26 to 27 false-positive alerts per colonoscopy, adding noise and cognitive load that has led some clinicians to turn the feature off altogether. Interviews conducted by Dr. Barkun’s team found delegation to CADe falls along a career-stage continuum: Early-career endoscopists tend to respond to most alerts, mid-career clinicians treat the tool as a safety net once they’ve learned to filter out false positives, and many late-career endoscopists disengage from it entirely, judging the added cognitive load not worth the benefit.
“False-positive alerts and associated cognitive load should be actively monitored, as they directly influence trust and sustained use. Understanding and designing [AI tools] for specific delegation patterns may constitute the key to consistent effectiveness,” Dr. Barkun and his co-authors said.
A third commentary, published in Digestive and Liver Disease by Federico Cabitza, PhD, and colleagues, argues that inconsistent evaluation practices are compounding the problem. Clinical evaluation of AI tools should be anchored to prespecified, clinically justified decision thresholds, the authors wrote, rather than judged on generic accuracy metrics that can obscure how a tool actually performs at the point of care.
One method the authors point to is what they call the K-heuristic: asking how many false positives, or unnecessary alarms and procedures, clinicians are willing to accept in order to avoid missing one true case of disease. Thresholds set using this approach are rarely above 50% and sometimes as low as 1%, depending on risk tolerance. The stakes of getting that threshold wrong are significant in low-prevalence settings. A CADe system with roughly 90% sensitivity and 80% specificity, applied to 1,000 patients at a 0.5% cancer prevalence, would produce about 199 false positives for every 0.5 false negatives avoided, according to the commentary.
At the Becker’s 32nd Annual Meeting: The Business and Operations of ASCs, taking place October 29-31 in Chicago, ASC leaders, surgeons and healthcare executives will explore strategies to drive growth, enhance operational performance, navigate reimbursement challenges and prepare for the future of ambulatory surgery. Apply for complimentary registration now.
