Dens invaginatus as a diagnostic challenge: evaluating large language models against expert endodontic reasoning
Damla Erkal; Turgut Felek; Oana-Paula Butean; Kürşat Er · 2025 · BMC Oral Health
WASTE classifies this as Negative / Null Result Report · AI classification, approximate
The study found no significant effect — useful as a negative control or null benchmark for your own design.
Abstract
Abstract Introduction This study hypothesized that large language models (LLMs) would underperform compared to expert clinicians in diagnosing and managing complex endodontic anomalies, such as dens invaginatus, when provided with periapical radiographs. Although LLMs have shown promise in dental education and basic diagnostics, their effectiveness in nuanced clinical reasoning has remained unclear. Methods Nineteen anonymized periapical radiographs depicting challenging endodontic conditions were paired with clinical vignettes. Six advanced LLMs and one expert endodontist independently answer
Abstract by Damla Erkal; Turgut Felek; Oana-Paula Butean; Kürşat Er, BMC Oral Health (2025) — licensed CC BY 4.0.
About to run something similar?
Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.
Related failures
Accuracy of four intraoral scanners in oral implantology: a comparative in vitro study
Negative / Null Result ReportBioactivity evaluation of three calcium silicate‐based endodontic materials
Negative / Null Result ReportIncidence of impacted mandibular and maxillary third molars-a radiographic study in a Southeast Iran population
Negative / Null Result ReportAnti‐infective therapy of peri‐implantitis with adjunctive local drug delivery or photodynamic therapy: 12‐month outcomes of a randomized controlled clinical trial
Negative / Null Result ReportExploring the genetic basis of chronic periodontitis: a genome-wide association study
Negative / Null Result ReportIn Vitro Biofilm Formation on Titanium and Zirconia Implant Surfaces
WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.
Metadata source: DOAJ · DOI 10.1186/s12903-025-06987-z
