Gains:
- Ability to explain the concepts of false positive and false negative with examples from dentistry and evaluate their clinical consequences on the patient.
- Ability to read indicators such as AI confidence score, sensitivity and specificity at a basic level and decide how much weight the output should carry
- Ability to establish a protocol that confirms each radiographic AI finding with clinical correlation, additional images, and second physician opinion when necessary
In the previous unit, we saw how AI radiography tools work. In this unit we get to the heart of the matter: how accurate are these tools, when do they go wrong, and what are the consequences of these mistakes for the patient? Because whether a tool is "helpful" or "dangerous" is determined not by when it works correctly, but by what happens when it goes wrong. Understanding the two basic error types of AI in dentistry — false positives and false negatives — is a prerequisite for safe use.
Two basic types of errors
False positive: When the AI marks something as existing when it doesn't actually exist. For example, it shows a healthy tooth as "decay". The result: unnecessary anxiety, unnecessary further examination, and even the risk of loss of healthy tissue due to incorrect treatment.
False negative: When the AI "ignores" something that actually happened, that is, does not flag it. For example, bypassing an existing periapical lesion (inflammatory area at the root tip). Result: missing pathology, delay in treatment, progression of the disease. This is often the most dangerous mistake in dentistry; because if the physician trusts the AI and says "clear", the real pathology will proceed silently.
Caution: Just because the AI doesn't flag anything does NOT mean "all is well". A false negative is always possible. Clinical examination and physician reading cannot be skipped under any circumstances.
Sensitivity and specificity: two key indicators
Two numbers summarize a vehicle's performance:
- Sensitivity: The rate of catching those who are truly sick. High sensitivity reduces false negative. It means "He does not miss the disease".
- Specificity: The rate at which the truly healthy ones are accurately considered "clean". High specificity reduces false positives. It means "He doesn't raise an alarm for nothing."
These two are often at odds with each other. If you set a tool too “sensitive” (it flags every suspicion) it will miss fewer pathologies but give a lot of false alarms. If you set it too "selective", false alarms are reduced but the risk of missing real pathology increases. Clinically, high sensitivity is often preferred in dentistry — because missing a lesion is riskier than a blank alarm. But the final say is again in the clinical examination.
concept
What measures
What does being high provide?
Sensitivity
Capture the patient
False negatives decrease
specificity
Cleansing the healthy
False positives decrease
Confidence score
The model's "confidence" in that sample
Tip for ranking/threshold only
Tip: If a tool only shows you a confidence score, ask the manufacturer for the sensitivity/specificity values behind that score. These values are meaningful in which population and how they are tested.
Validation protocol: for each finding
Verify each AI radiography finding with these steps:
- Clinical correlation: Does the finding match the patient's symptoms and examination?
- Additional imaging: Confirmation with periapical, bitewing or CBCT if necessary.
- Comparison over time: Compare with previous radiographs and evaluate change, if any.
- Second physician opinion: Colleague opinion in doubtful or high-risk cases.
- Record: Write the decision, rationale, and AI's role clearly in the patient chart.
Mini case 1: The cost of a false positive
The AI tool flags “caries” in the lower first molar in a 34-year-old patient with 88% confidence. On clinical examination, the tooth is intact, no catheter is inserted, and the patient has no complaints. The physician confirms with a bite film: there is no caries, the mark is due to the radiographic shadow of a restoration (filling). If the physician had relied directly on AI and intervened, there would have been loss of healthy tissue. False positives were prevented here only by the clinical control of the physician.
Mini case 2: The danger of false negative
A 62-year-old patient complains of a vague fullness in the lower jaw. AI does not mark anything in panoramic scan. The young doctor is relieved, "AI said it's clean." The senior physician insists and asks for clinical examination and CBCT; A lesion appears at the root tips that AI missed. Lesson: AI's silence is never a diagnosis; Clinical suspicion always takes precedence.
Mini case 3: Effect of threshold setting
A clinic tests the vehicle's marking threshold. When the threshold is lowered to 50%, the tool gives 900 signals per month; only 25% of these are clinically significant, the rest are false positives—physicians are overwhelmed. When the threshold is increased to 80%, the number of markers drops to 320, significance increases to 55%, but two true early lesions remain below the threshold and escape. The clinic finds balance at the 70% threshold and puts low-scoring areas on the “review” list. Lesson: no single threshold is perfect; The physician's eye closes the gaps.
Copiable templates
Use without adding real patient ID.
Role: Confirmation coach (non-diagnostic). Task: Produce a verification checklist for the following AI radiography finding: clinical correlation questions, additional images that may be recommended, differential diagnosis headings, note items for recording. Final decision WRITING. Finding: [anonymous]
Role: Error type descriptive.Task: Determine whether the following scenario carries the risk of a false positive or false negative and explain the likely clinical outcome to the patient. Recommend what action the physician should take.Scenario: [write]
Role: Indicator descriptor. Task: Interpret in plain language sensitivity, specificity, and test population information provided by a manufacturer; List questions for me to question whether these values fit our patient profile.Values: [paste]
Role: Second opinion request writer. Task: Draft a brief, professional, anonymized message requesting a second radiographic opinion from a colleague. Do not include patient identification. Context: [anonymous finding]
Weak prompt / Strong prompt
Weak: "AI said 90%, isn't this rotten certain?"
Why it's weak: Mistakes the confidence score for evidence, skips clinical validation, and delegates the decision to AI.
Strong: "For this area that the AI has marked with 90% confidence, what clinical and radiographic steps should I follow to confirm or exclude the diagnosis of caries? Also consider the possibility of a false positive."
Why it is powerful: It does not consider the score as absolute proof, it removes verification steps and takes the possibility of error into account.
Automation bias: the most insidious risk
False positives and false negatives are technical errors; But the most insidious risk is on the human side: automation bias. This is the tendency to trust the output of a tool more than our own judgment. As the tool turns out to be “mostly accurate” over time, the physician can stop questioning it and blindly follow the small number of critical cases where the tool is inaccurate. The paradox is this: the better the tool works, the more human attention risks atrophy and the more dangerous rare errors become.
The practical way to overcome this bias is to set up the workflow in a “me first, tool second” format: read the image yourself first, make your decision, then compare it with the AI. Thus, instead of guiding you, the tool becomes a second opinion that complements your reading. I also regularly ask “Could AI have been wrong in this case?” Asking, keeps attention alive. Asking this question is the strongest defense against automation bias, even when the tool has a high trust score.
Mini case 4: Atrophy of attention
In one clinic, the AI tool works very accurately for months, and physicians gradually begin to accept its output without questioning. One day the vehicle passes "clean" a real lesion that looks atypical; Because of the trust that has become a habit, the first doctor does not notice it either. Fortunately, a careful clinical examination performed upon the patient's persistent complaint reveals the lesion. The clinic then enforces the "physician reading first, then AI" rule. Lesson: the success of the tool should not replace human attention; the workflow must ensure this.
Common mistakes
- Downplaying the false negative and deeming AI silence as “healthy.”
- Interpreting the confidence score independently of sensitivity/specificity.
- Using the tool without knowing that it has been tested in a population different from your own patient profile.
- Not seeking a second opinion in high-risk findings.
- Not writing the role of the AI and the rationale for the decision in the patient chart.
In summary
AI radiography tools make two types of errors: false positive (showing what isn't there) and false negative (missing what is). In dentistry, a false negative is often more dangerous because the physician may trust and miss the pathology. Sensitivity and specificity are key to understanding these errors; The confidence score alone is not evidence. Each finding should be confirmed by clinical correlation, additional images, time comparison and, when necessary, a second opinion, and the decision and rationale should be recorded.
Application task
Choose 3 real examples from your own archive (anonymous): one where AI made a false positive, one where it made a false negative, one where it was correct. For each, write down the error type, clinical outcome, and correct verification steps. Turn the result into a one-page “radiographic AI validation protocol” to be implemented in your clinic.
checklist
- [ ] I can distinguish the concepts of false positive and false negative with examples.
- [ ] I can interpret sensitivity and specificity at a basic level.
- [ ] I perform the clinical correlation and additional image step for each finding.
- [ ] I get a second opinion in a high-risk situation.
- [ ] I record the decision, its rationale, and the AI's role in the patient chart.