Unit 7 / 11

Modality-Specific Artificial Intelligence: Lung X-ray, CT, MRI, Mammography and Stroke

Gains:

  • Ability to distinguish typical usage areas and modality-specific limits of artificial intelligence applications in different modalities (graphia, CT, MR, mammography)
  • Ability to evaluate the role of AI as a second-reader and clinical implications of decision thresholds in high-impact fields such as mammography and stroke CT
  • Understanding that performance may decrease (distribution shift) when the population and device on which artificial intelligence is trained in each modality are moved beyond

Artificial intelligence is not just one thing in radiology; For each modality (imaging method), it carries a different tool, a different power and a different limit. A chest X-ray model and a stroke CT model solve completely different problems and make different error patterns. In this unit, we will cover the five most common areas—chest radiography, CT, MRI, mammography, and stroke—with modality-specific uses and pitfalls. The aim is to use each tool for the right job, with the right limits.

The core principle does not change but takes on a new face in each modality: Each model is trained on a specific device, a specific protocol, and a specific patient population. When these conditions are exceeded, performance may decrease; This is called distribution shift. The diagnosis belongs to the radiologist, who is the first-reader, not the second-reader.

Use and limits by modality

Lung radiography (x-ray): The most abundant data is here; models generate flags for pneumothorax, consolidation (inflammatory fullness of the lung), nodule, pleural effusion (interpleural fluid), and heart size. Its strengths are speed and scanning; its weakness is the high false positives due to overlapping structures (bone, vessel, breast shadow) and the tendency to miss small nodules.

CT (computed tomography): Thanks to the cross-sectional image, the models are powerful for embolism, bleeding, nodule, fracture, organ volume. Its strength is three-dimensional measurement; The weakness is the high sensitivity to protocol/slice thickness and contrast phase—performance varies with different protocols.

MRI (magnetic resonance): Soft tissue contrast is high; There are models for stroke diffusion, prostate lesion, brain lesions, cartilage. Its weakness is that there is too much device/sequence diversity — a model loses reliability outside of the sequence on which it was trained.

Mammography: A high-impact field; The models act as second-readers for masses and microcalcifications. Its strength is reducing misses in breast cancer screening; Its weakness is the performance degradation in dense breast tissue and the recall overhead of false positives.

Stroke (brain CT/CT angiography/perfusion): This is the time-critical area; models provide rapid flags for large vessel occlusion (LVO), early signs of ischemia, and bleeding, accelerating treatment (thrombectomy) decisions. Its weaknesses are movement artifact and limited reliability in early/faint findings.

modality

Typical AI usage

strong side

Modality specific limit

Lung X-ray

Pneumothorax, nodule, effusion flag

speed, scan

Overlap, high false positives

IT

Embolism, bleeding, nodule, volume

3D measurement

Sensitive to protocol/contrast phase

MRI

Stroke diffusion, prostate, brain

soft tissue

Sequence/device diversity

mammography

Mass, microcalcification (2nd reader)

Reduce incontinence

Drop in dense breast, recall load

Stroke CT

LVO, ischemia, bleeding flag

time saving

Artifact, early/faint finding

Second-reader role and decision threshold

In high-impact fields such as mammography and stroke, the AI is often positioned as a second-reader: the radiologist reads, then what the AI marks is reviewed (or vice versa). This gives a chance to catch what the single radiologist missed. But there are two risks. First, the area not marked by the AI ​​is considered "clean" (automation bias). Second, the clinical consequence of the decision threshold: if the threshold is low in mammography, there will be more recalls (and unnecessary biopsies, anxiety); If it is high, leakage increases. This threshold is set by the institution's screening philosophy and the radiologist's interpretation; is not an engineering default.

Caution: Just because a model works well in one modality does not mean that it will work well in another modality or on a different device/protocol of the same modality. If the model is trained for "lung radiography", its performance may differ in portable critical care radiography. The scope of use is always limited to the scope for which the model has been approved.

three mini cases

Case 1 — Mammography gains second-readers. A radiologist reads a screening mammogram normally. YZ as the second-reader marks a cluster of microcalcifications in a region. The radiologist looks back and sees a really faint cluster; With additional imaging, an early ductal carcinoma in situ (early stage breast cancer) is detected. AI prevented escape as a “second eye”; The radiologist gave the diagnosis.

Case 2 — Stroke time gain. In an emergency brain CT angiogram, the AI ​​flags a large vessel occlusion (LVO) within 2 minutes and triggers the stroke team. The radiologist confirms, the patient is taken to thrombectomy; The door-to-needle time is significantly shortened. But in the same week, the model misses a faint early ischemia; The radiologist captures it with his own reading. AI brought speed, but the first-reader remained the radiologist.

Case 3 — Distribution shift. A chest x-ray model works perfectly on stationary equipment in a large hospital. When the same model is used on an intensive care unit's portable devices, false positives explode: different exposure, poor quality and overlapping equipment shadows confuse the model. The team realizes that they are using the model on portable graphs without local validation and narrows down the scope. Lesson: performance changes when device/protocol changes.

Weak prompt / Strong prompt

Weak prompt:

Is this MRI model good? Should we use it?

No modality details, sequence, population and intended use; A general and risky answer comes.

Powerful prompt:

Your role: CO-consultant in modality-specific AI selection. Decision making; Come up with what questions I should ask. We are considering a prostate MRI lesion model. Consider: (1) what sequences/device/magnetic field strength the model was trained on, does it match ours, (2) in which population has it been validated, (3) is it a second-reader or diagnostic aid, (4) is it at its known limit in dense/atypical cases, (5) is the risk of using it without local validation. The final decision lies with the institution.

The strong prompt questions modality-specific compliance, population, and local validation.

Copiable prompt templates

MODALITY COMPLIANCE TEMPLATEYour role: ASSISTANT. I will give you an AI model and our own device/protocol information. Compare ours with the modality, device, sequence/protocol, and population on which the model was trained; List differences at risk of distribution drift and evaluate whether local verification is required. The decision is in the institution. Information: [write]

SECOND-READER DISCIPLINE TEMPLATEProduce me a checklist that will maintain my reading discipline when using second-reader AI: should I read myself first, how do I protect the area that the AI does not mark, how do I confirm the AI mark, how do I avoid automation bias. Modality: [write]

DECISION THRESHOLD RESULT TEMPLATEI will give you a decision threshold scenario in a modality (e.g. mammography). Explain the clinical consequences of lowering/raising the threshold on recall, unnecessary biopsy, evasion, and anxiety. Remind us that the threshold is set by the institution's screening philosophy. Script: [write]

MODALITY LIMIT REMINDER TEMPLATEI will give you a modality and an AI task. List typical sources of error specific to this modality (e.g., overlap in chest X-ray, sequence difference in MRI, artifact in stroke CT) and points for which the radiologist should pay extra attention. Modality/task: [write]

Common mistakes

  • Using a modality-specific model in another condition. The stationary device model behaves differently in the portable; scope must be preserved.
  • Counting as "clean" what the second-reader did not mark. Automation bias; Their reading should be protected.
  • Mistaking the decision threshold as an engineering setting. The threshold is a clinical balance; It is adjusted by scanning philosophy.
  • Ignoring the sequence/protocol difference. MR/CT models lose reliability outside the protocol on which they were trained.
  • Dissemination without local verification. Distribution drift silently degrades performance.
Tip: Before using a new modality-specific model, ask three questions: "What device/protocol/population was this model trained on? Does ours look like it? Have we validated it on our own data?" If all three are not satisfactory, the model is not an accelerator but a hidden risk.

In summary

AI is a different tool in each modality: rapid screening in chest X-ray but false positive due to overlap; Powerful 3D measurement but protocol sensitivity in CT; Soft tissue strength but sequence diversity in MRI; second-reader reducing abduction but recall burden in mammography; Time gain in stroke but artifact limit. In high-impact fields such as mammography and stroke, the AI ​​becomes the second-reader, but the first-reader is always the radiologist and the decision threshold is a clinical balance. When the device, protocol and population on which each model is trained are moved out of range, its performance decreases (distribution shift); so scope must be maintained and local validation must be performed.

Application task

Choose a modality in your routine (radiography, CT, MRI, mammography, or stroke). Using the “Modality Limit Reminder” template, write down at least three typical sources of AI error specific to that modality and points of extra attention for the radiologist. Then evaluate a model you are considering for that modality with the “Modality Fit” template: does the device/protocol/population it is trained on match yours, is local validation required? If there is a second-reader role, adapt the "Second-Reader Discipline" list into your routine.

checklist

  • [ ] I learned the modality, device, protocol and population on which the model was trained.
  • [ ] I evaluated the compatibility with my own conditions (risk of distribution drift).
  • [ ] I use the model only to the extent it has been approved.
  • [ ] I maintain my own independent reading in the role of second-reader.
  • [ ] I understand that the decision threshold is a clinical balance and is set by screening philosophy.
  • [ ] I consider typical modality-specific error sources (artifact, overlap, sequence difference).
  • [ ] I don't propagate in new condition without local validation.