Episode Details

Back to Episodes

243: Cytopathology AI: Crowded-Cell Gaps and LLM Guardrails

Episode 243 Published 12 hours ago
Description

Send us Fan Mail

What happens when AI models that perform almost perfectly on scattered cervical cells become less reliable than a coin flip on crowded cell groups?

In DigiPath Digest #51, I examine what this performance gap tells us about artificial intelligence in cytopathology.

The first paper evaluated six convolutional neural network models trained to distinguish benign from high-grade lesions using scattered cervical cytology cells. The models achieved AUCs ranging from 0.950 to 0.996 on the original images.

When the same models were applied to hyperchromatic crowded cell groups without retraining or threshold recalibration, their AUCs fell to 0.385-0.683. Different architectures also failed differently. Some overcalled benign clusters, while others missed high-grade clusters.

This isn’t simply a technical problem. It illustrates a practical rule for pathology AI: a model should only be trusted for the morphology, specimen type, imaging system, and intended use on which it has been directly validated.

I also review the current evidence for large language models in cytopathology. Potential applications include structured reporting, diagnostic support, quality control, education, research, and workflow integration.

Some early results appear promising, but each comes with important limitations. One structured-reporting application achieved 99.4% accuracy at a single institution. A diagnostic-support model included the correct answer among its top 10 differentials in 59.1% of general medicine cases. A hybrid quality-control system flagged 84% of errors associated with amended reports, but its false-positive rate wasn’t reported.

Most importantly, the review found no language model specifically trained and clinically validated on cytopathology reports.

The takeaway is straightforward: we’re still working with narrow AI. Strong performance in one setting doesn’t guarantee performance when the cells, preparation, scanner, institution, or clinical task changes.

Low-risk applications may offer the most practical starting point. Text extraction, completeness checks, report consistency review, and quality-control flagging could reduce repetitive work without asking an unvalidated model to make the final diagnosis.

Highlights with timestamps

  • 00:00 - Welcome to DigiPath Digest #51 and the new lunch-and-learn time
  • 01:40 - How two image models performed worse than a coin flip on cell clusters
  • 02:29 - Image models, language models, and vision-language models
  • 04:40 - Why AI adoption in cytopathology remains low
  • 05:25 - Scattered single cells versus hyperchromatic crowded cell groups
  • 08:14 - AUCs fall from 0.950-0.996 to 0.385-0.683
  • 10:02 - How ResNet-50 and GoogLeNet failed differently
  • 11:04 - What the attention maps revealed
  • 13:22 - The intended-use lesson for pathology AI
  • 17:02 - Current applications of large language models in cytopathology
  • 18:39 - Structured reporting and the 99.4% accuracy result
  • 19:32 - Diagnostic support, AMIE, and the top-10 limitation
  • 21:09 - Quality-control applications and the missing false-positive rate
  • 23:30 - Why cytopathology still needs domain-specific language models
  • 24:54 - Retrieval-augmented generation, education, and research support
  • 27:07 - Two AI families, one requirement: direct validation
  • 28:35 - Context of use and intended-use validation
  • 29:12 - Human-reviewed training data and destructive book scanning
  • 32:02 - FDA and European approaches to evolving AI models
  • 35:00 - Protecting patient and practitioner well-being
  • 36:21 - Cy
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us