Healthcare & life sciences

Imaging triage for screening programmes

DICOM-native models that re-order the reading worklist, so scarce radiologist attention lands on the studies most likely to matter.

Concentric arcs with a small circular region marked for attention
Typical pilot
10–14 weeks to a retrospective validation
Benchmark
RSNA Screening Mammography Breast Cancer Detection
Benchmark host
Radiological Society of North America
Sector
Healthcare & life sciences
What this is — and is not

This page describes a problem class and the architecture we deploy for it, written against a named public benchmark — the open dataset or competition where the world's data scientists have tested approaches to this exact problem against a hard metric. It is not a client engagement, and the benchmark's results, prizes and rankings belong to its host and participants, not to us. Client work is confidential and is only ever published with written permission.

The problem

What it costs when this goes unsolved.

Screening programmes read enormous volumes to find rare positives, against a global shortage of radiologists. The failure modes pull in opposite directions — a missed finding is harm deferred to the patient, an over-call is an unnecessary biopsy — and both are worsened by fatigue. Triage AI does not replace the reader; it decides what the reader sees first, which is the deployment pattern that clinical imaging AI has actually earned.

The data reality

Medical imaging arrives as DICOM with vendor-specific quirks, embedded patient identifiers that must be stripped at ingest, and labels of varying epistemic quality — radiologist assessment is not pathology follow-up. Positives are severely rare, so a naive validation split flatters the model. Getting the pre-processing and the split right is most of the engineering; the architecture is the published part.

How we build it

The architecture, stage by stage.

01

DICOM-native ingest with de-identification

Pixel pipelines that handle windowing, vendor variation and identifier stripping at the boundary — PHI never reaches the training environment.

02

Region-of-interest cascade

Detection models localise candidate regions before classification networks grade them, with multi-view fusion where the modality provides it.

03

Leak-free validation by patient

Splits stratified at patient level and grouped so no individual appears on both sides — the validation discipline this benchmark class established.

04

Triage, not diagnosis

Output is a worklist ordering with per-study evidence overlays a clinician can verify, framed for second-reader support — the claim that matches both the evidence and the regulatory reality.

Model families on this problem class: DICOM pipelines · Detection → classification cascades · ConvNeXt / EfficientNet · Patient-level validation.

What you get

What an engagement hands over.

Everything below goes in the scope document before you sign it, with a fixed price or a rate with a ceiling — the same terms as every other engagement in the catalogue.

  • A triage model with per-study evidence overlays, validated leak-free at patient level
  • A DICOM ingest pipeline with de-identification your DPO can audit
  • A validation report written for clinical governance review
  • A deployment plan that keeps the radiologist as the decision-maker — including what the model must not be used for

Provenance

The benchmark behind this page.

RSNA Screening Mammography Breast Cancer Detection, run by Radiological Society of North America, is the public proving ground for this problem class. The figures below are the host's, cited as context for how seriously this problem is tested in the open — they are not our results and we do not claim them.

Benchmark scale
1,687 teams · US $50,000 prize · real screening-programme DICOM
Standing
RSNA's largest challenge since 2017; results analysed in the journal Radiology
Metric
Probabilistic F1 on a severely imbalanced target
Reference
Benchmark page ↗

Next step

Thirty minutes on whether this fits your problem.

Bring the constraint — the regulator, the data boundary, the latency budget. If your data cannot support this build yet, the call will conclude with what to fix first, not with a proposal.