
Benchmarking AI Risk Estimates in Prostate Cancer Against STAR-CAP
Daniel Spratt, MD, explains what the STAR-CAP cohort is, what an MMAI risk estimate represents, and why calibration benchmarking differs from clinical validation.
Daniel Spratt, MD, introduces the benchmarking analysis, which compares real-world multimodal AI (MMAI) results against the STAR-CAP cohort.1 He explains that STAR-CAP was developed with collaborators across dozens of community and academic centers, including VA hospitals, in the United States and several other countries. The cohort paired long-term clinical outcomes with standard risk factors—prostate-specific antigen level, T stage, nodal stage, number of positive biopsy cores, and grade—and was used to build the STAR-CAP staging system, which Spratt characterizes as a gold standard for risk stratification using historical tools.
Because STAR-CAP includes approximately 20,000 real patients with long-term outcomes who were treated with surgery, radiation, hormone therapy, or brachytherapy across many centers, Spratt says it provides a benchmark for testing whether a new tool's predicted risks match what patients actually experience.
He then explains the MMAI tool itself: a model that incorporates clinical and digital pathology information to generate a score, which can be thought of as ranging from 0 to 100 and grouped into low-, intermediate-, and high-risk categories. Beyond the category, Spratt notes, the report provides an absolute risk of developing metastatic disease or dying from prostate cancer at 10 or 15 years, which clinicians can use when weighing intensification or de-escalation.
Finally, Spratt distinguishes benchmarking from clinical validation. Validation studies—many of which, for MMAI, used randomized trial data—test whether a tool improves discrimination, typically reported as hazard ratios, area under the curve, or C-index. A tool can discriminate well, he explains, yet still misstate absolute risk, reporting a 20% risk of metastasis when the true risk is 5%. Benchmarking tests calibration: whether the number on the report is likely to be accurate.
Related to this article








