A research team led by Shannon L. Walston at Osaka Metropolitan University examined how commercial radiology AI products are validated for potential bias across different patient populations. The researchers reviewed 545 studies covering 252 commercial products to assess whether developers tested their AI systems for performance differences based on sex, age and ethnic demographics.
Only 77 of the 545 studies, validating 52 products, included detailed demographic analysis and subgroup performance results. When the researchers applied statistical methods to studies on AI designed to detect tuberculosis, they found that 67 percent of reported datasets risked being too small to reliably measure differences in performance between male and female patients.
The findings suggest that demographic reporting for commercial medical AI has not improved despite growing calls for such validation. "Reporting for both demographics and per-subgroup performance is inadequate for estimating subgroup bias," Walston said. The issue has potential consequences as AI-driven decisions become more common in patient care.
The research was published in the journal European Radiology. Walston said the problem requires action from multiple sectors, including researchers, regulatory agencies and manufacturers, to ensure "thorough reporting and commercial product validation to support physician and patient trust in medical AI products."
