Comparative accuracy of two commercial AI algorithms for musculoskeletal trauma detection in emergency radiographs - UTU Research Portal

A1 Refereed original research article in a scientific journal

Comparative accuracy of two commercial AI algorithms for musculoskeletal trauma detection in emergency radiographs

Authors: Huhtanen, Jarno T.; Nyman, Mikko; Sequeiros, Roberto, Blanco; Koskinen, Seppo K.; Pudas, Tomi K.; Kajander, Sami; Niemi, Pekka; Aronen, Hannu J.; Hirvonen, Jussi

Publisher: SPRINGER HEIDELBERG

Publishing place: HEIDELBERG

Publication year: 2025

Journal:: Emergency Radiology

Journal name in source: EMERGENCY RADIOLOGY

Journal acronym: EMERG RADIOL

Number of pages: 12

ISSN: 1070-3004

eISSN: 1438-1435

DOI: https://doi.org/10.1007/s10140-025-02353-2

Web address : https://link.springer.com/article/10.1007/s10140-025-02353-2

Self-archived copy’s web address: https://research.utu.fi/converis/portal/detail/Publication/498946096

Abstract

Purpose: Missed fractures are the primary cause of interpretation errors in emergency radiology, and artificial intelligence has recently shown great promise in radiograph interpretation. This study compared the diagnostic performance of two AI algorithms, BoneView and RBfracture, in detecting traumatic abnormalities (fractures and dislocations) in MSK radiographs.

Methods: AI algorithms analyzed 998 radiographs (585 normal, 413 abnormal), against the consensus of two MSK specialists. Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and interobserver agreement (Cohen's Kappa) were calculated. 95% confidence intervals (CI) assessed robustness, and McNemar's tests compared sensitivity and specificity between the AI algorithms.

Results: BoneView demonstrated a sensitivity of 0.893 (95% CI: 0.860-0.920), specificity of 0.885 (95% CI: 0.857-0.909), PPV of 0.846, NPV of 0.922, and accuracy of 0.889. RBfracture demonstrated a sensitivity of 0.872 (95% CI: 0.836-0.901), specificity of 0.892 (95% CI: 0.865-0.915), PPV of 0.851, NPV of 0.908, and accuracy of 0.884. No statistically significant differences were found in sensitivity (p = 0.151) or specificity (p = 0.708). Kappa was 0.81 (95% CI: 0.77-0.84), indicating almost perfect agreement between the two AI algorithms. Performance was similar in adults and children. Both AI algorithms struggled more with subtle abnormalities, which constituted 66% and 70% of false negatives but only 20% and 18% of true positives for the two AI algorithms, respectively (p < 0.001).

Conclusions: BoneView and RBfracture exhibited high diagnostic performance and almost perfect agreement, with consistent results across adults and children, highlighting the potential of AI in emergency radiograph interpretation.

Downloadable publication

This is an electronic reprint of the original article.
This reprint may differ from the original in pagination and typographic detail. Please cite the original version.

Huhtanen_etal_Comparative_accuracy_2025.pdf

Funding information in the publication:
Open Access funding provided by University of Turku (including Turku University Central Hospital). This research was supported by Radiological Society of Finland