-
Original Article
-
Korean J Intern Med. 2026;41(5):862-872. Published online September 1, 2026.
DOI: https://doi.org/10.3904/kjim.2024.245
- Clinical feasibility of deep learning-assisted classification of Helicobacter pylori infection in endoscopic imagery: a reader study
-
Jun-young Seo1,2, Jiseon Kang3, Do Hoon Kim1
, Namkug Kim4
-
1Department of Gastroenterology, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
2Digestive Disease Center, CHA Bundang Medical Center, CHA University School of Medicine, Seongnam, Korea
3Department of Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
4Department of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
- Corresponding author: Do Hoon Kim ,Tel: +82-2-3010-3197, Fax: +82-2-3010-6517, Email: dohoon.md@gmail.com
Namkug Kim ,Tel: +82-2-3010-6573, Fax: +82-2-476-4719, Email: namkugkim@gmail.com
- Received: July 14, 2024; Revised: January 20, 2026 Accepted: June 16, 2026.
- Abstract
- Background/Aims
Detecting Helicobacter pylori infection through endoscopic imaging is a preferred method to address the limitations of invasive diagnostics. We have developed a deep learning (DL) model to classify these images based on their H. pylori infection status and to evaluate its potential as a diagnostic aid.
Methods
We retrospectively enrolled H. pylori-positive patients (by rapid urease test and serum IgG) at Asan Medical Center between January 2015 and December 2020, establishing development and test datasets to evaluate model utility. We developed a DL model and assessed its performance at both the image and patient levels using metrics such as accuracy, sensitivity, specificity, and F1-score. Fourteen readers of varying experience levels interpreted the endoscopic images in two apsessions, before and after the integration of the DL model results. The diagnostic accuracy of the readers was analyzed using the McNemar test and generalized estimating equations.
Results
In the utility test set, the image-level accuracy, sensitivity, specificity, and F1-score of our DL model were 84.2%, 74.4%, 93.9%, and 82.5%, respectively. At the patient level, the corresponding values were 97.0%, 100%, 94.0%, and 97.1%. Endoscopists utilizing the DL model demonstrated significantly improved accuracy in classifying H. pylori infections, with classification rates of 83.2% versus 72.4% (p < 0.001).
Conclusions
Our DL model shows exceptional predictive capabilities for identifying H. pylori infections in gastroscopy images, suggesting the potential of DL-based tools to enhance clinical decision-making in endoscopy. Future research should focus on multicenter prospective studies to validate these findings.
Keywords :Artificial intelligence; Deep learning; Endoscope; Helicobacter pylori