The Korean Journal of Internal Medicine

Search

Close

Original Article
Korean J Intern Med. 2026;41(5):862-872. Published online September 1, 2026.
DOI: https://doi.org/10.3904/kjim.2024.245
Clinical feasibility of deep learning-assisted classification of Helicobacter pylori infection in endoscopic imagery: a reader study
Jun-young Seo1,2, Jiseon Kang3, Do Hoon Kim1  , Namkug Kim4 
1Department of Gastroenterology, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
2Digestive Disease Center, CHA Bundang Medical Center, CHA University School of Medicine, Seongnam, Korea
3Department of Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
4Department of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Korea
Corresponding author: Do Hoon Kim ,Tel: +82-2-3010-3197, Fax: +82-2-3010-6517, Email: dohoon.md@gmail.com
Namkug Kim ,Tel: +82-2-3010-6573, Fax: +82-2-476-4719, Email: namkugkim@gmail.com
Received: July 14, 2024; Revised: January 20, 2026   Accepted: June 16, 2026.
Abstract
Background/Aims
Detecting Helicobacter pylori infection through endoscopic imaging is a preferred method to address the limitations of invasive diagnostics. We have developed a deep learning (DL) model to classify these images based on their H. pylori infection status and to evaluate its potential as a diagnostic aid.

Methods
We retrospectively enrolled H. pylori-positive patients (by rapid urease test and serum IgG) at Asan Medical Center between January 2015 and December 2020, establishing development and test datasets to evaluate model utility. We developed a DL model and assessed its performance at both the image and patient levels using metrics such as accuracy, sensitivity, specificity, and F1-score. Fourteen readers of varying experience levels interpreted the endoscopic images in two apsessions, before and after the integration of the DL model results. The diagnostic accuracy of the readers was analyzed using the McNemar test and generalized estimating equations.

Results
In the utility test set, the image-level accuracy, sensitivity, specificity, and F1-score of our DL model were 84.2%, 74.4%, 93.9%, and 82.5%, respectively. At the patient level, the corresponding values were 97.0%, 100%, 94.0%, and 97.1%. Endoscopists utilizing the DL model demonstrated significantly improved accuracy in classifying H. pylori infections, with classification rates of 83.2% versus 72.4% (p < 0.001).

Conclusions
Our DL model shows exceptional predictive capabilities for identifying H. pylori infections in gastroscopy images, suggesting the potential of DL-based tools to enhance clinical decision-making in endoscopy. Future research should focus on multicenter prospective studies to validate these findings.

Keywords :Artificial intelligence; Deep learning; Endoscope; Helicobacter pylori
Memo patch ics samyangbiopharm
Hanmi yungjin

Go to Top