An AI stethoscope found heart defects in recordings doctors called unusable
A deep-learning model reached 80% accuracy on noisy heart sounds deemed non-diagnostic, while scoring 94.1% on its primary Bangladesh dataset.
OddBrief EditorialAI-assisted, human-reviewed
ScienceKey facts
- Input
- Raw 15-second phonocardiogram recordings
- Primary result
- 94.1% accuracy on the Bangladesh dataset
- Sensitivity and specificity: 92.7% and 96.3%
- Noisy subset
- 80% accuracy on recordings cardiologists deemed non-diagnostic
- Caveat
- Prospective clinical validation and broader population testing are still needed
A deep-learning system detected congenital heart disease in low-quality stethoscope recordings that cardiologists had judged non-diagnostic. The model reached 80 percent accuracy on that difficult subset, according to an open-access study published in Scientific Reports.
On the study's primary dataset from Bangladesh, the system achieved 94.1 percent accuracy, 92.7 percent sensitivity and 96.3 percent specificity. The researchers designed it as a screening tool for places where echocardiography and pediatric specialists are difficult to access.
Fifteen seconds of heart sound
The model analyzes phonocardiograms, digital recordings of the mechanical sounds produced by the heart. It works with raw 15-second segments rather than requiring a human to mark individual heart sounds in advance.
Researchers trained and evaluated the system using recordings from several points on the chest. Combining four listening locations produced the best result, but accuracy remained above 85 percent when the model received only one location.
That matters outside specialist clinics. Community health workers may have limited time and may not place a digital stethoscope with expert precision. Infants can cry or move, adding noise to an already subtle signal.
Performance survived some domain changes
The team also tested the model on public PhysioNet Challenge datasets. It reached 92 percent accuracy and a 94 percent area under the receiver-operating curve on the 2022 dataset.
Performance dropped on the 2016 collection, which includes adults, a wider range of heart disease and recordings from different devices and settings. The authors say reduced sensitivity under that domain shift is an important limitation.
The 80 percent result on recordings labeled unsatisfactory is striking, but it should not be read as proof that the system outperforms cardiologists in diagnosis. Experts declined to diagnose from those recordings because the signal was poor; the algorithm was evaluated against dataset labels in a research setting.
Screening is not a final diagnosis
Congenital heart disease often needs early treatment, while definitive diagnosis usually requires imaging and specialist assessment. A low-cost acoustic model could help decide which children should be referred first.
False negatives remain especially serious in a screening tool. The researchers propose broader validation across ages, devices and populations, along with signal-quality assessment and noise reduction. The current model also performs binary classification rather than identifying specific heart-defect types.
The work shows why digital stethoscopes are becoming an AI target: they turn an inexpensive clinical instrument into structured data. The next step is a prospective study in real clinics, where prevalence, operator training and referral capacity will determine whether benchmark accuracy becomes useful care.
Sources
- Congenital heart disease classification using phonocardiogramsScientific Reportsprimary source


