Back

Diagnosis of Pathological Speech with Efficient and Effective Features for Long Short-Term Memory Learning

Pham, T. D.; Holmes, S.; Zou, L.; Patel, M.; Coulthard, P.

2023-09-04 dentistry and oral medicine
10.1101/2023.09.04.23295008 medRxiv
Show abstract

The majority of voice disorders stem from improper vocal usage. Alterations in voice quality can also serve as indicators for a broad spectrum of diseases. Particularly, the significant correlation between voice disorders and dental health underscores the need for precise diagnosis through acoustic data. This paper introduces effective and efficient features for deep learning with speech signals to distinguish between two groups: individuals with healthy voices and those with pathological voice conditions. Using a public voice database, the ten-fold test results obtained from long short-term memory networks trained on the combination of time-frequency and time-space features with a data balance strategy achieved the following metrics: accuracy = 90%, sensitivity = 93%, specificity = 87%, precision = 88%, F1 score = 0.90, and area under the receiver operating characteristic curve = 0.96.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.