TY - GEN
T1 - Evaluation of deep clustering for diarization of aphasic speech
AU - Klischies, Daniel
AU - Kohlschein, Christian
AU - Werner, Cornelius J.
AU - Jonas, Stephan M.
N1 - Publisher Copyright:
© 2019 The authors, AIT Austrian Institute of Technology and IOS Press.
PY - 2019
Y1 - 2019
N2 - Speaker attribution and labeling of single channel, multi speaker audio files is an area of active research, since the underlying problems have not been solved satisfactorily yet. This especially holds true for non-standard voices and speech, such as children and impaired speakers. Being able to perform speaker labelling of pathological speech would potentially enable the development of computer assisted diagnosis and treatment systems and is thus a desirable research goal. In this manuscript we investigate on the applicability of embeddings of audio signals, in the form of time and frequency-band based segments, into arbitrary vector spaces on diarization of pathological speech. We focus on modifying an existing embedding estimator such that it can be used for diarization. This is mainly done via clustering the time and frequency band dependant vectors and subsequently performing a majority vote procedure on all frequency dependent vectors of the same time segment to assign a speaker label. The result is evaluated on recordings of interviews of aphasia patients and language therapists. We demonstrate general applicability, with error rates that are close to what has been previously achieved in diarizing children’s speech. Additionally, we propose to enhance the processing pipelines with smoothing and a more sophisticated, energy based, voting scheme.
AB - Speaker attribution and labeling of single channel, multi speaker audio files is an area of active research, since the underlying problems have not been solved satisfactorily yet. This especially holds true for non-standard voices and speech, such as children and impaired speakers. Being able to perform speaker labelling of pathological speech would potentially enable the development of computer assisted diagnosis and treatment systems and is thus a desirable research goal. In this manuscript we investigate on the applicability of embeddings of audio signals, in the form of time and frequency-band based segments, into arbitrary vector spaces on diarization of pathological speech. We focus on modifying an existing embedding estimator such that it can be used for diarization. This is mainly done via clustering the time and frequency band dependant vectors and subsequently performing a majority vote procedure on all frequency dependent vectors of the same time segment to assign a speaker label. The result is evaluated on recordings of interviews of aphasia patients and language therapists. We demonstrate general applicability, with error rates that are close to what has been previously achieved in diarizing children’s speech. Additionally, we propose to enhance the processing pipelines with smoothing and a more sophisticated, energy based, voting scheme.
KW - Diarization
KW - Expressive language disorders
KW - Machine learning
KW - Medical informatics
UR - https://www.scopus.com/pages/publications/85066461217
U2 - 10.3233/978-1-61499-971-3-81
DO - 10.3233/978-1-61499-971-3-81
M3 - Conference contribution
C2 - 31118322
AN - SCOPUS:85066461217
T3 - Studies in Health Technology and Informatics
SP - 81
EP - 88
BT - dHealth 2019 - From eHealth to dHealth - Proceedings of the 13th Health Informatics Meets Digital Health Conference
A2 - Hayn, Dieter
A2 - Eggerth, Alphons
A2 - Schreier, Gunter
PB - IOS Press
T2 - 13th Health Informatics Meets Digital Health Conference: From eHealth to dHealth, dHealth 2019
Y2 - 28 May 2019 through 29 May 2019
ER -