FaceReader 10 - Voice Analysis White Paper - Introduction
Last updated: Jul 31, 2026
What Is Voice Analysis?
Voice analysis is a powerful tool for understanding human emotions, offering a new dimension to affective computing. So, how does it work?
Speech Emotion Recognition (SER) refers to the process of identifying and classifying emotional expressions based on vocal characteristics, using machine learning models trained on diverse datasets. By analyzing features such as pitch, loudness, intonation, and speech rate, SER can estimate the expressed emotional state of a speaker, independent of the actual words spoken.
SER has a wide range of research applications, including human-computer interaction, psychological research, and human factors studies. It enables the investigation of emotional responses in social interactions, cognitive load during task performance, and user experience in conversational systems. By integrating voice analysis with existing affective computing technologies, researchers can gain a more comprehensive understanding of human emotions.
Voice Analysis in FaceReader
FaceReader, which already provides facial expression analysis, eye tracking, heart rate, and breathing rate estimation, now also supports voice-based emotion recognition.
This multimodal approach allows for a more holistic assessment of affective states by combining facial and vocal expressions. With this enhancement, FaceReader can better capture subtle emotional cues, making it an even more powerful tool for emotion research and applied behavioral studies.
This white paper discusses the methodology and expected performance of FaceReader's voice analysis techniques, including the assessment of potential bias and practical guidelines for optimal results.
Source: EthoVision XT 18 - THC - Trial and Hardware Control, Noldus Information Technology