menu_book

Knowledge Base

Documentation, guides, and resources for Noldus products.

FaceReader 10 - Voice Analysis White Paper - Applications

Last updated: Jul 31, 2026

Performance of Voice Analysis

Table 1 summarizes the performance of FaceReader's voice analysis on a diverse English-language test set. This test set was constructed using both publicly available and proprietary datasets. It was also explicitly excluded from training to ensure unbiased evaluation.

Table 1. Overall Performance Metrics on the English-Language Test Set.

  • Accuracy – 87.1: Percentage of predictions that were correct across all emotion classes.
  • Precision – 87.2: Measures how many of the predicted emotions were correct (focuses on minimizing false positives).
  • Recall – 87.1: Measures how many of the emotions were correctly identified (focuses on minimizing false negatives).
  • F1 Score – 87.1: Harmonic mean of precision and recall; a balanced measure of model accuracy.
  • UAR – 87.3: Unweighted Average Recall: average recall per class, giving equal weight to each emotion regardless of how often it occurs.

The test set includes a broad spectrum of recordings, covering both acted and spontaneous expressions of emotion. The dataset features over 4,000 recordings from speakers of varying gender, ethnicity, and dialects to ensure robustness across demographic variation.

Emotion labels were either predefined (in the case of scripted, acted speech) or obtained through crowd-sourced or expert annotation (for both spontaneous and acted speech). Inter-annotator agreement was used to assess and improve labeling reliability.

Table 2 breaks down performance by emotion category, highlighting consistently strong and balanced results across all four target emotions.

Table 2. Classification Performance Per Emotion Category on the Test Set.

  • Neutral: Precision 90.2, Recall 86.7, F1 Score 88.4
  • Happy: Precision 83.2, Recall 84.8, F1 Score 84.0
  • Sad: Precision 84.4, Recall 91.9, F1 Score 88.0
  • Angry: Precision 90.7, Recall 85.9, F1 Score 88.2

Voice Analysis

To minimize gender bias in FaceReader's voice analysis, deliberate steps are taken throughout data collection, training, and evaluation. Datasets are carefully balanced to include a representative mix of male and female voices across a range of speaking styles and dialects.

Model performance is also evaluated separately by gender to identify and address any disparities. The current model shows a performance difference of 2.7% in accuracy and 3.9% in unweighted average recall (UAR), with slightly better results for female speakers.

The training data primarily consists of English-speaking adult voices; voices from children or elderly individuals are currently underrepresented. Performance will be lower for other languages, especially those that are linguistically distant from English.

The approach is continuously monitored and refined to support fairness and inclusivity, while improving generalization across diverse linguistic contexts.

Optimizing Voice Analysis for Your Research

The following recommendations can help you get high-quality input and reliable results when using FaceReader's voice analysis in your research.

Setup

  • Use a high-quality microphone in a quiet environment: voice analysis only works on speech and background noise can negatively affect results.
  • Ensure the recording level is appropriate: neither too low (voice won't be detected) nor too high (to prevent clipping or distortion).
  • Avoid overlapping speech: voice analysis works best when one person speaks at a time.

Research

  • Try voice analysis on sample data to make sure your setup is clean and the results are stable.
  • Results are relative, so compare values across time or across speakers rather than relying on a single absolute number.
  • Use in combination with facial expressions and the measurement of vital signs for more complete emotional insights.

References

  1. Reynolds, D. (2009). Gaussian Mixture Models. In: Li, S.Z., Jain, A. (eds) Encyclopedia of Biometrics. Springer, Boston, MA.
  2. Arnfield, S., Roach, P., Setter, J., Greasley, P., Horton, D. (1995). Emotional stress and speech tempo variation. Proc. ESCA/NATO Workshop on Speech under Stress, 13–15.
  3. Braun, A. & Oba, R. (2007). Speaking tempo in emotional speech – a cross-cultural study using dubbed speech.

Source: EthoVision XT 18 - THC - Trial and Hardware Control, Noldus Information Technology

Looking for a quote?

Tell us about your lab and we will send you a tailored pricing proposal within 24 hours.

Noldus is here to assist you throughout the whole process.

shopping_bag
check_circle

Thank you!

We'll get back to you shortly.

error

Please correct the following errors:

error

error

error

error

By clicking Submit, you consent to Noldus processing your data as described in our privacy policy.