How music affects children’s development
Researchers dove deeper into building an understanding of the relationship between music and emotions and how music affects children’s development.
Read More arrow_forwardFaceReader 10 vs. OpenFace, LibreFace, and pyAFAR: compare emotion classification accuracy, Action Unit detection, and infant facial expression analysis.
In 2025, we compared FaceReader to OpenFace, at the time the leading open-source facial expression tool. Since then, there is a new version of OpenFace and two more open-source contenders have emerged, with one of them (LibreFace) releasing an updated version. Here's how FaceReader 10 stacks up against all of them today, on emotion classification, Action Unit detection, and infant facial expression analysis.
Facial expression analysis is one of several complementary ways researchers measure emotion, alongside facial EMG and manual FACS coding. Each has strengths and weaknesses. FACS coding by certified human coders is the most established gold standard, but it's slow. Researchers at the University of Bern (Stöckli, Schulte-Mecklenbeck, Borer, & Samson, 2018) put it this way:
Video recordings of participants' faces are often recorded with a resolution of 24 frames/s, meaning that for each second of recording the coder has to produce 24 ratings of the 46 AUs. So for one participant with only 1 min of video, 1,440 individual ratings are necessary. Assuming that a coder could rate one picture per second, this would add up to approximately 24 min of work for 1 min of video data.
FaceReader 10, by contrast, can process a comparable one-minute video in well under a minute in many cases, though exact speed depends on hardware and settings. Trade-offs between speed, cost, and validation are unavoidable, and automated tools of every kind, commercial and open-source, have progressed steadily in recent years.
We support the wider adoption of automated facial expression analysis generally, including open-source tools: more validated options are good for the field. But validation is also the main thing that separates these tools in practice. In our own conversations with customers, researchers occasionally come to us after a peer reviewer pushed back on an unvalidated open-source tool and asked for a FACS-validated alternative instead. That's a pattern we've seen firsthand, not a formal study, but it's a real, practical reason the validation data below matters beyond the numbers themselves.
When we last compared FaceReader to OpenFace, OpenFace v2.2 (2018) was the most current open-source option, and it was the only serious open-source alternative most researchers considered. That's no longer true. OpenFace 3 was released in 2025, and two newer tools, pyAFAR and LibreFace, have gained adoption in the research community. This update re-runs the comparison against all of them, using FaceReader 10.
Emotion classification was tested on the same two benchmark datasets as our original comparison: the Amsterdam Dynamic Facial Expression Set (ADFES) and the Warsaw Set of Emotional Facial Expression Pictures (WSEFEP).
Action Unit detection is harder to compare directly, because each tool detects a different number of AUs: FaceReader 10 outputs 20 (with intensities), LibreFace and LibreFace 2 each cover 17 (12 intensities and 5 binary), OpenFace 2.2 outputs 16, pyAFAR outputs 14, and OpenFace 3 outputs 8. To keep each comparison fair, we scored FaceReader only on the specific AU subset each competing tool supports, then compared F1 scores (the balanced measure of precision and recall) on that shared subset. That's why FaceReader's own score shifts slightly from one comparison to the next below, as it's being evaluated on a different, smaller slice of its 20 AUs each time, matched to what the competing tool can detect.
A note on the two LibreFace versions: they output graded intensity values for 12 of their 17 AUs and binary presence labels for the remaining 5. Where an AU offers both types of output, our analysis used whichever was more accurate for that AU.
Infant facial expression analysis was tested separately, on the BabyFACS manual test set (Maroulis et al., 2017; Oster, 2006), since it uses infant-specific Action Units that don't appear in the adult ADFES/WSEFEP data.
Classifying the six basic expressions plus neutral on ADFES + WSEFEP:
| Tool | Accuracy | F1 score |
|---|---|---|
| FaceReader 10 | 98% | 98% |
| LibreFace | 93% | 93% |
| LibreFace 2 | 91%* | 91%* |
| OpenFace 3 | 88% | 88% |
* LibreFace 2 was scored on five basic expressions plus neutral: at the time of testing, the tool never output the Disgust category, so Disgust was excluded from its evaluation. All other tools were scored on all seven categories.
F1 score on each tool's own supported AU subset, matched against FaceReader 10 on that same subset:
| Competing tool | AUs supported | Tool's F1 | FaceReader's F1 (same AUs) |
|---|---|---|---|
| OpenFace 3 | 8 | 81% | 88% |
| OpenFace 2.2 | 16 | 61% | 81% |
| pyAFAR | 14 | 41% | 80% |
| LibreFace | 12 intensity + 5 binary | 58% | 81% |
| LibreFace 2 | 12 intensity + 5 binary | 56% | 81% |
FaceReader 10 outperforms every tool on its own AU subset, and detects more AUs overall (20) than any of the five alternatives.
Of the five tools compared, only pyAFAR ships an infant-specific model, covering 7 Action Units. On the BabyFACS manual test set (a subset of these 7 AUs), FaceReader 10's baby model scores meaningfully higher:
| Tool | AUs (infant model) | F1 score |
|---|---|---|
| pyAFAR (infant) | 7 | 44% |
| FaceReader 10 (baby model) | 7 | 65% |
Accuracy isn't the only thing that determines how usable a tool is day-to-day:
Is OpenFace 3 better than OpenFace 2?
Not straightforwardly. OpenFace 3 is a lighter, faster rebuild, but it detects far fewer Action Units than OpenFace 2.2 (8 vs. 16), and on that smaller set, FaceReader's advantage over OpenFace 3 (81% vs. 88% F1) is narrower than its advantage over OpenFace 2.2 (61% vs. 81% F1). Whether OpenFace 3 is a net upgrade depends on which AUs a given study actually needs.
Is LibreFace 2 an upgrade over the original LibreFace?
Partly. LibreFace 2.0 retrains its models on a large synthetic dataset to improve generalizability and adds gaze estimation. On our benchmarks its overall Action Unit score came out close to the original's, with most relative improvement seen in binary prediction of AU presence (without intensities). As with OpenFace 3 versus 2.2, whether it's a net upgrade depends on what a given study needs.
Why does FaceReader's score change in each Action Unit comparison?
Because each competing tool supports a different, smaller number of AUs than FaceReader's full 20. To keep the comparison fair, FaceReader is scored only on the specific AUs each tool detects, so its score reflects a different subset each time, not an inconsistent result.
Can any open-source tool analyze infant facial expressions?
Of the tools compared here, only pyAFAR has an infant-specific model (7 Action Units, added in 2024). FaceReader 10's baby model covers the same 7 AUs and scores higher on the BabyFACS benchmark.
Do I need to write code to use these open-source tools?
For OpenFace 3, LibreFace, and LibreFace 2, yes. All three are distributed primarily as Python libraries. OpenFace 2.2 and pyAFAR include a minimal graphical interface, though still far more limited than a purpose-built application.
What datasets were used for this comparison?
Emotion classification was tested on ADFES and WSEFEP, the same datasets used in our original FaceReader-vs-OpenFace comparison. Infant Action Unit detection was tested separately on the BabyFACS manual test set.
Talk to our team or request a demo to see how FaceReader 10 handles your specific study design.
Stöckli, S., Schulte-Mecklenbeck, M., Borer, S., & Samson, A. C. (2018). Facial expression analysis with AFFDEX and FACET: A validation study. Behavior Research Methods (50), 1446–1460. https://doi.org/10.3758/s13428-017-0996-1
Baltrusaitis, T., Zadeh, A., Lim, Y. C., & Morency, L.-P. (2018). OpenFace 2.0: Facial Behavior Analysis Toolkit. IEEE FG 2018.
Hu, J., Mathur, L., Liang, P. P., & Morency, L.-P. (2025). OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis. arXiv:2506.02891.
Chang, D., Yin, Y., Li, Z., Tran, M., & Soleymani, M. (2024). LibreFace: An Open-Source Toolkit for Deep Facial Expression Analysis. WACV 2024.
Guan, X., Chaubey, A., Siniukov, M., Hsieh, A., Li, Z., & Soleymani, M. (2026). LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic Data. IEEE FG 2026.
Cohn, J. F., et al. (2023). PyAFAR: Python-based Automated Facial Action Recognition library for use in Infants and Adults. IEEE ACII 2023.
Ertugrul, I. O., Hinduja, S., Bilalpur, M., Messinger, D. S., & Cohn, J. F. (2024). Expanding PyAFAR: A Novel Privacy-Preserving Infant AU Detector. IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG) 2024. https://doi.org/10.1109/FG59268.2024.10581868
Van der Schalk, J., Hawk, S., Fischer, A., & Doosje, B. (2011). Moving faces, looking places: Validation of the Amsterdam Dynamic Facial Expression Set (ADFES). Emotion (11), 907–920.
Olszanowski, M., Pochwatko, G., Kuklinski, K., Scibor-Rylski, M., Lewinski, P., & Ohme, R. (2014). Warsaw set of emotional facial expression pictures: a validation study of facial display photographs. Frontiers in Psychology (5).
Oster, H. (2006). Baby FACS: Facial Action Coding System for Infants and Young Children.
Maroulis, A., Spink, A. J., Theuws, J. J. M., Oster, H., & Buitelaar, J. (2017). Sweet or sour. Validating Baby FaceReader to analyse infant responses to food. Poster presentation, 12th Pangborn Sensory Science Symposium, 20–24 August 2017.
Researchers dove deeper into building an understanding of the relationship between music and emotions and how music affects children’s development.
Read More arrow_forward
Think about some of your favorite holiday foods – what are they? Maybe gingerbread, candy canes, or pies?
Read More arrow_forward
As a researcher, one of my biggest thrills was being able to predict how someone was going to behave, especially without asking him or her.
Read More arrow_forwardWe'll get back to you shortly.
Please correct the following errors:
We look forward to seeing you at the NoldusViso launch event.
Please correct the following errors: