FaceReader 10 vs. Open-Source Facial Expression Tools

FaceReader 10 vs. OpenFace, LibreFace, and pyAFAR: compare emotion classification accuracy, Action Unit detection, and infant facial expression analysis.

calendar_today Tue 29 Sep. 2026
folder Emotion
FaceReader 10 vs. Open-Source Facial Expression Tools

In 2025, we compared FaceReader to OpenFace, at the time the leading open-source facial expression tool. Since then, there is a new version of OpenFace and two more open-source contenders have emerged, with one of them (LibreFace) releasing an updated version. Here's how FaceReader 10 stacks up against all of them today, on emotion classification, Action Unit detection, and infant facial expression analysis.

Bar chart comparing FaceReader 10 to LibreFace, LibreFace 2, and OpenFace 3 on emotion classification accuracy

1. Why this comparison, and why now

Facial expression analysis is one of several complementary ways researchers measure emotion, alongside facial EMG and manual FACS coding. Each has strengths and weaknesses. FACS coding by certified human coders is the most established gold standard, but it's slow. Researchers at the University of Bern (Stöckli, Schulte-Mecklenbeck, Borer, & Samson, 2018) put it this way:

Video recordings of participants' faces are often recorded with a resolution of 24 frames/s, meaning that for each second of recording the coder has to produce 24 ratings of the 46 AUs. So for one participant with only 1 min of video, 1,440 individual ratings are necessary. Assuming that a coder could rate one picture per second, this would add up to approximately 24 min of work for 1 min of video data.

FaceReader 10, by contrast, can process a comparable one-minute video in well under a minute in many cases, though exact speed depends on hardware and settings. Trade-offs between speed, cost, and validation are unavoidable, and automated tools of every kind, commercial and open-source, have progressed steadily in recent years.

We support the wider adoption of automated facial expression analysis generally, including open-source tools: more validated options are good for the field. But validation is also the main thing that separates these tools in practice. In our own conversations with customers, researchers occasionally come to us after a peer reviewer pushed back on an unvalidated open-source tool and asked for a FACS-validated alternative instead. That's a pattern we've seen firsthand, not a formal study, but it's a real, practical reason the validation data below matters beyond the numbers themselves.

When we last compared FaceReader to OpenFace, OpenFace v2.2 (2018) was the most current open-source option, and it was the only serious open-source alternative most researchers considered. That's no longer true. OpenFace 3 was released in 2025, and two newer tools, pyAFAR and LibreFace, have gained adoption in the research community. This update re-runs the comparison against all of them, using FaceReader 10.

2. Meet the contenders

  • OpenFace 2.2 (2018): the toolkit from our original comparison, included again here as the aging baseline. Baltrusaitis, T., Zadeh, A., Lim, Y. C., & Morency, L.-P. (2018). OpenFace 2.0: Facial Behavior Analysis Toolkit. IEEE FG 2018.
  • OpenFace 3.0 (2025): a full rebuild from CMU's MultiComp Lab, a lightweight, multitask model for facial landmarks, Action Units, gaze, and emotion recognition. Hu, J., Mathur, L., Liang, P. P., & Morency, L.-P. (2025). OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis. FG 2025 / arXiv:2506.02891.
  • LibreFace (2024): an open-source deep-learning toolkit for AU detection, AU intensity, and expression recognition. Chang, D., Yin, Y., Li, Z., Tran, M., & Soleymani, M. (2024). LibreFace: An Open-Source Toolkit for Deep Facial Expression Analysis. WACV 2024.
  • LibreFace 2.0 (2026): an update of LibreFace that retrains its models with a large synthetic dataset to improve generalizability, and adds gaze estimation. Guan, X., Chaubey, A., Siniukov, M., Hsieh, A., Li, Z., & Soleymani, M. (2026). LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic Data. IEEE FG 2026.
  • pyAFAR (2023, infant extension 2024): a Python library for AU detection in adults and, since 2024, infants. Cohn, J. F., et al. (2023). PyAFAR: Python-based Automated Facial Action Recognition library for use in Infants and Adults. IEEE ACII 2023; infant extension: Ertugrul, I. O., Hinduja, S., Bilalpur, M., Messinger, D. S., & Cohn, J. F. (2024). Expanding PyAFAR: A Novel Privacy-Preserving Infant AU Detector. IEEE FG 2024.

3. How we tested

Emotion classification was tested on the same two benchmark datasets as our original comparison: the Amsterdam Dynamic Facial Expression Set (ADFES) and the Warsaw Set of Emotional Facial Expression Pictures (WSEFEP).

Action Unit detection is harder to compare directly, because each tool detects a different number of AUs: FaceReader 10 outputs 20 (with intensities), LibreFace and LibreFace 2 each cover 17 (12 intensities and 5 binary), OpenFace 2.2 outputs 16, pyAFAR outputs 14, and OpenFace 3 outputs 8. To keep each comparison fair, we scored FaceReader only on the specific AU subset each competing tool supports, then compared F1 scores (the balanced measure of precision and recall) on that shared subset. That's why FaceReader's own score shifts slightly from one comparison to the next below, as it's being evaluated on a different, smaller slice of its 20 AUs each time, matched to what the competing tool can detect.

A note on the two LibreFace versions: they output graded intensity values for 12 of their 17 AUs and binary presence labels for the remaining 5. Where an AU offers both types of output, our analysis used whichever was more accurate for that AU.

Infant facial expression analysis was tested separately, on the BabyFACS manual test set (Maroulis et al., 2017; Oster, 2006), since it uses infant-specific Action Units that don't appear in the adult ADFES/WSEFEP data.

4. Basic emotion classification accuracy

Classifying the six basic expressions plus neutral on ADFES + WSEFEP:

ToolAccuracyF1 score
FaceReader 1098%98%
LibreFace93%93%
LibreFace 291%*91%*
OpenFace 388%88%

* LibreFace 2 was scored on five basic expressions plus neutral: at the time of testing, the tool never output the Disgust category, so Disgust was excluded from its evaluation. All other tools were scored on all seven categories.

5. Action Unit detection

F1 score on each tool's own supported AU subset, matched against FaceReader 10 on that same subset:

Competing toolAUs supportedTool's F1FaceReader's F1 (same AUs)
OpenFace 3881%88%
OpenFace 2.21661%81%
pyAFAR1441%80%
LibreFace12 intensity + 5 binary58%81%
LibreFace 212 intensity + 5 binary56%81%

FaceReader 10 outperforms every tool on its own AU subset, and detects more AUs overall (20) than any of the five alternatives.

6. Infant facial expression analysis

Of the five tools compared, only pyAFAR ships an infant-specific model, covering 7 Action Units. On the BabyFACS manual test set (a subset of these 7 AUs), FaceReader 10's baby model scores meaningfully higher:

ToolAUs (infant model)F1 score
pyAFAR (infant)744%
FaceReader 10 (baby model)765%

7. Beyond the numbers: usability and support

Accuracy isn't the only thing that determines how usable a tool is day-to-day:

  • Interface: OpenFace 2.2 and pyAFAR ship with a minimal graphical interface. OpenFace 3 and LibreFace are primarily Python libraries, so using either means writing and maintaining your own code. FaceReader is a complete application with a GUI, visualization, and reporting built in, no coding required.
  • Update cadence: open-source tools have released major versions at very irregular intervals. OpenFace, for example, went seven years between 2.2 (2018) and 3.0 (2025). FaceReader, by contrast, ships a new version roughly every one to two years.
  • Support: FaceReader includes direct vendor support and training; open-source tools rely on community forums and, for the academic projects here, the availability of the original research groups.

8. FAQ

Is OpenFace 3 better than OpenFace 2?

Not straightforwardly. OpenFace 3 is a lighter, faster rebuild, but it detects far fewer Action Units than OpenFace 2.2 (8 vs. 16), and on that smaller set, FaceReader's advantage over OpenFace 3 (81% vs. 88% F1) is narrower than its advantage over OpenFace 2.2 (61% vs. 81% F1). Whether OpenFace 3 is a net upgrade depends on which AUs a given study actually needs.

Is LibreFace 2 an upgrade over the original LibreFace?

Partly. LibreFace 2.0 retrains its models on a large synthetic dataset to improve generalizability and adds gaze estimation. On our benchmarks its overall Action Unit score came out close to the original's, with most relative improvement seen in binary prediction of AU presence (without intensities). As with OpenFace 3 versus 2.2, whether it's a net upgrade depends on what a given study needs.

Why does FaceReader's score change in each Action Unit comparison?

Because each competing tool supports a different, smaller number of AUs than FaceReader's full 20. To keep the comparison fair, FaceReader is scored only on the specific AUs each tool detects, so its score reflects a different subset each time, not an inconsistent result.

Can any open-source tool analyze infant facial expressions?

Of the tools compared here, only pyAFAR has an infant-specific model (7 Action Units, added in 2024). FaceReader 10's baby model covers the same 7 AUs and scores higher on the BabyFACS benchmark.

Do I need to write code to use these open-source tools?

For OpenFace 3, LibreFace, and LibreFace 2, yes. All three are distributed primarily as Python libraries. OpenFace 2.2 and pyAFAR include a minimal graphical interface, though still far more limited than a purpose-built application.

What datasets were used for this comparison?

Emotion classification was tested on ADFES and WSEFEP, the same datasets used in our original FaceReader-vs-OpenFace comparison. Infant Action Unit detection was tested separately on the BabyFACS manual test set.

Curious how FaceReader fits your research?

Talk to our team or request a demo to see how FaceReader 10 handles your specific study design.

References

Stöckli, S., Schulte-Mecklenbeck, M., Borer, S., & Samson, A. C. (2018). Facial expression analysis with AFFDEX and FACET: A validation study. Behavior Research Methods (50), 1446–1460. https://doi.org/10.3758/s13428-017-0996-1

Baltrusaitis, T., Zadeh, A., Lim, Y. C., & Morency, L.-P. (2018). OpenFace 2.0: Facial Behavior Analysis Toolkit. IEEE FG 2018.

Hu, J., Mathur, L., Liang, P. P., & Morency, L.-P. (2025). OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis. arXiv:2506.02891.

Chang, D., Yin, Y., Li, Z., Tran, M., & Soleymani, M. (2024). LibreFace: An Open-Source Toolkit for Deep Facial Expression Analysis. WACV 2024.

Guan, X., Chaubey, A., Siniukov, M., Hsieh, A., Li, Z., & Soleymani, M. (2026). LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic Data. IEEE FG 2026.

Cohn, J. F., et al. (2023). PyAFAR: Python-based Automated Facial Action Recognition library for use in Infants and Adults. IEEE ACII 2023.

Ertugrul, I. O., Hinduja, S., Bilalpur, M., Messinger, D. S., & Cohn, J. F. (2024). Expanding PyAFAR: A Novel Privacy-Preserving Infant AU Detector. IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG) 2024. https://doi.org/10.1109/FG59268.2024.10581868

Van der Schalk, J., Hawk, S., Fischer, A., & Doosje, B. (2011). Moving faces, looking places: Validation of the Amsterdam Dynamic Facial Expression Set (ADFES). Emotion (11), 907–920.

Olszanowski, M., Pochwatko, G., Kuklinski, K., Scibor-Rylski, M., Lewinski, P., & Ohme, R. (2014). Warsaw set of emotional facial expression pictures: a validation study of facial display photographs. Frontiers in Psychology (5).

Oster, H. (2006). Baby FACS: Facial Action Coding System for Infants and Young Children.

Maroulis, A., Spink, A. J., Theuws, J. J. M., Oster, H., & Buitelaar, J. (2017). Sweet or sour. Validating Baby FaceReader to analyse infant responses to food. Poster presentation, 12th Pangborn Sensory Science Symposium, 20–24 August 2017.

Related Posts

shopping_bag
check_circle

Thank you!

We'll get back to you shortly.

error

Please correct the following errors:

error

error

error

error

By clicking Submit, you consent to Noldus processing your data as described in our privacy policy.

check_circle

Thank you for registering!

We look forward to seeing you at the NoldusViso launch event.

error

Please correct the following errors:

error

error

error

error

error

By clicking Register, you consent to Noldus processing your data as described in our privacy policy.