Facial attractiveness datasets: what AI beauty scores are actually trained on, 14 datasets compared
SCUT-FBP5500 averages 60 young volunteers per face; compare 14 facial attractiveness datasets, labels, rater pools and licenses.

You are writing about an AI beauty score, or staring at one after an app has given you a number. The useful question is not simply “how accurate is it?” It is: whose opinion was the model trained to forecast?
The clearest example is SCUT-FBP5500: every face is summarized from ratings by 60 volunteers aged 18–27, on a 1–5 scale. CelebA is much larger, with 202,599 images, but its “Attractive” field is only one yes/no attribute. Those are very different kinds of evidence.
This page compares 14 datasets and their labels, rater pools and permissions. A beauty score forecasts one crowd’s opinion, not a measurement.
Key numbers
- 60 volunteers aged 18–27 rated every SCUT-FBP5500 face on a 1–5 scale. Source
- 202,599 images make up CelebA, whose yes/no “Attractive” field is one of 40 binary attributes. Source
- Fleiss’ kappa 0.514 was reported for CelebA’s Attractive label, compared with 0.9789 for Male in one audit. Source
- 27.9% versus 67.9% of CelebA images were labeled Attractive among images labeled Male versus images not labeled Male. Source
- 2,513 raters aged 17–90 rated neutral front faces in the openly licensed London Set. Source
- 213 OpenAlex citations made SCUT-FBP5500 the most-cited purpose-built beauty-prediction dataset paper in this ledger on 10 September 2026. Source
Table of contents
- What datasets are AI attractiveness scores trained on?
- Who rated the faces, and why does that matter more than how many?
- Why do beauty datasets cover such a narrow slice of faces?
- Is CelebA’s Attractive label a real beauty score?
- Can a commercial face-rating app use these datasets?
- Which face datasets are the most cited, and does that make them beauty AI?
- What should you ask before trusting an AI attractiveness score?
- How to cite this page
- The bottom line
- Sources
What datasets are AI attractiveness scores trained on?
AI attractiveness scores can be trained on purpose-built beauty datasets, general face-attribute collections, or psychology stimulus sets with attractiveness norms. The decisive difference is not the dataset’s fame; it is the label definition, rater pool and permission attached to the images.
| Dataset (year) | Faces | Who rated | Terms of use |
|---|---|---|---|
| SCUT-FBP5500 (2018) | 5,500 | 60 volunteers aged 18–27, 1–5 | Non-commercial research |
| SCUT-FBP (2015) | 500 Asian women | 75-rater pool, about 70 per image | Not verified |
| MEBeauty (2021) | 2,550 | About 300 raters | Non-commercial research |
| LSAFBD (2020) | 20,000 | 200 volunteers, mostly 20–35 | Not verified |
| CelebA (2015) | 202,599 images | Professional labeling company, yes/no | Non-commercial research only |
| Chicago Face Database (2015; v3.0) | 158 → 597 people | 1,087 raters (2015), 1–7 | Non-commercial scientific research |
| CFD-MR (2020) | 88 | 499 MTurk raters, 1–7 | CFD terms |
| CFD-INDIA (2021) | 142 | U.S. and Indian rater samples | CFD terms |
| 10k US Adult Faces (2013) | 10,168; 2,222 rated | 15 AMT raters per survey version, 9-point | Research only |
| London Set (2021) | 102 | 2,513 raters aged 17–90, 1–7 | CC BY 4.0 |
| AMFD (2020) | 110 | 2,123 raters, about 50 per face | Academic researchers; exact terms not verified |
| RaFD (2010) | 67 models | Not verified | Accredited universities, non-commercial |
| Oslo Face Database | About 200 | Not verified | Research use on request |
| HotOrNot (2010) | 2,056 faces in one secondary account | Accounts conflict | Not verified |
SCUT-FBP5500 directly targets beauty; CelebA is larger but general-purpose, with Attractive as one binary field among 40. SCUT-FBP5500, CelebA
Caveat: a table can show provenance and permissions, not whether any particular commercial app used a dataset. We have no evidence for what any named app trained on.
Who rated the faces, and why does that matter more than how many?
Sixty ratings per face is not automatically a weak sample. The AMFD authors cite earlier work saying face-based ratings stabilize after about 45 independent observations; the sharper limitation is often who those observers were and what population the model is expected to represent.
The pools differ: SCUT-FBP5500 used 60 volunteers aged 18–27; LSAFBD, 200 mostly aged 20–35; the London Set, 2,513 aged 17–90; CFD, 1,087; and AMFD, 2,123, about 50 per face. AMFD, Chicago Face Database, London Set
A ground-truth score is usually an average, compressing a crowd into one target. That helps training but hides disagreement, age and cultural context, and the cues different raters notice.

Caveat: bigger rater pools can stabilize an average without making that average universal. Reliability and representativeness are separate questions.
Why do beauty datasets cover such a narrow slice of faces?
Beauty datasets are narrow for practical and methodological reasons: some were designed around a single demographic, some deliberately selected attractive faces, and some inherited the bias of online image search. A model can learn a clean pattern from a narrow collection and still meet the wrong real-world population.
SCUT-FBP contains 500 Asian female subjects by design. SCUT-FBP5500 contains 2,000 Asian females, 2,000 Asian males, 750 Caucasian females and 750 Caucasian males; its faces are frontal, unoccluded and neutral-expression, aged 15–60. It also deliberately contains a higher proportion of beautiful faces to make beauty easier to learn. SCUT-FBP, SCUT-FBP5500 release
LSAFBD contains 20,000 images of Asian women collected from the web. The 10k US Adult Faces database is broader in raw size, but its 10,168 photographs came from Google Image Search using random first-and-last-name pairings sampled from the 1990 U.S. Census name distribution. Of those people, 57.1% were male and 83.7% White.
Are face-rating apps Eurocentric? cannot be answered from model output alone; inspect the faces and labels behind it.
Caveat: narrow datasets are not useless. They can be valuable for controlled experiments; the mistake is presenting their forecast as a universal human judgment.
Is CelebA’s Attractive label a real beauty score?
CelebA’s Attractive field is a real dataset label, but it is not a continuous beauty score. Each of the 202,599 celebrity images carries 40 binary attributes, and Attractive is one yes/no annotation made by a professional labeling company, with the labeler count, demographics and instructions undisclosed.
An audit reported Fleiss’ kappa of 0.5140 for Attractive, compared with 0.9789 for Male. Fleiss’ kappa measures agreement beyond chance; 1 means perfect agreement. It also found at least one third of CelebA images had one or more incorrect attribute labels, though that covers all 40 attributes rather than Attractive specifically. CelebA paper, label audit
The audit also found a gender skew: 27.9% of images labeled Male were also labeled Attractive, compared with 67.9% of images not labeled Male. The authors say labeling bias and data-selection bias are difficult to separate.
So a binary annotation for general face attributes should not be treated as a universal beauty instrument.
Caveat: kappa in this audit was computed by treating different images of the same identity as reviewers, so it is not ordinary inter-annotator agreement between two people labeling the same image.
Can a commercial face-rating app use these datasets?
Usually, the answer is no unless the license clearly permits it. In the sources checked here, SCUT-FBP5500, MEBeauty, CelebA, CFD, RaFD and the 10k US Adult Faces are restricted to research or non-commercial use; the London Set is released under CC BY 4.0.
Licensing also goes beyond “can I download it?” CelebA’s agreement bars commercial exploitation of any portion of its images or derived data. CFD terms forbid redistribution, including to file-hosting services and cloud-based AI image-manipulation services; its April 2024 update says uploading CFD materials to such services violates them. CelebA agreement, CFD terms, CFD version history
RaFD is stricter: non-commercial scientific use is for researchers at an officially accredited university, and its site said downloads were unavailable on 10 September 2026. The London Set is the contrast: CC BY 4.0 permits commercial reuse with attribution.
Do not infer dataset use from marketing. This ledger has no evidence of what any named app trained on or whether it complied with dataset terms. Do face-rating apps work? is the right follow-up for judging an output.

Caveat: licensing is a legal and permission question, not a quality score. A restricted dataset can be scientifically useful; an open dataset can still carry sampling and label problems.
Which face datasets are the most cited, and does that make them beauty AI?
Citation count measures how often a paper is cited, not whether it is a good beauty predictor. CelebA is the most cited paper in this list, while SCUT-FBP5500 is the most cited purpose-built facial beauty prediction paper; many highly cited datasets are used for emotion, social psychology or memorability instead.
| Rank | Dataset | OpenAlex citations | Main use |
|---|---|---|---|
| 1 | CelebA (2015) | 7,815 | General face attributes |
| 2 | RaFD (2010) | 2,495 | Emotional expressions |
| 3 | Chicago Face Database (2015) | 1,884 | Social-psychology stimuli |
| 4 | 10k US Adult Faces (2013) | 440 | Memorability |
| 5 | SCUT-FBP5500 (2018) | 213 | Beauty prediction |
| 6 | SCUT-FBP (2015) | 139 | Beauty prediction |
| 7 | HotOrNot (2010) | 127 | Beauty prediction |
| 8 | CFD-MR (2020) | 85 | Multiracial face perception |
| 9 | AMFD (2020) | 72 | Multiracial face perception |
| 10 | CFD-INDIA (2021) | 54 | Cross-cultural impressions |
| 11 | London Set | 41 | Face research |
| 12 | LSAFBD (2020) | 33 | Beauty prediction |
| 13 | MEBeauty (2021) | 28 | Beauty prediction |
These OpenAlex counts were retrieved on 10 September 2026 and drift over time. They measure broad use: CelebA’s 7,815 citations are not 7,815 attractiveness studies, and the London Set count may undercount inconsistent citations of other versions. OpenAlex SCUT-FBP5500 record, OpenAlex CelebA record
Caveat: “most cited” is a discoverability signal, not evidence that a dataset captures real-world attractiveness well.
What should you ask before trusting an AI attractiveness score?
Before trusting a face score, audit the target rather than admiring the decimal. Five questions usually reveal more than the model’s claimed accuracy.
- Whose ratings produced the target? SCUT-FBP5500 used 60 volunteers aged 18–27; the London Set used 2,513 raters aged 17–90. Those forecasts need not agree.
- Which faces were included? SCUT-FBP was 500 Asian female subjects; SCUT-FBP5500 used frontal, unoccluded, neutral-expression faces and over-represented beautiful faces.
- What kind of label is it? CelebA’s Attractive field is binary. SCUT-FBP5500 averages a 1–5 rating; CFD norms use a 1–7 Likert scale; the 10k database uses a 9-point scale.
- What does the score mean outside the dataset? A model may predict a crowd average, not dating outcomes, chemistry or a particular reaction. Why AI can’t measure attractiveness explains the missing context.
- Can the data legally support the product? Research-only terms are not commercial training permission. Check the agreement, especially for derived data.
This is also where looksmaxxing and appearance anxiety can spiral: a crowd forecast becomes a verdict on your face. Ask instead what first impression a photo gives, what is changeable, and what the data cannot tell you. Real World Appeal’s free photo read uses that frame: it describes a photo’s first impression, shows results before any payment decision, and is not a clinical or scientific instrument.
Caveat: even a transparent dataset audit cannot predict every real-life interaction. Human responses remain contextual and individual.
How to cite this page
Suggested citation: Real World Appeal. “Facial attractiveness datasets: what AI beauty scores are actually trained on, 14 datasets compared.” realworldappeal.com, September 2026. https://realworldappeal.com/blog/facial-attractiveness-datasets
The sources below are linked so journalists, developers and students can cite the primary papers and official dataset pages directly. Dataset details and OpenAlex citation counts were retrieved on 10 September 2026; citation counts drift.
The bottom line
The number everyone wants is “how accurate is the AI?” The better question comes first: whose opinion is it predicting, from which faces, under which label definition and license? SCUT-FBP5500’s 60 young volunteers are not a trivial sample, but they are one crowd; CelebA’s 202,599 images are not a beauty scale, but a general attribute collection with a yes/no Attractive field.
A beauty score is a forecast of one crowd’s opinion. It can be useful as a narrow benchmark. It cannot be promoted into a verdict on a person.
For a practical next step, try Real World Appeal’s free photo read. It is not a clinical or scientific instrument or a universal measurement of you.
Sources
- Xie, Duorui, Lingyu Liang, Lianwen Jin, Jie Xu and Mengru Li (2015). “SCUT-FBP: A Benchmark Dataset for Facial Beauty Perception.” 2015 IEEE International Conference on Systems, Man, and Cybernetics, pp. 1821–1826. https://doi.org/10.1109/SMC.2015.319
- Xie, Duorui, Lingyu Liang, Lianwen Jin, Jie Xu and Mengru Li (2015). “SCUT-FBP: A Benchmark Dataset for Facial Beauty Perception.” arXiv:1511.02459. https://arxiv.org/abs/1511.02459
- Liang, Lingyu, Luojun Lin, Lianwen Jin, Duorui Xie and Mengru Li (2018). “SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty Prediction.” 2018 24th International Conference on Pattern Recognition, pp. 1598–1603. https://doi.org/10.1109/ICPR.2018.8546038
- HCII Lab, South China University of Technology (2018). “SCUT-FBP5500-Database-Release README.” https://github.com/HCIILAB/SCUT-FBP5500-Database-Release
- Lebedeva, Irina, Yi Guo and Fangli Ying (2021). “MEBeauty: a multi-ethnic facial beauty dataset in-the-wild.” Neural Computing and Applications 34(17), 14169–14183. https://doi.org/10.1007/s00521-021-06535-0
- fbplab (2022). “MEBeauty-database README.” https://github.com/fbplab/MEBeauty-database
- Zhai, Yikui et al. (2020). “Asian Female Facial Beauty Prediction Using Deep Neural Networks via Transfer Learning and Multi-Channel Feature Fusion.” IEEE Access 8, 56892–56907. https://doi.org/10.1109/ACCESS.2020.2980248
- Ma, Debbie S., Joshua Correll and Bernd Wittenbrink (2015). “The Chicago face database: A free stimulus set of faces and norming data.” Behavior Research Methods 47(4), 1122–1135. https://doi.org/10.3758/s13428-014-0532-5
- University of Chicago, Center for Decision Research (2026). “Chicago Face Database.” https://www.chicagofaces.org/
- University of Chicago, Center for Decision Research (2026). “Chicago Face Database Terms of Use.” https://www.chicagofaces.org/download/
- University of Chicago, Center for Decision Research (2024). “Chicago Face Database Version History.” https://www.chicagofaces.org/#version
- Ma, Debbie S., Justin Kantner and Bernd Wittenbrink (2020). “Chicago Face Database: Multiracial expansion.” Behavior Research Methods 53(3), 1289–1300. https://doi.org/10.3758/s13428-020-01482-5
- Ma, Debbie S., Justin Kantner and Bernd Wittenbrink (2020). “Chicago Face Database: Multiracial expansion.” Europe PMC full text. https://europepmc.org/article/PMC/PMC8219557
- Lakshmi, Anjana, Bernd Wittenbrink, Joshua Correll and Debbie S. Ma (2021). “The India Face Set: International and Cultural Boundaries Impact Face Impressions and Perceptions of Category Membership.” Frontiers in Psychology 12. https://doi.org/10.3389/fpsyg.2021.627678
- Bainbridge, Wilma A., Phillip Isola and Aude Oliva (2013). “The intrinsic memorability of face photographs.” Journal of Experimental Psychology: General 142(4), 1323–1334. https://doi.org/10.1037/a0033872
- Bainbridge, Wilma (2026). “10k US Adult Faces Database.” https://www.wilmabainbridge.com/facememorability2.html
- Liu, Ziwei, Ping Luo, Xiaogang Wang and Xiaoou Tang (2015). “Deep Learning Face Attributes in the Wild.” 2015 IEEE International Conference on Computer Vision, pp. 3730–3738. https://doi.org/10.1109/ICCV.2015.425
- MMLAB, The Chinese University of Hong Kong (2026). “Large-scale CelebFaces Attributes (CelebA) Dataset.” https://mmlab.ie.cuhk.edu.hk/projects/CelebA.html
- Lingenfelter, Bryson, Sara R. Davis and Emily M. Hand (2022). “A Quantitative Analysis of Labeling Issues in the CelebA Dataset.” Advances in Visual Computing, LNCS, pp. 129–141. https://doi.org/10.1007/978-3-031-20713-6_10
- DeBruine, Lisa and Benedict Jones (2021). “Face Research Lab London Set.” figshare, version 5. https://doi.org/10.6084/m9.figshare.5047666.v5
- Chen, Jacqueline M., Jasmine B. Norman and Yeseul Nam (2020). “Broadening the stimulus set: Introducing the American Multiracial Faces Database.” Behavior Research Methods 53(1), 371–389. https://doi.org/10.3758/s13428-020-01447-8
- Chen, Jacqueline M., Jasmine B. Norman and Yeseul Nam (2020). “American Multiracial Faces Database.” PsyArXiv preprint. https://doi.org/10.31234/osf.io/s4hke
- Langner, Oliver et al. (2010). “Presentation and validation of the Radboud Faces Database.” Cognition and Emotion 24(8), 1377–1388. https://doi.org/10.1080/02699930903485076
- Behavioural Science Institute, Radboud University Nijmegen (2026). “Radboud Faces Database.” https://rafd.nl/
- Affective Brain Lab, University of Oslo (2026). “Oslo Face Database.” https://affectivebrains.com/oslo-face-database/
- Gray, Douglas, Kai Yu, Wei Xu and Yihong Gong (2010). “Predicting Facial Beauty without Landmarks.” ECCV 2010, Lecture Notes in Computer Science, pp. 434–447. https://doi.org/10.1007/978-3-642-15567-3_32
- Xu, Lu, Jinhai Xiang and Xiaohui Yuan (2018). “Transferring Rich Deep Features for Facial Beauty Prediction.” arXiv:1803.07253. https://arxiv.org/abs/1803.07253
- Kalra, Peterson et al. (2019). “Photofeeler-D3: A Neural Network with Voter Modeling for Dating Photo Impression Prediction.” arXiv:1904.07435. https://arxiv.org/abs/1904.07435
- OpenAlex (2026). “Work records and citation counts retrieved 10 September 2026.” https://api.openalex.org/works/doi:10.1109/icpr.2018.8546038, https://api.openalex.org/works/doi:10.1109/iccv.2015.425
Frequently asked questions
What is the most-cited facial attractiveness dataset?
SCUT-FBP5500 is the most-cited purpose-built beauty-prediction dataset paper in this ledger, with 213 OpenAlex citations retrieved on 10 September 2026. Its scores are averages from 60 volunteers aged 18–27. See our AI face analysis glossary for the label vocabulary.
What is CelebA's Attractive label?
CelebA contains 202,599 images with 40 binary attributes, one of which is Attractive. That means the label is yes/no, not a continuous beauty scale, and it was annotated by a professional labeling company according to the paper.
Can a commercial app use a facial attractiveness dataset?
Often not. SCUT-FBP5500, MEBeauty, CelebA, CFD, RaFD and the 10k US Adult Faces are restricted to research or non-commercial use in the sources checked here; the London Set is CC BY 4.0. Read whether face-rating apps work before treating a public dataset as a product license.
Why do face-rating apps give different scores?
They may be forecasting different rater pools, labels, face collections or scales. A model trained on one crowd predicts that crowd's average opinion, not a universal property of a face. Our guide to why face-rating apps give different scores breaks down the mechanism.
Is an AI attractiveness score a verdict on my face?
No. It is, at best, a forecast of one labeled crowd's opinion and the dataset may have narrow faces, labels or licensing limits. Real World Appeal's free photo read describes the first impression a photo gives, with results visible before any payment decision; it is not a clinical or scientific instrument.
