Attractiveness research methods glossary: how to read a study from inter-rater reliability to p-hacking
A plain-English attractiveness research methods glossary: 27 terms, from r = 0.99 crowd agreement to stated versus revealed preference.

You have seen the post: “science says women prefer X.” Before you believe it, you need the words to read the study behind it. This glossary gives you 27 of them, each with a plain definition, why it matters for attractiveness research, and a real example.
The short answer to the biggest question: crowd averages can agree almost perfectly, while two random people agree on only 48% of face preferences. Agreement belongs to crowds, not to people.
Key numbers
- r = 0.99 — men’s and women’s average ratings agreed for 102 female faces. Source
- 48% — two randomly selected participants agreed on face preferences. Source
- 22%, 0%, 78% — genetic, shared-family, and unique-environment or measurement-error components in individual face preferences. Source
- r = 0.83 — photo and video attractiveness ratings correlated for 60 male faces. Source
- 36% versus 97% — replication and original significance rates across 100 general psychology studies. Source
- d = 0.60 versus 0.15 — median effect size in Many Labs 2 originals and replications. Source
- 5–10% — the plainness penalty reported in labor-market research. Source
Table of contents
- How do researchers measure attractiveness?
- What do participants actually look at?
- Is beauty in the eye of the beholder?
- Who was studied, and does it generalise?
- How big is the finding?
- Should you trust a single finding?
- How to cite this page
- The bottom line
- Sources
How do researchers measure attractiveness?
Researchers measure attractiveness with ratings, choices, and agreement statistics; ask what task they performed.
Rating scale
A rating scale places judgments numerically. Rhodes et al. used a 10-point scale from 1 = extremely unattractive to 10 = extremely attractive; studies may use 5-point or 7-point scales. APA definition
Likert scale
A Likert scale rates statements from strongly negative to strongly positive; a 1-to-10 judgment is not necessarily Likert. APA definition
Forced choice
Forced choice requires choosing between two options without “I do not know.” Jones et al. showed women 10 masculinized-versus-feminized pairs and recorded strength. Study
Inter-rater reliability
Inter-rater reliability is how similarly independent evaluators rate targets. Cohen’s kappa corrects for agreement that would happen by chance: in McHugh’s worked example, 0.94 raw agreement became a kappa of 0.85. McHugh
Cronbach’s alpha
Cronbach’s alpha estimates internal consistency from 0 to 1, and it rises as more raters or items are added. Rhodes et al. reported alpha = .92 from 13 raters of 60 male faces; Hönekopp argues experts have read high alpha as proof of shared taste when it is not. Tavakol and Dennick
Intraclass correlation
The intraclass correlation (ICC) reflects within-group homogeneity and ranges from 0 to 1. There are 10 forms, so papers should say which one they used; Koo and Li’s guideline calls values below 0.5 poor and above 0.90 excellent, judged on the 95% confidence interval. Koo and Li
Ceiling effect
A ceiling effect crowds observations at a scale’s upper limit, leaving little variation; all-high ratings may not distinguish faces. APA definition
Caveat: reliability is consistency, not truth; raters may share bias.
What do participants actually look at?
Attractiveness depends on the stimulus: average, photo, video, or glimpse. “Face preference” is not one task.
Composite faces and averageness
Composite faces are averaged images; averageness is closeness to the population average. Langlois and Roggman found composites more attractive than most components; Meta-analytic averages summarised by Rhodes et al. put the correlation with attractiveness at r = 0.40 for averageness, r = 0.23 for symmetry, and r = 0.35 for masculinity. Rhodes et al.
Standardized face stimuli
Standardized face stimuli are controlled comparison images. Chicago Face Database offered 158 photographs of Black and White adults aged 18–40; a 2021 expansion added 88 multiracial images. See our facial attractiveness datasets guide.
Static versus dynamic stimuli
Static stimuli are still; dynamic stimuli show movement. Rhodes et al. found 10-second video and still-frame ratings correlated r = 0.83 for 60 male faces; earlier studies found r = 0.19–0.38. Study
Thin slices and 100 ms judgments
Thin slicing means judging from brief clips of expressive behaviour; Ambady and Rosenthal’s meta-analysis found accuracy of r = .39 across 38 results, with no gain from longer observation. Willis and Todorov found 100-ms judgments highly correlated with untimed judgments. Willis and Todorov
Caveat: controlled photos isolate facial cues but cannot represent movement, expression, clothing, and context.
Is beauty in the eye of the beholder?
Partly. Groups of raters, and raters from different cultures, agree on averages, while individual taste stays substantial.
Rater consensus
Rater consensus is similarity of judgments across evaluators. Langlois et al. found cross-cultural agreement across 11 meta-analyses; Cunningham et al. reported r = .93 across Asian, Hispanic, and White American averages. Langlois et al.
Shared versus private taste
Shared taste aligns with average ratings; private taste is person-specific. Germine et al. found group averages correlated r = 0.99 for 102 female faces, yet two random participants agreed on only 48%; Hönekopp’s three experiments found private taste about as powerful as shared taste. Germine et al.
| Claim you will see | What the study actually measured | Term to check |
|---|---|---|
| People agree on who is attractive | Group averages, r = 0.99 | Shared versus private taste |
| Nobody agrees about faces | Two random people, 48% agreement | Individual agreement |
| Beauty is inherited | 22% genetic component in one twin model | Heritability |
| Women prefer masculine faces | Choices between 10 masculinized and feminized pairs | Forced choice |
Twin study and heritability
A twin study compares identical and fraternal twins to estimate heredity and environment. Germine et al. rated 547 identical and 214 non-identical pairs; the ACE model estimated 22% additive genetic, 0% shared-family, and 78% unique-environment or measurement-error components. Study
Caveat: 0.99 is average agreement, not a promise strangers share the crowd reaction.
Who was studied, and does it generalise?
Generalisation depends on how the sample, information, and task resemble life outside the lab. Narrow samples remain narrow.
Convenience sample
A convenience sample recruits available people; results do not automatically generalise. Rhodes et al. recruited 60 male targets and 58 female raters from a university community or researchers’ families and friends: useful, not universal. APA definition
WEIRD samples
WEIRD describes Western, Educated, Industrialized, Rich, Democratic societies; their participants often underpin broad psychology claims. Henrich, Heine, and Norenzayan argued WEIRD subjects can be unusual; the review was not about attractiveness. Review
Statistical power
Statistical power is the probability a test rejects the null when an effect exists; .80 or above is generally acceptable. Low power can miss effects or inflate effects. APA definition
Demand characteristics
Demand characteristics are cues that can change behaviour by suggesting the expected response. An attraction-theory study may measure hint compliance, not preference. APA definition
Ecological validity
Ecological validity is how well results represent wider-world conditions. Student-only samples or isolated face ratings may not generalise. APA definition
Stated versus revealed preferences
Stated preferences are reported wants; revealed preferences are inferred from actual choices. Eastwick and Finkel found stated ideals failed to predict speed-dating desire; the attractiveness studies index separates questionnaires from behaviour.
Self-enhancement in self-ratings
Self-enhancement is seeing oneself more favourably than an observer may. Epley and Whitchurch found people picked attractively morphed versions of their own face as the real one; the same enhancement applied to a friend’s face but not a stranger’s. Study

Caveat: convenience samples are useful; local results become problematic when advertised as universal.
How big is the finding?
A finding can be persuasive yet tiny. Read p value and effect size separately.
Statistical significance
Statistical significance describes how difficult it is to attribute an outcome to chance under a model, expressed with a p value; a large sample can make a small difference significant. APA definition
Effect size
Effect size measures relationship magnitude; Cohen’s d is the standard-deviation distance between means. Conventionally, d = 0.2 is small, d = 0.5 medium, and d ≥ 0.8 large; “significant” is not enough. Effect-size explanation
Meta-analysis
A meta-analysis combines effect-size estimates across studies; there is no minimum number. The Langlois et al. review comprised 11 meta-analyses, not individual studies. APA definition
Halo effect
The halo effect is a rating bias in which a general evaluation spills into specific traits. Dion, Berscheid, and Walster described “what is beautiful is good”; Eagly et al. found a moderate effect, largest for social competence and near zero for integrity. See our halo effect in attractiveness guide.
Beauty premium
The beauty premium is an earnings advantage for better-looking people; the plainness penalty is its disadvantage. Hamermesh and Biddle reported a 5–10% plainness penalty from interviewer ratings in US and Canadian surveys; this is correlational, not causal proof. NBER paper
Caveat: effect-size conventions are rough guides; a “large” association does not identify mechanism or behaviour.
Should you trust a single finding?
Treat one result as a clue, not a verdict. Check unpublished results, flexibility, preregistration, and replication.
Publication bias and the file drawer
Publication bias is the tendency for published results to differ from unpublished results because significant findings are more publishable. Rosenthal imagined journals holding 5% of Type I-error studies while file drawers held 95% of non-significant studies; illustrative, not an estimate. Rosenthal
P-hacking
P-hacking uses flexible analyses to find significant patterns. Simmons, Nelson, and Simonsohn showed undisclosed analytic flexibility raises false-positive rates. Study
Preregistration and registered reports
Preregistration records questions and an analysis plan before outcomes, separating prediction from postdiction. Registered Reports add peer review before data collection, with acceptance when the method is followed. Nosek et al. and Center for Open Science
Replication and Many Labs
Replication repeats a study to check its result through exact, modified, or conceptual forms. The Open Science Collaboration found 97% of 100 originals significant versus 36% of replications; Many Labs 1 replicated 10 of 13 effects across 36 samples, and Many Labs 2 found 15 of 28 significant. Replication definition
Many Labs 2 found median d fell from 0.60 in originals to 0.15 in replications. Many Labs 2
Worked example: the ovulation preference debate
The ovulation debate shows why evidence needs a ledger, not a summary. Gildersleeve et al. meta-analysed 134 effects from 38 published and 12 unpublished studies, reporting genetic-quality shifts; Wood et al. analysed 58 reports and found the evidence largely nonsupportive: effects shrank over time and appeared only with looser fertile-window definitions and only in published research.
Gildersleeve et al.’s reply argued p-curves supported genuine effects. Jones et al.’s largest longitudinal study, N = 584, found no compelling evidence that facial-masculinity preferences tracked hormonal changes; women generally preferred masculinized faces, especially for short-term relationships. Jones et al.
Caveat: conflicting meta-analyses require keeping the claim, sample, measure, and analysis visible.
How to cite this page
Suggested citation: Real World Appeal. “Attractiveness research methods glossary: how to read a study from inter-rater reliability to p-hacking.” realworldappeal.com, September 2026. https://realworldappeal.com/blog/attractiveness-research-methods-glossary
The primary sources are linked throughout, so students, journalists, and AI answer engines can cite the original papers directly. Sources were retrieved 10 September 2026.
The bottom line
The honest answer is not “beauty is objective” or “beauty is entirely personal.” Crowd averages show consensus; individual preferences are private and shaped mostly by each person’s own experience.
Agreement belongs to crowds, not to people. A population average is not a verdict on your face, and a single viral finding is a clue, not a rule. Before letting a study feed an appearance-anxiety spiral, check the sample, the measure, the effect size and whether it replicated; our guide to whether looksmaxxing is pseudoscience applies the same test to forum claims.
For a practical complement, Real World Appeal’s free photo read describes your photo’s first impression before payment. It is not a scientific instrument and does not report a beauty score.
Sources
- American Psychological Association (2018). APA Dictionary of Psychology: rating scale. Link
- American Psychological Association (2018). APA Dictionary of Psychology: Likert scale. Link
- Jones BC, Hahn AC, Fisher CI, et al. (2018). No Compelling Evidence that Preferences for Facial Masculinity Track Changes in Women’s Hormonal Status. Psychological Science, 29(6), 996–1005. Link
- American Psychological Association (2018). APA Dictionary of Psychology: ceiling effect. Link
- American Psychological Association (2018). APA Dictionary of Psychology: interrater reliability. Link
- McHugh ML (2012). Interrater reliability: the kappa statistic. Biochemia Medica, 22(3), 276–282. Link
- American Psychological Association (2018). APA Dictionary of Psychology: Cronbach’s alpha. Link
- Tavakol M, Dennick R (2011). Making sense of Cronbach’s alpha. International Journal of Medical Education, 2, 53–55. Link
- Hönekopp J (2006). Once more: is beauty in the eye of the beholder? Journal of Experimental Psychology: Human Perception and Performance, 32(2), 199–209. Link
- American Psychological Association (2018). APA Dictionary of Psychology: intraclass correlation. Link
- Koo TK, Li MY (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine, 15(2), 155–163. Link
- Langlois JH, Roggman LA (1990). Attractive Faces Are Only Average. Psychological Science, 1(2), 115–121. Link
- Rhodes G (2006). The evolutionary psychology of facial beauty. Annual Review of Psychology, 57, 199–226. Link
- Rhodes G, Lie HC, Thevaraja N, et al. (2011). Facial attractiveness ratings from video-clips and static images tell the same story. PLoS One, 6(11), e26653. Link
- Ma DS, Correll J, Wittenbrink B (2015). The Chicago face database: A free stimulus set of faces and norming data. Behavior Research Methods, 47(4), 1122–1135. Link
- Ma DS, Kantner J, Wittenbrink B (2021). Chicago Face Database: Multiracial expansion. Behavior Research Methods, 53(3), 1289–1300. Link
- Ambady N, Rosenthal R (1992). Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2), 256–274. Link
- Willis J, Todorov A (2006). First impressions: making up your mind after a 100-ms exposure to a face. Psychological Science, 17(7), 592–598. Link
- Langlois JH, Kalakanis L, Rubenstein AJ, et al. (2000). Maxims or myths of beauty? A meta-analytic and theoretical review. Psychological Bulletin, 126(3), 390–423. Link
- Cunningham MR, Roberts AR, Barbee AP, et al. (1995). Their ideas of beauty are, on the whole, the same as ours. Journal of Personality and Social Psychology, 68(2), 261–279. Link
- Germine L, Russell R, Bronstad PM, et al. (2015). Individual Aesthetic Preferences for Faces Are Shaped Mostly by Environments, Not Genes. Current Biology, 25(20), 2684–2689. Link
- American Psychological Association (2018). APA Dictionary of Psychology: twin study. Link
- American Psychological Association (2018). APA Dictionary of Psychology: convenience sampling. Link
- Henrich J, Heine SJ, Norenzayan A (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83. Link
- Arnett JJ (2008). The neglected 95%: why American psychology needs to become less American. American Psychologist, 63(7), 602–614. Link
- Thalmayer AG, Toscanelli C, Arnett JJ (2021). The neglected 95% revisited: Is American psychology becoming less American? American Psychologist, 76(1), 116–129. Link
- American Psychological Association (2018). APA Dictionary of Psychology: power. Link
- Button KS, Ioannidis JP, Mokrysz C, et al. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. Link
- American Psychological Association (2018). APA Dictionary of Psychology: demand characteristics. Link
- American Psychological Association (2018). APA Dictionary of Psychology: ecological validity. Link
- Eastwick PW, Finkel EJ (2008). Sex differences in mate preferences revisited: do people know what they initially desire in a romantic partner? Journal of Personality and Social Psychology, 94(2), 245–264. Link
- Wikipedia contributors (n.d.). Revealed preference. Link
- Epley N, Whitchurch E (2008). Mirror, mirror on the wall: enhancement in self-recognition. Personality and Social Psychology Bulletin, 34(9), 1159–1170. Link
- American Psychological Association (2018). APA Dictionary of Psychology: statistical significance. Link
- Sullivan GM, Feinn R (2012). Using Effect Size—or Why the P Value Is Not Enough. Journal of Graduate Medical Education, 4(3), 279–282. Link
- American Psychological Association (2018). APA Dictionary of Psychology: effect size. Link
- American Psychological Association (2018). APA Dictionary of Psychology: meta-analysis. Link
- American Psychological Association (2018). APA Dictionary of Psychology: halo effect. Link
- Dion K, Berscheid E, Walster E (1972). What is beautiful is good. Journal of Personality and Social Psychology, 24(3), 285–290. Link
- Eagly AH, Ashmore RD, Makhijani MG, Longo LC (1991). What is beautiful is good, but…: A meta-analytic review of research on the physical attractiveness stereotype. Psychological Bulletin, 110(1), 109–128. Link
- Hamermesh DS, Biddle JE (1994). Beauty and the Labor Market. American Economic Review, 84(5), 1174–1194. Link
- Hamermesh DS, Biddle JE (1993). Beauty and the Labor Market. NBER Working Paper No. 4518. Link
- American Psychological Association (2018). APA Dictionary of Psychology: publication bias. Link
- Rosenthal R (1979). The file drawer problem and tolerance for null results. Psychological Bulletin, 86(3), 638–641. Link
- Wikipedia contributors (n.d.). Data dredging. Link
- Simmons JP, Nelson LD, Simonsohn U (2011). False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. Link
- Nosek BA, Ebersole CR, DeHaven AC, Mellor DT (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. Link
- Center for Open Science (n.d.). Registered Reports. Link
- Wikipedia contributors (n.d.). Registered report. Link
- American Psychological Association (2018). APA Dictionary of Psychology: replication. Link
- Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. Link
- Klein RA, Ratliff KA, et al. (2014). Investigating Variation in Replicability: A “Many Labs” Replication Project. Social Psychology, 45(3), 142–152. Link
- Klein RA, et al. (2018). Many Labs 2: Investigating Variation in Replicability Across Samples and Settings. Advances in Methods and Practices in Psychological Science, 1(4). Link
- Gildersleeve K, Haselton MG, Fales MR (2014). Do women’s mate preferences change across the ovulatory cycle? A meta-analytic review. Psychological Bulletin, 140(5), 1205–1259. Link
- Wood W, Kressel L, Joshi PD, Louie B (2014). Meta-Analysis of Menstrual Cycle Effects on Women’s Mate Preferences. Emotion Review, 6(3), 229–249. Link
- Gildersleeve K, Haselton MG, Fales MR (2014). Meta-analyses and p-curves support robust cycle shifts in women’s mate preferences. Psychological Bulletin, 140(5), 1272–1280. Link
- Wikipedia contributors (n.d.). Two-alternative forced choice. Link
Frequently asked questions
What does inter-rater reliability mean in attractiveness studies?
It measures how similarly independent evaluators rate the same faces. A study can show high agreement among group averages while individual people still disagree often; the attractiveness studies index puts that distinction in context.
Is beauty in the eye of the beholder according to research?
Both parts are true: groups often agree about average attractiveness, while private taste remains substantial. Germine et al. found group-mean agreement of r = 0.99 but only 48% agreement between two random participants; see the first-impression psychology glossary.
What is the difference between stated and revealed preferences?
Stated preferences are what people say they want; revealed preferences are inferred from actual choices. Eastwick and Finkel found that ideal preferences stated before speed dating failed to predict whom participants actually desired; our attractiveness myths fact check explains why viral claims often miss this distinction.
Does a significant attractiveness finding prove the effect is important?
No. Statistical significance concerns compatibility with chance, while effect size describes magnitude, and large samples can make small differences significant. Check the effect size and replication evidence in this attractiveness research statistics guide.
What should I check before believing a study about attraction?
Check the sample, measurement, effect size, preregistration, publication context, and whether the result replicated. A population average is not a verdict on your face; if you want a practical complement, Real World Appeal offers a free first-impression photo read at /test, not a research instrument.

