Real World Appeal
ResearchSeptember 10, 202616 min read

Attractiveness research methods glossary: how to read a study from inter-rater reliability to p-hacking

A plain-English attractiveness research methods glossary: 27 terms, from r = 0.99 crowd agreement to stated versus revealed preference.

A woman reading research papers at a library desk
Photo: Ron Lach

You have seen the post: “science says women prefer X.” Before you believe it, you need the words to read the study behind it. This glossary gives you 27 of them, each with a plain definition, why it matters for attractiveness research, and a real example.

The short answer to the biggest question: crowd averages can agree almost perfectly, while two random people agree on only 48% of face preferences. Agreement belongs to crowds, not to people.

Key numbers

  • r = 0.99 — men’s and women’s average ratings agreed for 102 female faces. Source
  • 48% — two randomly selected participants agreed on face preferences. Source
  • 22%, 0%, 78% — genetic, shared-family, and unique-environment or measurement-error components in individual face preferences. Source
  • r = 0.83 — photo and video attractiveness ratings correlated for 60 male faces. Source
  • 36% versus 97% — replication and original significance rates across 100 general psychology studies. Source
  • d = 0.60 versus 0.15 — median effect size in Many Labs 2 originals and replications. Source
  • 5–10% — the plainness penalty reported in labor-market research. Source

Table of contents

How do researchers measure attractiveness?

Researchers measure attractiveness with ratings, choices, and agreement statistics; ask what task they performed.

Rating scale

A rating scale places judgments numerically. Rhodes et al. used a 10-point scale from 1 = extremely unattractive to 10 = extremely attractive; studies may use 5-point or 7-point scales. APA definition

Likert scale

A Likert scale rates statements from strongly negative to strongly positive; a 1-to-10 judgment is not necessarily Likert. APA definition

Forced choice

Forced choice requires choosing between two options without “I do not know.” Jones et al. showed women 10 masculinized-versus-feminized pairs and recorded strength. Study

Inter-rater reliability

Inter-rater reliability is how similarly independent evaluators rate targets. Cohen’s kappa corrects for agreement that would happen by chance: in McHugh’s worked example, 0.94 raw agreement became a kappa of 0.85. McHugh

Cronbach’s alpha

Cronbach’s alpha estimates internal consistency from 0 to 1, and it rises as more raters or items are added. Rhodes et al. reported alpha = .92 from 13 raters of 60 male faces; Hönekopp argues experts have read high alpha as proof of shared taste when it is not. Tavakol and Dennick

Intraclass correlation

The intraclass correlation (ICC) reflects within-group homogeneity and ranges from 0 to 1. There are 10 forms, so papers should say which one they used; Koo and Li’s guideline calls values below 0.5 poor and above 0.90 excellent, judged on the 95% confidence interval. Koo and Li

Ceiling effect

A ceiling effect crowds observations at a scale’s upper limit, leaving little variation; all-high ratings may not distinguish faces. APA definition

Caveat: reliability is consistency, not truth; raters may share bias.

What do participants actually look at?

Attractiveness depends on the stimulus: average, photo, video, or glimpse. “Face preference” is not one task.

Composite faces and averageness

Composite faces are averaged images; averageness is closeness to the population average. Langlois and Roggman found composites more attractive than most components; Meta-analytic averages summarised by Rhodes et al. put the correlation with attractiveness at r = 0.40 for averageness, r = 0.23 for symmetry, and r = 0.35 for masculinity. Rhodes et al.

Standardized face stimuli

Standardized face stimuli are controlled comparison images. Chicago Face Database offered 158 photographs of Black and White adults aged 18–40; a 2021 expansion added 88 multiracial images. See our facial attractiveness datasets guide.

Static versus dynamic stimuli

Static stimuli are still; dynamic stimuli show movement. Rhodes et al. found 10-second video and still-frame ratings correlated r = 0.83 for 60 male faces; earlier studies found r = 0.19–0.38. Study

Thin slices and 100 ms judgments

Thin slicing means judging from brief clips of expressive behaviour; Ambady and Rosenthal’s meta-analysis found accuracy of r = .39 across 38 results, with no gain from longer observation. Willis and Todorov found 100-ms judgments highly correlated with untimed judgments. Willis and Todorov

Hands working through charts and notes on a desk
Photo: Lukas Blazek / Pexels

Caveat: controlled photos isolate facial cues but cannot represent movement, expression, clothing, and context.

Is beauty in the eye of the beholder?

Partly. Groups of raters, and raters from different cultures, agree on averages, while individual taste stays substantial.

Rater consensus

Rater consensus is similarity of judgments across evaluators. Langlois et al. found cross-cultural agreement across 11 meta-analyses; Cunningham et al. reported r = .93 across Asian, Hispanic, and White American averages. Langlois et al.

Shared versus private taste

Shared taste aligns with average ratings; private taste is person-specific. Germine et al. found group averages correlated r = 0.99 for 102 female faces, yet two random participants agreed on only 48%; Hönekopp’s three experiments found private taste about as powerful as shared taste. Germine et al.

Claim you will seeWhat the study actually measuredTerm to check
People agree on who is attractiveGroup averages, r = 0.99Shared versus private taste
Nobody agrees about facesTwo random people, 48% agreementIndividual agreement
Beauty is inherited22% genetic component in one twin modelHeritability
Women prefer masculine facesChoices between 10 masculinized and feminized pairsForced choice

Twin study and heritability

A twin study compares identical and fraternal twins to estimate heredity and environment. Germine et al. rated 547 identical and 214 non-identical pairs; the ACE model estimated 22% additive genetic, 0% shared-family, and 78% unique-environment or measurement-error components. Study

Caveat: 0.99 is average agreement, not a promise strangers share the crowd reaction.

Who was studied, and does it generalise?

Generalisation depends on how the sample, information, and task resemble life outside the lab. Narrow samples remain narrow.

Convenience sample

A convenience sample recruits available people; results do not automatically generalise. Rhodes et al. recruited 60 male targets and 58 female raters from a university community or researchers’ families and friends: useful, not universal. APA definition

WEIRD samples

WEIRD describes Western, Educated, Industrialized, Rich, Democratic societies; their participants often underpin broad psychology claims. Henrich, Heine, and Norenzayan argued WEIRD subjects can be unusual; the review was not about attractiveness. Review

Statistical power

Statistical power is the probability a test rejects the null when an effect exists; .80 or above is generally acceptable. Low power can miss effects or inflate effects. APA definition

Demand characteristics

Demand characteristics are cues that can change behaviour by suggesting the expected response. An attraction-theory study may measure hint compliance, not preference. APA definition

Ecological validity

Ecological validity is how well results represent wider-world conditions. Student-only samples or isolated face ratings may not generalise. APA definition

Stated versus revealed preferences

Stated preferences are reported wants; revealed preferences are inferred from actual choices. Eastwick and Finkel found stated ideals failed to predict speed-dating desire; the attractiveness studies index separates questionnaires from behaviour.

Self-enhancement in self-ratings

Self-enhancement is seeing oneself more favourably than an observer may. Epley and Whitchurch found people picked attractively morphed versions of their own face as the real one; the same enhancement applied to a friend’s face but not a stranger’s. Study

A printed bar graph of research data
Photo: RDNE Stock project / Pexels

Caveat: convenience samples are useful; local results become problematic when advertised as universal.

How big is the finding?

A finding can be persuasive yet tiny. Read p value and effect size separately.

Statistical significance

Statistical significance describes how difficult it is to attribute an outcome to chance under a model, expressed with a p value; a large sample can make a small difference significant. APA definition

Effect size

Effect size measures relationship magnitude; Cohen’s d is the standard-deviation distance between means. Conventionally, d = 0.2 is small, d = 0.5 medium, and d ≥ 0.8 large; “significant” is not enough. Effect-size explanation

Meta-analysis

A meta-analysis combines effect-size estimates across studies; there is no minimum number. The Langlois et al. review comprised 11 meta-analyses, not individual studies. APA definition

Halo effect

The halo effect is a rating bias in which a general evaluation spills into specific traits. Dion, Berscheid, and Walster described “what is beautiful is good”; Eagly et al. found a moderate effect, largest for social competence and near zero for integrity. See our halo effect in attractiveness guide.

Beauty premium

The beauty premium is an earnings advantage for better-looking people; the plainness penalty is its disadvantage. Hamermesh and Biddle reported a 5–10% plainness penalty from interviewer ratings in US and Canadian surveys; this is correlational, not causal proof. NBER paper

Caveat: effect-size conventions are rough guides; a “large” association does not identify mechanism or behaviour.

Should you trust a single finding?

Treat one result as a clue, not a verdict. Check unpublished results, flexibility, preregistration, and replication.

Publication bias and the file drawer

Publication bias is the tendency for published results to differ from unpublished results because significant findings are more publishable. Rosenthal imagined journals holding 5% of Type I-error studies while file drawers held 95% of non-significant studies; illustrative, not an estimate. Rosenthal

P-hacking

P-hacking uses flexible analyses to find significant patterns. Simmons, Nelson, and Simonsohn showed undisclosed analytic flexibility raises false-positive rates. Study

Preregistration and registered reports

Preregistration records questions and an analysis plan before outcomes, separating prediction from postdiction. Registered Reports add peer review before data collection, with acceptance when the method is followed. Nosek et al. and Center for Open Science

Replication and Many Labs

Replication repeats a study to check its result through exact, modified, or conceptual forms. The Open Science Collaboration found 97% of 100 originals significant versus 36% of replications; Many Labs 1 replicated 10 of 13 effects across 36 samples, and Many Labs 2 found 15 of 28 significant. Replication definition

Many Labs 2 found median d fell from 0.60 in originals to 0.15 in replications. Many Labs 2

Worked example: the ovulation preference debate

The ovulation debate shows why evidence needs a ledger, not a summary. Gildersleeve et al. meta-analysed 134 effects from 38 published and 12 unpublished studies, reporting genetic-quality shifts; Wood et al. analysed 58 reports and found the evidence largely nonsupportive: effects shrank over time and appeared only with looser fertile-window definitions and only in published research.

Gildersleeve et al.’s reply argued p-curves supported genuine effects. Jones et al.’s largest longitudinal study, N = 584, found no compelling evidence that facial-masculinity preferences tracked hormonal changes; women generally preferred masculinized faces, especially for short-term relationships. Jones et al.

Caveat: conflicting meta-analyses require keeping the claim, sample, measure, and analysis visible.

How to cite this page

Suggested citation: Real World Appeal. “Attractiveness research methods glossary: how to read a study from inter-rater reliability to p-hacking.” realworldappeal.com, September 2026. https://realworldappeal.com/blog/attractiveness-research-methods-glossary

The primary sources are linked throughout, so students, journalists, and AI answer engines can cite the original papers directly. Sources were retrieved 10 September 2026.

The bottom line

The honest answer is not “beauty is objective” or “beauty is entirely personal.” Crowd averages show consensus; individual preferences are private and shaped mostly by each person’s own experience.

Agreement belongs to crowds, not to people. A population average is not a verdict on your face, and a single viral finding is a clue, not a rule. Before letting a study feed an appearance-anxiety spiral, check the sample, the measure, the effect size and whether it replicated; our guide to whether looksmaxxing is pseudoscience applies the same test to forum claims.

For a practical complement, Real World Appeal’s free photo read describes your photo’s first impression before payment. It is not a scientific instrument and does not report a beauty score.

Sources

  • American Psychological Association (2018). APA Dictionary of Psychology: rating scale. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: Likert scale. Link
  • Jones BC, Hahn AC, Fisher CI, et al. (2018). No Compelling Evidence that Preferences for Facial Masculinity Track Changes in Women’s Hormonal Status. Psychological Science, 29(6), 996–1005. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: ceiling effect. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: interrater reliability. Link
  • McHugh ML (2012). Interrater reliability: the kappa statistic. Biochemia Medica, 22(3), 276–282. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: Cronbach’s alpha. Link
  • Tavakol M, Dennick R (2011). Making sense of Cronbach’s alpha. International Journal of Medical Education, 2, 53–55. Link
  • Hönekopp J (2006). Once more: is beauty in the eye of the beholder? Journal of Experimental Psychology: Human Perception and Performance, 32(2), 199–209. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: intraclass correlation. Link
  • Koo TK, Li MY (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine, 15(2), 155–163. Link
  • Langlois JH, Roggman LA (1990). Attractive Faces Are Only Average. Psychological Science, 1(2), 115–121. Link
  • Rhodes G (2006). The evolutionary psychology of facial beauty. Annual Review of Psychology, 57, 199–226. Link
  • Rhodes G, Lie HC, Thevaraja N, et al. (2011). Facial attractiveness ratings from video-clips and static images tell the same story. PLoS One, 6(11), e26653. Link
  • Ma DS, Correll J, Wittenbrink B (2015). The Chicago face database: A free stimulus set of faces and norming data. Behavior Research Methods, 47(4), 1122–1135. Link
  • Ma DS, Kantner J, Wittenbrink B (2021). Chicago Face Database: Multiracial expansion. Behavior Research Methods, 53(3), 1289–1300. Link
  • Ambady N, Rosenthal R (1992). Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2), 256–274. Link
  • Willis J, Todorov A (2006). First impressions: making up your mind after a 100-ms exposure to a face. Psychological Science, 17(7), 592–598. Link
  • Langlois JH, Kalakanis L, Rubenstein AJ, et al. (2000). Maxims or myths of beauty? A meta-analytic and theoretical review. Psychological Bulletin, 126(3), 390–423. Link
  • Cunningham MR, Roberts AR, Barbee AP, et al. (1995). Their ideas of beauty are, on the whole, the same as ours. Journal of Personality and Social Psychology, 68(2), 261–279. Link
  • Germine L, Russell R, Bronstad PM, et al. (2015). Individual Aesthetic Preferences for Faces Are Shaped Mostly by Environments, Not Genes. Current Biology, 25(20), 2684–2689. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: twin study. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: convenience sampling. Link
  • Henrich J, Heine SJ, Norenzayan A (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83. Link
  • Arnett JJ (2008). The neglected 95%: why American psychology needs to become less American. American Psychologist, 63(7), 602–614. Link
  • Thalmayer AG, Toscanelli C, Arnett JJ (2021). The neglected 95% revisited: Is American psychology becoming less American? American Psychologist, 76(1), 116–129. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: power. Link
  • Button KS, Ioannidis JP, Mokrysz C, et al. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: demand characteristics. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: ecological validity. Link
  • Eastwick PW, Finkel EJ (2008). Sex differences in mate preferences revisited: do people know what they initially desire in a romantic partner? Journal of Personality and Social Psychology, 94(2), 245–264. Link
  • Wikipedia contributors (n.d.). Revealed preference. Link
  • Epley N, Whitchurch E (2008). Mirror, mirror on the wall: enhancement in self-recognition. Personality and Social Psychology Bulletin, 34(9), 1159–1170. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: statistical significance. Link
  • Sullivan GM, Feinn R (2012). Using Effect Size—or Why the P Value Is Not Enough. Journal of Graduate Medical Education, 4(3), 279–282. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: effect size. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: meta-analysis. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: halo effect. Link
  • Dion K, Berscheid E, Walster E (1972). What is beautiful is good. Journal of Personality and Social Psychology, 24(3), 285–290. Link
  • Eagly AH, Ashmore RD, Makhijani MG, Longo LC (1991). What is beautiful is good, but…: A meta-analytic review of research on the physical attractiveness stereotype. Psychological Bulletin, 110(1), 109–128. Link
  • Hamermesh DS, Biddle JE (1994). Beauty and the Labor Market. American Economic Review, 84(5), 1174–1194. Link
  • Hamermesh DS, Biddle JE (1993). Beauty and the Labor Market. NBER Working Paper No. 4518. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: publication bias. Link
  • Rosenthal R (1979). The file drawer problem and tolerance for null results. Psychological Bulletin, 86(3), 638–641. Link
  • Wikipedia contributors (n.d.). Data dredging. Link
  • Simmons JP, Nelson LD, Simonsohn U (2011). False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. Link
  • Nosek BA, Ebersole CR, DeHaven AC, Mellor DT (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. Link
  • Center for Open Science (n.d.). Registered Reports. Link
  • Wikipedia contributors (n.d.). Registered report. Link
  • American Psychological Association (2018). APA Dictionary of Psychology: replication. Link
  • Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. Link
  • Klein RA, Ratliff KA, et al. (2014). Investigating Variation in Replicability: A “Many Labs” Replication Project. Social Psychology, 45(3), 142–152. Link
  • Klein RA, et al. (2018). Many Labs 2: Investigating Variation in Replicability Across Samples and Settings. Advances in Methods and Practices in Psychological Science, 1(4). Link
  • Gildersleeve K, Haselton MG, Fales MR (2014). Do women’s mate preferences change across the ovulatory cycle? A meta-analytic review. Psychological Bulletin, 140(5), 1205–1259. Link
  • Wood W, Kressel L, Joshi PD, Louie B (2014). Meta-Analysis of Menstrual Cycle Effects on Women’s Mate Preferences. Emotion Review, 6(3), 229–249. Link
  • Gildersleeve K, Haselton MG, Fales MR (2014). Meta-analyses and p-curves support robust cycle shifts in women’s mate preferences. Psychological Bulletin, 140(5), 1272–1280. Link
  • Wikipedia contributors (n.d.). Two-alternative forced choice. Link

Frequently asked questions

What does inter-rater reliability mean in attractiveness studies?

It measures how similarly independent evaluators rate the same faces. A study can show high agreement among group averages while individual people still disagree often; the attractiveness studies index puts that distinction in context.

Is beauty in the eye of the beholder according to research?

Both parts are true: groups often agree about average attractiveness, while private taste remains substantial. Germine et al. found group-mean agreement of r = 0.99 but only 48% agreement between two random participants; see the first-impression psychology glossary.

What is the difference between stated and revealed preferences?

Stated preferences are what people say they want; revealed preferences are inferred from actual choices. Eastwick and Finkel found that ideal preferences stated before speed dating failed to predict whom participants actually desired; our attractiveness myths fact check explains why viral claims often miss this distinction.

Does a significant attractiveness finding prove the effect is important?

No. Statistical significance concerns compatibility with chance, while effect size describes magnitude, and large samples can make small differences significant. Check the effect size and replication evidence in this attractiveness research statistics guide.

What should I check before believing a study about attraction?

Check the sample, measurement, effect size, preregistration, publication context, and whether the result replicated. A population average is not a verdict on your face; if you want a practical complement, Real World Appeal offers a free first-impression photo read at /test, not a research instrument.

Test your own first-impression score

1 minute, two photos + a few quick details. Concrete improvement levers ranked by how much they actually move the dial.

Start the test

Related reading