Does the engine behind it actually work?

Fair question — most "face raters" never show any evidence. So here's ours, nothing hidden. We took thousands of real faces that large groups of people had already rated, ran the model that powers our analysis, and checked how often it agreed with them. We also checked it's fair across ethnicities. (Your own report is descriptive — it shows your proportions, not this score.)

How accurate is the engine? (the numbers)

The model behind our analysis is trained on a large public research benchmark (5,500 faces rated by large panels). Below is its accuracy on held-out faces it never saw during training. "Agreement" is how often it ranks a pair of faces the same way people did; ρ is rank correlation (50% / 0.0 = chance). Your own report is descriptive and does not show this score.

WomenFacesAgreementρ
All women2,75088%0.92
Asian2,00087%0.92
White75087%0.91
MenFacesAgreementρ
All men2,75087%0.90
Asian2,00087%0.91
White75084%0.87

Is it fair across ethnicities? (tested on faces from a different database)

The toughest test: we ran it on a separate, independent face database — a completely separate set the model never trained on, spanning six ethnic groups — and measured how often it agreed with that database's raters within each group.

Group (unseen data)WomenMen
Asian66%77%
White75%70%
Black78%64%
Latino65%69%
Indian67%62%
Multiracial70%67%

Honest caveats: on this unseen database accuracy is lower than in-domain (different cameras, rater pools and tastes), and the non-Asian groups are small (n ~25–100), so those figures are noisier. The model also carries small level offsets between groups. We use that database only to check fairness, not to set scores — its photos are deliberately neutral and range-restricted at the top. This is a guide, not an objective verdict.

For the curious — how this was measured

The model behind our analysis produces a within-sex score we use here purely to validate it against human ratings. Your own report does not show this score — it describes the classical proportions, measured separately, as an explainable breakdown. Accuracy is measured on faces the model never saw during training, and fairness is checked on a separate, independent set balanced across ethnic groups. We don't publish the underlying datasets or raw faces here. This measures the engine, not a verdict on anyone's appearance.