2026/05/31 by Caio Petrucci, Leo Sampaio Ferraz Ribeiro, Sandra Avila
Computer Science · #cs.CV #cs.AI
13 pages; 3 figures; 8 tables; 1 algorithm
arxiv created 2026/07/30 · arxiv updated 2026/07/31
Age estimation from facial images typically relies on training data that includes images of minors, a practice that raises ethical, legal, and privacy concerns and that child-data governance frameworks explicitly advise against. While the task remains relevant (e.g., for detecting child sexual abuse imagery), we advocate against using data from minors entirely and quantify what the exclusion costs in accuracy. We formalize age estimation without children's training data as a generalized zero-shot learning (GZSL) problem: age intervals present during training are seen classes and withheld intervals are unseen, with models evaluated jointly on both. The generalized setting, rather than conventional zero-shot evaluation on unseen classes alone, is the appropriate one here because a deployed estimator must operate across the entire lifespan, not only on the interval withheld from it. Revisiting six widely used datasets, we introduce standardized splits with strict age-group separation. For datasets with identity annotations, subject-age-exclusive splits prevent identity leakage across the seen/unseen boundary. Evaluating nine state-of-the-art age estimation methods under this protocol reveals that all of them fail to generalize to unseen age groups, suffering substantial degradation --- on average 46.4%, and up to 52.8% --- relative to the supervised baseline. Moreover, models do not simply degrade: they systematically anchor predictions for unseen ages to nearby seen classes, a manifestation of the well-known seen-class bias in generalized zero-shot learning.