Release Date: 10/2/2026
Format: PNG, JSON, CSV
Size: 2.79 GB
Contains face-cropped and aligned FHIBE images and their corresponding JSON annotations. The face crops were created from the full-resolution original images, not from the downsampled versions. The aligned face crops were constructed by cropping oriented rectangles based on two eye landmarks and two mouth landmarks. These rectangles were first resized to 4,096 x 4,096 pixels using bilinear filtering, and then downsampled to 512 x 512 pixels using Lanczos filtering. Only faces with visible eye and mouth landmarks were included in the final cropped-and-aligned set. The dataset package contains a filepaths.csv file which summarizes all the face image and annotation file paths contained in the nested sub-directories. All paths are relative to the location of filepaths.csv. This file can be used for automatically reading/loading images and annotations. The filepaths.csv file contains the following columns: uid: A unique identifier of the image. Can be used to map data across all FHIBE variants. img: A list represented as string, containing the relative paths to all face image files extracted from the original image. json: A list represented as string, containing the relative paths to all face annotation files corresponding to the extracted face images. For support related to this dataset, including consent revocation, please contact [email protected] You can find FHIBE's data card and transparency documentation here: fairnessbenchmark.ai.sony/data-card
Restrictions/Special Constraints
As set out in Section 2.1 of the Terms of Use, FHIBE is primarily an evaluation dataset and may only be used as a training dataset for the development of bias diagnostic and mitigation tools or methods.
Forbidden Usage
As a user of FHIBE, you may not, and may not permit or assist others, to:
Attempt to re-identify any individuals in FHIBE.
Attempt to infer, predict, or label any sensitive or objectionable attributes to the Licensed Images, such as race or ethnicity, gender or sexual orientation, political opinions, religion or religious beliefs, genetic data, propensity towards crime, personality, attractiveness, etc., except in connection with detecting bias in model analysis.
Use biometric data in any Licensed Images to perform any Processing activities (as defined in the Data Sharing Addendum in Exhibit B of the Terms of Use) unrelated to bias evaluation and mitigation, including without limitation facial recognition.
Use FHIBE in any way that causes reputational harm to the individuals in the Licensed Images or other parties.
Use FHIBE in any manner that violates any applicable law, such as privacy or security laws or regulations as well as including accessing FHIBE from a country that is subject to a U.S. Government embargo.
Use FHIBE for any law enforcement, military, arms or surveillance purpose or any other purpose other than the Purpose.
Use FHIBE for evaluating AI systems that (are banned in, or violate restrictions of, the United States, the European Union and in your applicable jurisdiction.
For more restrictions please refer to the Terms of Use and Acceptable Use Policy.
Ethical Review
Data collection commenced after April 23, 2023, following Institutional Review Board (IRB) approval from WCG Clinical, Inc. (study number 1352290). All participants provided informed consent, and image subjects additionally consented to their identifiable images being included in the dataset.
Intended Use
FHIBE is intended to be used as a bias/fairness evaluation dataset for machine learning models. The diverse set of images and rich annotations enable evaluations over a wide range of computer vision tasks, including person detection and segmentation, face detection and verification, keypoint (pose) estimation, and visual question answering. Applications for which FHIBE has been used include surfacing gender biases in vision language model question answering and identifying hair style variability as a leading contributor to gender bias in popular face verification models.
Total files: 17895
Total images: 8947
Total annotation files: 8947
Number of unique human subjects: 1903
Image format: PNG
Annotation file format: JSON
Number of rows in filepaths.csv: 10901
Number of columns in filepaths.csv: 3
Columns of filepaths.csv: [‘uid’, ‘img’, ‘json’]
Compensated · similar languages
High-specularity optical edge-case dataset (425 assets) for benchmarking computer vision and autonomous navigation against severe specular reflections.
Versioned software-engineering corpus with code, configuration, workflows, and AI-tooling artifacts across frontend, backend, cloud, and infrastructure.