IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

2026-07-27Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors summarize a competition where teams adapted a foundation model called CLIP for face recognition using only synthetic (computer-generated) training data. The competition had two parts: one with lots of data and one with limited data to test different adaptation styles. They judged entries by how well they recognized faces across various test sets and also checked fairness for different demographic groups. The results showed that adapting the foundation model with synthetic data improved its performance compared to the original model. Different fine-tuning methods worked best depending on how much data was available.

Foundation modelFace recognitionSynthetic dataFine-tuningCLIPIDPERTURBVerificationIdentificationBorda countFairness evaluation
Authors
Tahar Chettaoui, Guray Ozgur, Eduarda Caldeira, Arturas Nakvosas, Hatef Otroshi Shahreza, Sébastien Marcel, Rishabh Shukla, Aditya Takkar, Rushil Khullar, Lalak Yadav, Gourav Gupta, Anant Gupta, Shiqi Yu, Vitomir Struc, Naser Damer, Fadi Boutros
Abstract
This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams across two complementary tracks: a Full Data Track, in which participants adapt the CLIP ViT-L/14 foundation model using large-scale synthetic identity data, and a Limited Data Track, designed to reflect more resource-constrained adaptation regimes. All training data was generated exclusively using IDPERTURB. Submitted solutions are ranked based on verification and identification performance across a diverse suite of benchmarks, including LFW, CFP-FP, AgeDB-30, CALFW, CPLFW, IJB-B, IJB-C, and TinyFace, using the Borda count method. Fairness evaluation is additionally conducted on the RFW dataset across four demographic groups. The results demonstrate that adaptation of the CLIP foundation model with synthetic training data substantially improves over the off-the-shelf model and, in several cases, surpasses the baseline. Notably, full fine-tuning with Sub-Center ArcFace (DMSTI-Neurotechnology) leads the Full Data Track, while rank-stabilized LoRA adaptation (Idiap-BSP) proves most effective under limited-data conditions.