The face and voice of a person have distinctive characteristics and are widely used as biometric cues for person authentication, either independently or jointly in multimodal systems. The strong perceptual correspondence humans establish between faces and voices has motivated a large body of work on automated face-voice association. Most of these approaches, however, are evaluated under controlled experimental settings that neglect the varying conditions of real-world scenarios. Two important sources of this variation are the language spoken by the speakers and the distribution of demographic traits across speakers. Since a substantial proportion of the global population is bilingual or multilingual, it matters whether face-voice association models remain reliable when the same speaker communicates in different languages. At the same time, gender-controlled evaluation is necessary to determine whether models rely on speaker-specific cross-modal cues rather than simply exploiting demographic differences between the speakers forming the positive and negative pairs.

The goal of the Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge is therefore two-fold: to evaluate face-voice association methods in multilingual conditions while preventing models from using gender as a proxy for identity, and to foster the development of multilingual and gender-aware methods that learn identity-specific associations. FLAG 2027 builds on the prior FAME 2026 and FAME 2024 Grand Challenges hosted at IEEE ICASSP and ACM Multimedia, and has been accepted for support under the IEEE Signal Processing Society Challenge Program.

TIMELINE

Tentative — dates may be updated.

  1. Teams register for the challenge via the registration form.

    + Add to Google Calendar
  2. Training and development sets, pre-extracted features and the baseline model are available. Teams compare results on the CodaBench leaderboard — up to 150 submissions, at most 15 per day.

    + Add to Google Calendar
  3. Evaluation inputs are released on the first day, without ground-truth labels. The total number of submissions is limited to 15.

    + Add to Google Calendar
  4. Final rankings are announced, based on the overall score (sum of the four EERs divided by 4).

    + Add to Google Calendar
  5. Teams submit a 2-page system description in the ICASSP template, together with a link to a working version of their code. Teams without a system description will be disqualified.

    + Add to Google Calendar
  6. Deadline for challenge paper submission.

    + Add to Google Calendar

EVALUATION TRACKS

Language impact

Face-voice association is established through a cross-modal verification task: given a face and a voice, the system decides whether both belong to the same identity. The dataset is split following the unseen-unheard configuration, in which train and test splits consist of disjoint speakers. Each video shows a speaker using one language only, while every speaker appears in videos of at least two distinct languages. Evaluation is carried out on both a heard language, available during training, and an unheard language, not available during training. This track measures whether a voice recorded in a different language can still be associated with the corresponding speaker's face.

Gender impact

This track evaluates gender-controlled verification. Negative samples are restricted to speakers of the same gender as the speaker of the positive sample, so a model can no longer infer identity from demographic differences alone. This yields a more rigorous assessment of whether models learn genuinely identity-specific face-voice associations rather than using gender as a shortcut.

FLAG 2027 Challenge overview

Figure 1: Overview of the FLAG 2027 Challenge. The "language impact" track focuses on whether a voice recorded in different languages can be correctly associated with the corresponding speaker's face. The "gender impact" track evaluates gender-controlled verification, where negative samples are restricted to speakers of the same gender as the positive sample speaker.

DATASET

Building on the prior FAME 2024 and FAME 2026 Grand Challenges, the MAV-Celeb corpus has been extended with an additional language: a newly curated split of 100 English–Bengali speakers (MAV-Celeb v4). This release provides both language and gender annotations, which is what makes the two FLAG tracks possible.

100Speakers
2Languages
557Videos
71.5 hSpeech

Corpus composition

Alongside the new split, the earlier releases remain available: MAV-Celeb v1 covers English and Urdu speakers, and MAV-Celeb v3 covers English and German speakers. All splits contain distinct, non-overlapping speaker identities.

Data collection and curation

As in previous releases, the audio-visual samples are obtained from YouTube videos of celebrities appearing in interviews, talk shows, and television debates. The visual data spans a wide range of recording setups, including variation in pose, motion blur, background clutter, video quality, occlusion, and lighting. Because the material originates from unconstrained, real-world recordings, it reproduces the conditions encountered when deploying face–voice association systems in practice: acoustic noise, background chatter and music, overlapping speech, and compression artefacts.

Each speaker appears in more than one video, for a total of 557 videos. Face–voice pairs are curated by sampling one video frame per second of active speech, paired with the corresponding audio segment. Table 1 summarises the characteristics of the English–Bengali split.

StatisticBengaliEnglish
# of celebrities100100
# of male celebrities5757
# of female celebrities4343
# of videos301256
# of hours41.629.9
# of utterances32,34122,823
Avg. # of videos per celebrity3.02.6
Avg. # of utterances per celebrity323.4228.2
Avg. length of utterances (in s)4.64.7

Table 1: Characteristics of the MAV-Celeb v4 English–Bengali split.

Audio-visual samples from the English-Bengali split

Figure 2: Audio-visual samples selected from the dataset. The visual data contains variations such as pose, lighting condition, and motion. The left part shows celebrities speaking English; the right block presents data of the same celebrity in Bengali.

Release schedule

The training and development sets are released to participants on the first day of the progress phase. The evaluation input data is released on the first day of the evaluation phase without ground-truth labels, so that the final assessment measures the generalisation capability of systems developed during the progress phase.

Access

The download folder contains the training and development sets, including both the raw data files and the corresponding pre-extracted audio-visual features. The meta file for the training set with gender annotations is distributed separately.

Directory structure

Training set

train_set/
├── features/
│   ├── faces/
│   │   └── train_English_faces.csv      # 6,485 rows × 4,097 cols (4,096-d face embedding + class label)
│   └── voices/
│       └── train_English_voices.csv     # 6,485 rows × 193 cols (192-d voice embedding + class label)
└── train_set/
    ├── train_English.txt                # 6,485 lines, file list with labels
    ├── faces/
    │   └── English/
    │       ├── id001/
    │       │   ├── 00036.jpg
    │       │   ├── 00056.jpg
    │       │   └── ...
    │       ├── id002/
    │       └── ... id070/               # 70 identities, 6,485 .jpg total
    └── voices/
        └── English/
            ├── id001/
            │   ├── 00036.wav
            │   ├── 00056.wav
            │   └── ...
            ├── id002/
            └── ... id070/               # 70 identities, 6,485 .wav total

Development set

dev_set/
├── gender/
│   ├── English_test.txt                 # 982 lines
│   ├── Bangla_test.txt                  # 1,468 lines
│   ├── features/
│   │   ├── English_test_faces.csv       # 982 × 4,096
│   │   ├── English_test_voices.csv      # 982 × 192
│   │   ├── Bangla_test_faces.csv        # 1,468 × 4,096
│   │   └── Bangla_test_voices.csv       # 1,468 × 192
│   ├── English_test/
│   │   ├── faces/    00000.jpg ... (982 files)
│   │   └── voices/   00000.wav ... (982 files)
│   └── Bangla_test/
│       ├── faces/    00000.jpg ... (1,468 files)
│       └── voices/   00000.wav ... (1,468 files)
└── no_gender/
    ├── English_test.txt                 # 1,008 lines
    ├── Bangla_test.txt                  # 1,406 lines
    ├── features/
    │   ├── English_test_faces.csv       # 1,008 × 4,096
    │   ├── English_test_voices.csv      # 1,008 × 192
    │   ├── Bangla_test_faces.csv        # 1,406 × 4,096
    │   └── Bangla_test_voices.csv       # 1,406 × 192
    ├── English_test/
    │   ├── faces/    00000.jpg ... (1,008 files)
    │   └── voices/   00000.wav ... (1,008 files)
    └── Bangla_test/
        ├── faces/    00000.jpg ... (1,406 files)
        └── voices/   00000.wav ... (1,406 files)

BASELINE MODEL & STARTER KIT

To allow participants to benchmark their results, we release a pretrained instance of a competitive multimodal method for the face-voice association task, together with a leaderboard on CodaBench where teams can compare their performance during the progress phase.

The model

The baseline is Fusion and Orthogonal Projection (FOP), a two-branch network taking face and voice embeddings as input (see Fig. 3). Face embeddings are obtained with a popular convolutional network pretrained on a large-scale facial recognition dataset; voice embeddings are obtained with an audio encoding network for speaker recognition, trained on the language available in the training set (i.e. the heard language). The multimodal model then combines the face and voice embeddings and is optimized with a loss function that imposes orthogonality constraints on the multimodal embeddings of different speakers.

Figure 3: Diagram showing the baseline methodology.

Starter kit

Baseline results

Table 2 reports the baseline face-voice association results (EER, %, lower is better) in the standard and gender-constrained settings, for the heard (English) and unheard (Bengali) languages. Performance degrades both under the language shift and in the gender-constrained setting — exactly the gap the challenge invites participants to close.

MethodPhaseConfig.Standard
Eng. test
Standard
Bengali test
Gender-Constrained
Eng. test
Gender-Constrained
Bengali test
Overall Score
FOPDevEng. train32.5438.1232.9944.0136.92
FOPEvalEng. train29.1032.9039.4039.9035.32

Table 2: Cross-modal verification between face and voice across the test configurations of the MAV-Celeb v4 dataset.

EVALUATION METRICS

As commonly done for face-voice association tasks, submissions are evaluated with the Equal Error Rate (EER), computed from the confidence scores a system assigns to every test pair. Within a submission file, a pair with a higher score is interpreted as one the system is more confident about.

Equal Error Rate Lower is better

The operating point at which the false acceptance rate (FAR) — wrongly accepting a non-matching face-voice pair — equals the false rejection rate (FRR), that is, wrongly rejecting a matching one. As for FAR and FRR, a low EER indicates good performance.

The four evaluation conditions

Each system is evaluated on two languages and in two settings, which yields four EERs per system.

Standard
Gender-constrained
Englishheard language
EER1
EER3
Bengaliunheard language
EER2
EER4

Overall score

The score used to decide the challenge winners is the mean of those four values:

Overall Score = ( EER1 + EER2 + EER3 + EER4 ) / 4

Why the Equal Error Rate?

Comparable across systems

EER does not require the confidence scores of different systems to span the same range, so teams are free to calibrate their outputs as they wish.

Independent of a threshold

Unlike accuracy or precision, EER does not depend on a fixed decision threshold. In a real deployment that threshold is tuned to the application: a high threshold gives a low FAR and a high FRR.

Consistent with earlier editions

The same metric was used in the FAME 2024 and FAME 2026 Grand Challenges, which keeps results comparable across editions.

SUBMISSION

Submissions are made through the CodaBench platform, where they are scored automatically against the withheld ground truth. This section describes the test pair files, the score files a system must produce, the archive layout expected by the scorer, and the submission limits of each phase.

Submission platform FLAG 2027 on CodaBench Predictions are evaluated and scored automatically, and results appear on the public leaderboard.
Open the competition

1. Test pair files

Alongside the audio (.wav) and image (.jpg) files, the dataset includes .txt files of the face-voice test pairs. The first entry of each line is the ID of the pair; the remaining two are the local paths of the corresponding audio and of a visual frame extracted from the video:

ljAnhn41 English_test/voices/00000.wav English_test/faces/00000.jpg
neHzLCeC English_test/voices/00001.wav English_test/faces/00001.jpg

The ground-truth files contain the ID of the pair and a 1 or a 0, depending on whether the face and voice correspond to the same speaker. Ground truth is withheld for fair evaluation during the challenge.

2. Score files

For every pair, systems output a confidence score that the face and the voice belong to the same person. Each score file contains one line per pair, with the pair ID followed by the score:

ljAnhn41 1.162691
neHzLCeC 1.235319

Scores are not required to lie in a fixed range: only their ordering matters, since evaluation is based on the Equal Error Rate.

3. Archive layout

Submit a ZIP archive containing one score file per protocol cell, placed in two folders named after the tracks. From within the directory holding the two folders, run:

zip -r submission.zip no_gender gender

The archive is expected to have the following layout:

submission.zip
├── no_gender/
│   ├── sub_score_v4_English_heard.txt
│   └── sub_score_v4_Bangla_unheard.txt
└── gender/
    ├── sub_score_v4_English_heard.txt
    └── sub_score_v4_Bangla_unheard.txt

4. Submission limits

Progress phase 150 submissions in total, at most 15 per day
Evaluation phase 15 submissions in total

RULES FOR SYSTEM DEVELOPMENT

Since FLAG 2027 aims at analyzing whether face-voice association capabilities translate across language and gender constraints, the following rules apply:

Research ethics and responsible use

Participants agree to the following:

Teams that break these rules will be disqualified.

REGISTRATION

Participation in the FLAG 2027 Challenge is open to teams from academia and industry, and there is no registration fee. One team member registers on behalf of the whole team using the form below. Registration closes on 15 October 2026.

1
Register your team

Submit the registration form with your team name, affiliation and the contact details of the corresponding member.

2
Access the data

Download the training and development sets, the pre-extracted features and the baseline, and create an account on CodaBench.

3
Submit your systems

Upload your score files to CodaBench during the progress and evaluation phases, then send a 2-page system description and a link to your code.

Registration is open Join the FLAG 2027 Challenge
Register your team

Questions about the challenge, the data or the submission process are welcome at mavceleb@gmail.com.

PRIZES

The FLAG 2027 Challenge has been accepted for support under the IEEE Signal Processing Society Challenge Program. The prize pool is distributed among the top five ranked teams.

$5,000 Total prize pool
$1,200 2nd place
$2,000 1st place
$800 3rd place
$600 4th place
$400 5th place

ORGANIZERS

Marta Moscati

Marta Moscati

Institute of Computational Perception, Johannes Kepler University Linz, Austria

Swapnil Khandoker

Swapnil Khandoker

Institute of Computational Perception, Johannes Kepler University Linz, Austria

Muhammad Saad Saeed

Muhammad Saad Saeed

University of Michigan, USA

Shah Nawaz

Shah Nawaz

Institute of Computational Perception, Johannes Kepler University Linz, Austria

Fatima Noor

University of Engineering and Technology Taxila, Pakistan

Rohan Kumar Das

Rohan Kumar Das

Fortemedia Singapore, Singapore

Mubashir Noman

Mubashir Noman

Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates

Junaid Mir

Junaid Mir

University of Engineering and Technology Taxila, Pakistan

Muhammad Haroon Yousaf

Muhammad Haroon Yousaf

University of Engineering and Technology Taxila, Pakistan

Khalid Malik

Khalid Malik

University of Michigan, USA

Markus Schedl

Markus Schedl

Institute of Computational Perception, Johannes Kepler University Linz, Austria | Linz Institute of Technology, Austria

COLLABORATING INSTITUTES