INTRODUCTION

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association between faces and voices across multiple languages spoken by the same set of persons. The aim of this paper is to answer two closely related questions: “Is face-voice association language independent?” and “Can a speaker be recognized irrespective of the spoken language?”.

These two questions are important to understand effectiveness and to boost development of multilingual biometric systems. To answer these, we collected a Multilingual Audio-Visual dataset, containing human speech clips of 154 identities with 3 language annotations extracted from various videos uploaded online. Extensive experiments on the two splits of the proposed dataset have been performed to investigate and answer these novel research questions that clearly point out the relevance of the multilingual problem.

DATASET

The data is obtained from YouTube videos, consisting of celebrity interviews along with talk shows, and television debates. The visual data spans over a vast range of variations including poses, motion blur, background clutter, video quality, occlusions and lighting conditions. Moreover, most videos contain real-world noise like background chatter, music, over-lapping speech, and compression artifacts, resulting into a challenging dataset to evaluate multimedia systems.

The dataset is available on the following links:





PUBLICATIONS

Dataset Paper

Cross-modal Speaker Verification and Recognition: A Multilingual Perspective

Authors: Nawaz, Shah and Saeed, Muhammad Saad and Morerio, Pietro and Mahmood, Arif and Gallo, Ignazio and Yousaf, Muhammad Haroon and Del Bue, Alessio

Baseline Paper

Fusion and Orthogonal Projection for Improved Face-Voice Association

Authors: Saeed, Muhammad Saad and Khan, Muhammad Haris and Nawaz, Shah and Yousaf, Muhammad Haroon and Del Bue, Alessio

CHALLENGE

The FLAG Challenge 2027 — Face-voice Association across LAnguages and Gender — studies face-voice association along two dimensions that are usually neglected under controlled experimental settings: the language spoken by a person, and their demographic traits. The challenge runs two evaluation tracks. The language impact track asks whether a voice recorded in one language can still be associated with the corresponding speaker's face when the model was trained on another language. The gender impact track uses gender-controlled verification, where negative pairs are restricted to speakers of the same gender as the positive pair, so that models cannot fall back on gender as a shortcut for identity. The aim is to foster multilingual and gender-aware methods that learn genuinely identity-specific cross-modal associations.


Click to see more details!

FLAG Challenge 2027



FAME Challenge 2026



FAME Challenge 2024