INTRODUCTION
Recent years have seen a surge in finding association between faces and voices within a cross-modal
biometric application along with speaker recognition. Inspired from this, we introduce a challenging
task in establishing association between faces and voices across multiple languages spoken by the same
set of persons. The aim of this paper is to answer two closely related questions: “Is face-voice
association language independent?” and “Can a speaker be recognized irrespective of the spoken language?”.
These two questions are important to understand effectiveness and to boost development of multilingual biometric
systems. To answer these, we collected a Multilingual Audio-Visual dataset, containing human speech clips of 154
identities with 3 language annotations extracted from various videos uploaded online. Extensive experiments
on the two splits of the proposed dataset have been performed to investigate and answer these novel research
questions that clearly point out the relevance of the multilingual problem.
DATASET
The data is obtained from YouTube videos, consisting of celebrity interviews along with talk shows, and television debates. The visual data spans over a vast range of variations including poses, motion blur, background clutter, video quality, occlusions and lighting conditions. Moreover, most videos contain real-world noise like background chatter, music, over-lapping speech, and compression artifacts, resulting into a challenging dataset to evaluate multimedia systems.
The dataset is available on the following links:
PUBLICATIONS
Dataset Paper
Cross-modal Speaker Verification and Recognition: A Multilingual PerspectiveAuthors: Nawaz, Shah and Saeed, Muhammad Saad and Morerio, Pietro and Mahmood, Arif and Gallo, Ignazio and Yousaf, Muhammad Haroon and Del Bue, Alessio
Baseline Paper
Fusion and Orthogonal Projection for Improved Face-Voice AssociationAuthors: Saeed, Muhammad Saad and Khan, Muhammad Haris and Nawaz, Shah and Yousaf, Muhammad Haroon and Del Bue, Alessio
CHALLENGE
The FLAG Challenge 2027 — Face-voice Association across LAnguages and Gender — studies face-voice association along two dimensions that are usually neglected under controlled experimental settings: the language spoken by a person, and their demographic traits. The challenge runs two evaluation tracks. The language impact track asks whether a voice recorded in one language can still be associated with the corresponding speaker's face when the model was trained on another language. The gender impact track uses gender-controlled verification, where negative pairs are restricted to speakers of the same gender as the positive pair, so that models cannot fall back on gender as a shortcut for identity. The aim is to foster multilingual and gender-aware methods that learn genuinely identity-specific cross-modal associations.
Click to see more details!