Diagnostic Index

Benchmark scenarios

Each scenario groups related trial sets around a deployment condition and provides a local ranking for diagnosing system behavior.

11 scenarios 26 trial sets 18 ranked systems
Scenario Rankings

Evaluation coverage

Open a scenario to compare systems using the corresponding group-level EER, minDCF, and trial-set evidence.

Scenario Evaluation Focus Trial Sets Included Trial Sets Current Leader
Accent / Dialect Robustness
Accent/dialect
Accent and dialect variation 2
GLOBE 3D-Speaker Dialect
wespeaker/w2vbert2
Aging Robustness
Aging
Speaker aging and time gaps 2
VoxKnesset VoxPopuli Aging
wespeaker/samresnet100
Channel / Device Robustness
Channel/device
Device and channel variation 2
3D-Speaker Device FFSVC 2022 Cross-Channel
iic/eres2net-large-3dspeaker
Cross-Lingual Robustness
Cross-lingual
Language mismatch 1
TidyVoiceX2-ASV
wespeaker/w2vbert2
Distance Robustness
Distance
Distance mismatch 4
3D-Speaker Distance AliMeeting Near/Far CHiME-6 Domestic Far-Field +1
iic/eres2net-large-3dspeaker
Genre-Shift Robustness
Genre shift
Source-genre variation 1
CN-Celeb Genre
iic/eres2netv2-zh
In-The-Wild Robustness
In-the-wild
Open-domain media speech 4
CN-Celeb VoxCeleb1-O VoxCeleb1-E +1
wespeaker/w2vbert2
Noise / Reverb Robustness
Noise/reverb
Noise and reverberation 1
VOiCES Noise/Reverb
wespeaker/w2vbert2
Overlap Robustness
Overlap
Overlapping speakers 2
AliMeeting Overlap CHiME-6 Overlap
wespeaker/w2vbert2
Short-Duration Robustness
Short-duration
Short-duration speech 4
HI-MIA GSC Short CN-Celeb Short +1
wespeaker/w2vbert2
Speaking-Style Robustness
Speaking style
Speaking style shift 3
ESD Whisper40 Whisper Lombard Grid Lombard
wespeaker/w2vbert2