OpenSVBench Scenario-driven SV leaderboard
System Detail

wespeaker/ w2vbert2

vb2-w2vbert

ranked review: approved WeSpeaker w2v-BERT 2.0

This page shows the global rank, scenario ranks, and trial-set results for this submission.

Global Position

#1/18

Public ranking position among reviewed full-core submissions.

Scenario-Macro EER
9.9604
Scenario-Macro minDCF
0.360767
Coverage

26/26

7 scenarios ranked #1 9 scenarios in top 3
Public Listing

Ranked Publicly

This full-core result is approved and included in the public ranking.

Model Provenance

Source and architecture

  • Displayed name: wespeaker/w2vbert2
  • Source: WeSpeaker
  • Architecture: w2v-BERT 2.0
  • Submitted alias: voxceleb_voxblink2_w2v_bert2_lora_adapterMFA_lm
  • Created at: 2026-05-29T14:24:45.970729+00:00
Training And Links

Training data and references

Higher-Ranked Scenarios

Scenario ranks above the global position

Show 3 supporting trial sets
  • VoxCeleb1-O: #1 on its trial-set ranking.
  • VoxCeleb1-E: #1 on its trial-set ranking.
  • VoxCeleb1-H: #1 on its trial-set ranking.
Lower-Ranked Scenarios

Scenario ranks below the global position

Show 3 supporting trial sets
  • FFSVC 2022 Cross-Domain: #10 on its trial-set ranking.
  • FFSVC 2022 Cross-Channel: #10 on its trial-set ranking.
  • CN-Celeb Short: #8 on its trial-set ranking.
Scenario Rankings

Full ranking across scenarios

This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.

Scenario Rank Lens Evidence Score
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
#1/18 Leader
Noise and reverberation 0.0000 EER from leader
2.5800
0.074040 minDCF
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
#1/18 Leader
Speaking style shift 0.0000 EER from leader
2.7895
0.157223 minDCF
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
#1/18 Leader
Language mismatch 0.0000 EER from leader
4.4676
0.330255 minDCF
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
#1/18 Leader
Open-domain media speech 0.0000 EER from leader
5.1834
0.162046 minDCF
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
#1/18 Leader
Overlapping speakers 0.0000 EER from leader
10.1383
0.467346 minDCF
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
#1/18 Leader
Short-duration speech 0.0000 EER from leader
12.5144
0.441680 minDCF
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
#1/18 Leader
Accent and dialect variation 0.0000 EER from leader
20.4352
0.641166 minDCF
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
#2/18 Top 3
Speaker aging and time gaps 0.0150 EER from leader
2.0846
0.052897 minDCF
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
#2/18 Top 3
Distance mismatch 0.9246 EER from leader
10.1885
0.403238 minDCF
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
#4/18 Competitive
Device and channel variation 6.9525 EER from leader
13.2628
0.581970 minDCF
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
#6/18 Competitive
Source-genre variation 11.6475 EER from leader
25.9200
0.656577 minDCF
Trial-Set Results

Raw ranking by trial set

Expand this section to inspect the detailed trial-set rankings.

Show
Trial Set Rank Scenarios Trials Score
TidyVoiceX2-ASV
TidyVoiceX2-ASV
#1/18
Cross-lingual
200000
4.4676
0.330255 minDCF
HI-MIA
HI-MIA
#8/18
Short-duration
660000
7.0400
0.349597 minDCF
GSC Short
Google Speech Commands
#2/18
Short-duration
220000
6.0250
0.388025 minDCF
GLOBE
GLOBE
#1/18
Accent/dialect
56848
24.5937
0.512152 minDCF
CN-Celeb
CN-Celeb
#4/18
In-the-wild
3484292
9.9690
0.300308 minDCF
CN-Celeb Genre
CN-Celeb
#6/18
Genre shift
440000
25.9200
0.656577 minDCF
CN-Celeb Short
CN-Celeb
#8/18
Short-duration
546964
25.9130
0.658933 minDCF
VoxCeleb1-O
VoxCeleb1
#1/18
In-the-wild
37611
0.1489
0.010954 minDCF
VoxCeleb1-E
VoxCeleb1
#1/18
In-the-wild
579818
0.3146
0.018750 minDCF
VoxCeleb1-H
VoxCeleb1
#1/18
In-the-wild
550894
0.7298
0.041650 minDCF
VoxCeleb Short
VoxCeleb1
#2/18
Short-duration
394724
11.0796
0.370165 minDCF
3D-Speaker Device
3D-Speaker
#2/18
Channel/device
180000
15.3867
0.634773 minDCF
3D-Speaker Distance
3D-Speaker
#3/18
Distance
175163
16.0613
0.619447 minDCF
3D-Speaker Dialect
3D-Speaker
#2/18
Accent/dialect
180000
16.2767
0.770180 minDCF
FFSVC 2022 Cross-Channel
FFSVC 2022
#10/18
Channel/device
72000
11.1389
0.529167 minDCF
FFSVC 2022 Cross-Domain
FFSVC 2022
#10/18
Distance
66546
11.3097
0.539530 minDCF
Whisper40 Whisper
Whisper40
#1/18
Speaking style
17600
5.8125
0.291250 minDCF
Lombard Grid Lombard
Lombard Grid
#2/18
Speaking style
29524
0.0559
0.005104 minDCF
VOiCES Noise/Reverb
VOiCES
#1/18
Noise/reverb
55000
2.5800
0.074040 minDCF
CHiME-6 Domestic Far-Field
CHiME-6
#1/18
Distance
36487
6.6928
0.260295 minDCF
CHiME-6 Overlap
CHiME-6
#1/18
Overlap
39600
12.3500
0.603472 minDCF
ESD
ESD
#1/18
Speaking style
437408
2.5000
0.175313 minDCF
AliMeeting Near/Far
AliMeeting
#3/18
Distance
220000
6.6900
0.193680 minDCF
AliMeeting Overlap
AliMeeting
#4/18
Overlap
165000
7.9267
0.331220 minDCF
VoxKnesset
VoxKnesset
#2/18
Aging
158312
2.6533
0.081689 minDCF
VoxPopuli Aging
VoxPopuli
#1/18
Aging
146575
1.5159
0.024105 minDCF