System Detail
wespeaker/ res152-voxceleb
wespeaker_resnet152
This page shows the global rank, scenario ranks, and trial-set results for this submission.
Global Position
#4/18
Public ranking position among reviewed full-core submissions.
Scenario-Macro EER
11.2275
Scenario-Macro minDCF
0.442748
Coverage
26/26
0 scenarios ranked #1
1 scenario in top 3
Public Listing
Ranked Publicly
This full-core result is approved and included in the public ranking.
Model Provenance
Source and architecture
- Displayed name:
wespeaker/res152-voxceleb - Source: WeSpeaker
- Architecture: ResNet
- Submitted alias:
wespeaker-voxceleb-resnet152-LM - Created at:
2026-05-29T14:53:46.991718+00:00
Training And Links
Training data and references
- Training data: VoxCeleb2 dev, 5,994 speakers
- Training setup: ResNet152-TSTP-emb256 r-vector, ArcMargin, 150-epoch VoxCeleb2 recipe with speed perturbation and MUSAN/RIRS augmentation.
- Submission mode:
full-core - Paper: Deep Residual Learning for Image Recognition
Higher-Ranked Scenarios
Scenario ranks above the global position
Show 3 supporting trial sets
- Lombard Grid Lombard: #3 on its trial-set ranking.
- ESD: #3 on its trial-set ranking.
- Whisper40 Whisper: #3 on its trial-set ranking.
Lower-Ranked Scenarios
Scenario ranks below the global position
Show 3 supporting trial sets
- 3D-Speaker Device: #8 on its trial-set ranking.
- AliMeeting Overlap: #8 on its trial-set ranking.
- CN-Celeb Genre: #7 on its trial-set ranking.
Scenario Rankings
Full ranking across scenarios
This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.
| Scenario | Rank | Lens | Evidence | Score |
|---|---|---|---|---|
|
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
|
#2/18
Top 3
|
Speaking style shift
1.3434 EER from leader
|
4.1329
0.268357 minDCF
|
|
|
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
|
#4/18
Competitive
|
Open-domain media speech
0.9177 EER from leader
|
6.1011
0.224787 minDCF
|
|
|
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
|
#4/18
Competitive
|
Accent and dialect variation
1.0902 EER from leader
|
21.5254
0.675759 minDCF
|
|
|
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
|
#5/18
Competitive
|
Speaker aging and time gaps
0.6780 EER from leader
|
2.7476
0.111015 minDCF
|
|
|
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
|
#5/18
Competitive
|
Language mismatch
0.3592 EER from leader
|
4.8269
0.305505 minDCF
|
|
|
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
|
#5/18
Competitive
|
Noise and reverberation
2.6800 EER from leader
|
5.2600
0.154640 minDCF
|
|
|
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
|
#5/18
Competitive
|
Overlapping speakers
2.8169 EER from leader
|
12.9553
0.565796 minDCF
|
|
|
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
|
#5/18
Competitive
|
Short-duration speech
0.9914 EER from leader
|
13.5058
0.545781 minDCF
|
|
|
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
|
#7/18
Competitive
|
Distance mismatch
2.6914 EER from leader
|
11.9552
0.523450 minDCF
|
|
|
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
|
#7/18
Competitive
|
Device and channel variation
7.9353 EER from leader
|
14.2456
0.688961 minDCF
|
|
|
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
|
#7/18
Competitive
|
Source-genre variation
11.9747 EER from leader
|
26.2472
0.806172 minDCF
|
Variant Compare
Compare ResNet variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
Show
Compare ResNet variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
| Variant | Source | Training Data | Training Setup | Global Rank | Macro EER | Macro minDCF | Open |
|---|---|---|---|---|---|---|---|
|
wespeaker/res152-voxceleb
Current
|
WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ResNet152-TSTP-emb256 r-vector, ArcMargin, 150-epoch VoxCeleb2 recipe with speed perturbation and MUSAN/RIRS augmentation. |
#4
ranked
|
11.2275 | 0.442748 | Open |
| wespeaker/res293-voxceleb | WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ResNet293-TSTP-emb256 r-vector, large-margin fine-tuned; official card reports 28.62M parameters and 28.10G FLOPs. |
#5
ranked
|
11.3130 | 0.439607 | Open |
| wespeaker/res34-voxceleb | WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ResNet34-TSTP-emb256 r-vector, large-margin fine-tuned; official card reports 6.63M parameters and 4.55G FLOPs. |
#7
ranked
|
12.0516 | 0.477054 | Open |
| wespeaker/res34-cnceleb | WeSpeaker | CN-Celeb train | ResNet34 r-vector with TSTP pooling and large-margin fine-tuning on the CN-Celeb WeSpeaker recipe. |
#14
ranked
|
14.4910 | 0.607964 | Open |
Trial-Set Results
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
Show
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
| Trial Set | Rank | Scenarios | Trials | Score |
|---|---|---|---|---|
|
TidyVoiceX2-ASV
TidyVoiceX2-ASV
|
#5/18
|
Cross-lingual
|
200000 |
4.8269
0.305505 minDCF
|
|
HI-MIA
HI-MIA
|
#4/18
|
Short-duration
|
660000 |
5.8453
0.284905 minDCF
|
|
GSC Short
Google Speech Commands
|
#5/18
|
Short-duration
|
220000 |
10.8860
0.601260 minDCF
|
|
GLOBE
GLOBE
|
#5/18
|
Accent/dialect
|
56848 |
25.0967
0.533224 minDCF
|
|
CN-Celeb
CN-Celeb
|
#6/18
|
In-the-wild
|
3484292 |
11.3672
0.399395 minDCF
|
|
CN-Celeb Genre
CN-Celeb
|
#7/18
|
Genre shift
|
440000 |
26.2472
0.806172 minDCF
|
|
CN-Celeb Short
CN-Celeb
|
#5/18
|
Short-duration
|
546964 |
24.9437
0.851112 minDCF
|
|
VoxCeleb1-O
VoxCeleb1
|
#4/18
|
In-the-wild
|
37611 |
0.4998
0.035524 minDCF
|
|
VoxCeleb1-E
VoxCeleb1
|
#6/18
|
In-the-wild
|
579818 |
0.7174
0.043120 minDCF
|
|
VoxCeleb1-H
VoxCeleb1
|
#5/18
|
In-the-wild
|
550894 |
1.2879
0.071894 minDCF
|
|
VoxCeleb Short
VoxCeleb1
|
#4/18
|
Short-duration
|
394724 |
12.3481
0.445848 minDCF
|
|
3D-Speaker Device
3D-Speaker
|
#8/18
|
Channel/device
|
180000 |
18.9800
0.832393 minDCF
|
|
3D-Speaker Distance
3D-Speaker
|
#4/18
|
Distance
|
175163 |
17.5655
0.786467 minDCF
|
|
3D-Speaker Dialect
3D-Speaker
|
#5/18
|
Accent/dialect
|
180000 |
17.9540
0.818293 minDCF
|
|
FFSVC 2022 Cross-Channel
FFSVC 2022
|
#7/18
|
Channel/device
|
72000 |
9.5111
0.545528 minDCF
|
|
FFSVC 2022 Cross-Domain
FFSVC 2022
|
#7/18
|
Distance
|
66546 |
9.6636
0.556279 minDCF
|
|
Whisper40 Whisper
Whisper40
|
#3/18
|
Speaking style
|
17600 |
9.2500
0.567375 minDCF
|
|
Lombard Grid Lombard
Lombard Grid
|
#3/18
|
Speaking style
|
29524 |
0.0745
0.011736 minDCF
|
|
VOiCES Noise/Reverb
VOiCES
|
#5/18
|
Noise/reverb
|
55000 |
5.2600
0.154640 minDCF
|
|
CHiME-6 Domestic Far-Field
CHiME-6
|
#5/18
|
Distance
|
36487 |
12.0320
0.480796 minDCF
|
|
CHiME-6 Overlap
CHiME-6
|
#5/18
|
Overlap
|
39600 |
16.5639
0.642111 minDCF
|
|
ESD
ESD
|
#3/18
|
Speaking style
|
437408 |
3.0742
0.225959 minDCF
|
|
AliMeeting Near/Far
AliMeeting
|
#4/18
|
Distance
|
220000 |
8.5600
0.270260 minDCF
|
|
AliMeeting Overlap
AliMeeting
|
#8/18
|
Overlap
|
165000 |
9.3467
0.489480 minDCF
|
|
VoxKnesset
VoxKnesset
|
#5/18
|
Aging
|
158312 |
3.7467
0.148597 minDCF
|
|
VoxPopuli Aging
VoxPopuli
|
#5/18
|
Aging
|
146575 |
1.7486
0.073433 minDCF
|