System Detail
wespeaker/ samresnet100
vb2-sam100
This page shows the global rank, scenario ranks, and trial-set results for this submission.
Global Position
#3/18
Public ranking position among reviewed full-core submissions.
Scenario-Macro EER
10.5780
Scenario-Macro minDCF
0.390239
Coverage
26/26
1 scenario ranked #1
6 scenarios in top 3
Public Listing
Ranked Publicly
This full-core result is approved and included in the public ranking.
Model Provenance
Source and architecture
- Displayed name:
wespeaker/samresnet100 - Source: WeSpeaker
- Architecture: SAM-ResNet100
- Submitted alias:
voxblink2_samresnet100_ft - Created at:
2026-05-29T14:22:47.375445+00:00
Training And Links
Training data and references
- Training data: VoxBlink2 (pretrain) + VoxCeleb2 (finetune)
- Training setup: SimAM-ResNet100 ASP release fine-tuned on VoxCeleb2 with ArcMargin, speed perturbation, and MUSAN/RIRS augmentation.
- Submission mode:
full-core - Paper: Deep Residual Learning for Image Recognition
Higher-Ranked Scenarios
Scenario ranks above the global position
Show 3 supporting trial sets
- Lombard Grid Lombard: #1 on its trial-set ranking.
- VoxKnesset: #1 on its trial-set ranking.
- VoxCeleb1-O: #2 on its trial-set ranking.
Lower-Ranked Scenarios
Scenario ranks below the global position
Show 3 supporting trial sets
- 3D-Speaker Distance: #8 on its trial-set ranking.
- 3D-Speaker Dialect: #7 on its trial-set ranking.
- CN-Celeb Short: #6 on its trial-set ranking.
Scenario Rankings
Full ranking across scenarios
This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.
| Scenario | Rank | Lens | Evidence | Score |
|---|---|---|---|---|
|
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
|
#1/18
Leader
|
Speaker aging and time gaps
0.0000 EER from leader
|
2.0696
0.068782 minDCF
|
|
|
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
|
#2/18
Top 3
|
Noise and reverberation
0.4000 EER from leader
|
2.9800
0.098660 minDCF
|
|
|
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
|
#2/18
Top 3
|
Language mismatch
0.2501 EER from leader
|
4.7178
0.362506 minDCF
|
|
|
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
|
#2/18
Top 3
|
Open-domain media speech
0.0653 EER from leader
|
5.2487
0.170274 minDCF
|
|
|
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
|
#3/18
Top 3
|
Speaking style shift
1.5147 EER from leader
|
4.3042
0.238095 minDCF
|
|
|
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
|
#3/18
Top 3
|
Overlapping speakers
0.9529 EER from leader
|
11.0912
0.502019 minDCF
|
|
|
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
|
#4/18
Competitive
|
Distance mismatch
1.9041 EER from leader
|
11.1679
0.421520 minDCF
|
|
|
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
|
#4/18
Competitive
|
Short-duration speech
0.9335 EER from leader
|
13.4479
0.497755 minDCF
|
|
|
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
|
#4/18
Competitive
|
Source-genre variation
10.8475 EER from leader
|
25.1200
0.674765 minDCF
|
|
|
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
|
#6/18
Competitive
|
Device and channel variation
7.7067 EER from leader
|
14.0169
0.591790 minDCF
|
|
|
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
|
#7/18
Competitive
|
Accent and dialect variation
1.7588 EER from leader
|
22.1939
0.666463 minDCF
|
Trial-Set Results
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
Show
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
| Trial Set | Rank | Scenarios | Trials | Score |
|---|---|---|---|---|
|
TidyVoiceX2-ASV
TidyVoiceX2-ASV
|
#2/18
|
Cross-lingual
|
200000 |
4.7178
0.362506 minDCF
|
|
HI-MIA
HI-MIA
|
#6/18
|
Short-duration
|
660000 |
5.9033
0.273133 minDCF
|
|
GSC Short
Google Speech Commands
|
#6/18
|
Short-duration
|
220000 |
11.2450
0.611560 minDCF
|
|
GLOBE
GLOBE
|
#2/18
|
Accent/dialect
|
56848 |
24.8065
0.531753 minDCF
|
|
CN-Celeb
CN-Celeb
|
#5/18
|
In-the-wild
|
3484292 |
9.9803
0.311502 minDCF
|
|
CN-Celeb Genre
CN-Celeb
|
#4/18
|
Genre shift
|
440000 |
25.1200
0.674765 minDCF
|
|
CN-Celeb Short
CN-Celeb
|
#6/18
|
Short-duration
|
546964 |
25.4042
0.714774 minDCF
|
|
VoxCeleb1-O
VoxCeleb1
|
#2/18
|
In-the-wild
|
37611 |
0.2287
0.011966 minDCF
|
|
VoxCeleb1-E
VoxCeleb1
|
#3/18
|
In-the-wild
|
579818 |
0.4564
0.026098 minDCF
|
|
VoxCeleb1-H
VoxCeleb1
|
#2/18
|
In-the-wild
|
550894 |
0.8665
0.049074 minDCF
|
|
VoxCeleb Short
VoxCeleb1
|
#3/18
|
Short-duration
|
394724 |
11.2390
0.391553 minDCF
|
|
3D-Speaker Device
3D-Speaker
|
#5/18
|
Channel/device
|
180000 |
18.9367
0.711080 minDCF
|
|
3D-Speaker Distance
3D-Speaker
|
#8/18
|
Distance
|
175163 |
18.4835
0.669019 minDCF
|
|
3D-Speaker Dialect
3D-Speaker
|
#7/18
|
Accent/dialect
|
180000 |
19.5813
0.801173 minDCF
|
|
FFSVC 2022 Cross-Channel
FFSVC 2022
|
#6/18
|
Channel/device
|
72000 |
9.0972
0.472500 minDCF
|
|
FFSVC 2022 Cross-Domain
FFSVC 2022
|
#6/18
|
Distance
|
66546 |
9.2164
0.478007 minDCF
|
|
Whisper40 Whisper
Whisper40
|
#6/18
|
Speaking style
|
17600 |
10.1250
0.485938 minDCF
|
|
Lombard Grid Lombard
Lombard Grid
|
#1/18
|
Speaking style
|
29524 |
0.0373
0.004657 minDCF
|
|
VOiCES Noise/Reverb
VOiCES
|
#2/18
|
Noise/reverb
|
55000 |
2.9800
0.098660 minDCF
|
|
CHiME-6 Domestic Far-Field
CHiME-6
|
#2/18
|
Distance
|
36487 |
8.4112
0.319174 minDCF
|
|
CHiME-6 Overlap
CHiME-6
|
#2/18
|
Overlap
|
39600 |
13.7778
0.612972 minDCF
|
|
ESD
ESD
|
#2/18
|
Speaking style
|
437408 |
2.7503
0.223690 minDCF
|
|
AliMeeting Near/Far
AliMeeting
|
#5/18
|
Distance
|
220000 |
8.5605
0.219880 minDCF
|
|
AliMeeting Overlap
AliMeeting
|
#5/18
|
Overlap
|
165000 |
8.4047
0.391067 minDCF
|
|
VoxKnesset
VoxKnesset
|
#1/18
|
Aging
|
158312 |
2.5933
0.101669 minDCF
|
|
VoxPopuli Aging
VoxPopuli
|
#3/18
|
Aging
|
146575 |
1.5460
0.035895 minDCF
|