System Detail
iic/ eres2net-large-3dspeaker
eres2net_large_3dspeaker
This page shows the global rank, scenario ranks, and trial-set results for this submission.
Global Position
#12/18
Public ranking position among reviewed full-core submissions.
Scenario-Macro EER
14.1682
Scenario-Macro minDCF
0.593516
Coverage
26/26
2 scenarios ranked #1
3 scenarios in top 3
Public Listing
Ranked Publicly
This full-core result is approved and included in the public ranking.
Model Provenance
Source and architecture
- Displayed name:
iic/eres2net-large-3dspeaker - Source: IIC / ModelScope
- Architecture: ERes2Net
- Submitted alias:
speech_eres2net_large_sv_zh-cn_3dspeaker_16k - Created at:
2026-05-29T14:58:14.607926+00:00
Training And Links
Training data and references
- Training data: 3D-Speaker open dataset, about 10k Chinese speakers
- Training setup: Official 16 kHz Chinese ERes2Net-Large release; the local model card reports 18.3M parameters and training on 3D-Speaker.
- Submission mode:
full-core - Paper: An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
This checkpoint is trained on 3D-Speaker, which is also one of the OpenSVBench source datasets; interpret 3D-Speaker-derived trial sets with that overlap in mind.
Higher-Ranked Scenarios
Scenario ranks above the global position
Show 3 supporting trial sets
- HI-MIA: #1 on its trial-set ranking.
- AliMeeting Near/Far: #1 on its trial-set ranking.
- AliMeeting Overlap: #1 on its trial-set ranking.
Lower-Ranked Scenarios
Scenario ranks below the global position
Show 3 supporting trial sets
- GLOBE: #18 on its trial-set ranking.
- VoxPopuli Aging: #18 on its trial-set ranking.
- VoxCeleb1-H: #18 on its trial-set ranking.
Scenario Rankings
Full ranking across scenarios
This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.
| Scenario | Rank | Lens | Evidence | Score |
|---|---|---|---|---|
|
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
|
#1/18
Leader
|
Device and channel variation
0.0000 EER from leader
|
6.3103
0.401249 minDCF
|
|
|
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
|
#1/18
Leader
|
Distance mismatch
0.0000 EER from leader
|
9.2638
0.454491 minDCF
|
|
|
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
|
#2/18
Top 3
|
Accent and dialect variation
0.0095 EER from leader
|
20.4447
0.733333 minDCF
|
|
|
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
|
#9/18
Competitive
|
Overlapping speakers
4.2283 EER from leader
|
14.3667
0.552132 minDCF
|
|
|
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
|
#11/18
Needs work
|
Source-genre variation
14.2125 EER from leader
|
28.4850
0.888960 minDCF
|
|
|
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
|
#12/18
Needs work
|
Noise and reverberation
6.1240 EER from leader
|
8.7040
0.382320 minDCF
|
|
|
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
|
#14/18
Needs work
|
Short-duration speech
5.0564 EER from leader
|
17.5708
0.654520 minDCF
|
|
|
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
|
#15/18
Needs work
|
Language mismatch
4.0860 EER from leader
|
8.5536
0.499376 minDCF
|
|
|
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
|
#15/18
Needs work
|
Speaking style shift
5.8963 EER from leader
|
8.6857
0.465867 minDCF
|
|
|
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
|
#15/18
Needs work
|
Open-domain media speech
9.0729 EER from leader
|
14.2563
0.581636 minDCF
|
|
|
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
|
#17/18
Needs work
|
Speaker aging and time gaps
17.1397 EER from leader
|
19.2094
0.914797 minDCF
|
Variant Compare
Compare ERes2Net variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
Show
Compare ERes2Net variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
| Variant | Source | Training Data | Training Setup | Global Rank | Macro EER | Macro minDCF | Open |
|---|---|---|---|---|---|---|---|
| iic/eres2net-en | IIC | VoxCeleb2 development set, 5,994 speakers | Official 16 kHz English ERes2Net release. |
#8
ranked
|
12.0955 | 0.472802 | Open |
|
iic/eres2net-large-3dspeaker
Current
|
IIC / ModelScope | 3D-Speaker open dataset, about 10k Chinese speakers | Official 16 kHz Chinese ERes2Net-Large release; the local model card reports 18.3M parameters and training on 3D-Speaker. |
#12
ranked
|
14.1682 | 0.593516 | Open |
Trial-Set Results
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
Show
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
| Trial Set | Rank | Scenarios | Trials | Score |
|---|---|---|---|---|
|
TidyVoiceX2-ASV
TidyVoiceX2-ASV
|
#15/18
|
Cross-lingual
|
200000 |
8.5536
0.499376 minDCF
|
|
HI-MIA
HI-MIA
|
#1/18
|
Short-duration
|
660000 |
2.8963
0.178460 minDCF
|
|
GSC Short
Google Speech Commands
|
#11/18
|
Short-duration
|
220000 |
13.3500
0.658090 minDCF
|
|
GLOBE
GLOBE
|
#18/18
|
Accent/dialect
|
56848 |
28.8893
0.801800 minDCF
|
|
CN-Celeb
CN-Celeb
|
#11/18
|
In-the-wild
|
3484292 |
14.0524
0.486731 minDCF
|
|
CN-Celeb Genre
CN-Celeb
|
#11/18
|
Genre shift
|
440000 |
28.4850
0.888960 minDCF
|
|
CN-Celeb Short
CN-Celeb
|
#10/18
|
Short-duration
|
546964 |
26.8023
0.858622 minDCF
|
|
VoxCeleb1-O
VoxCeleb1
|
#18/18
|
In-the-wild
|
37611 |
13.0837
0.638176 minDCF
|
|
VoxCeleb1-E
VoxCeleb1
|
#18/18
|
In-the-wild
|
579818 |
13.1915
0.655802 minDCF
|
|
VoxCeleb1-H
VoxCeleb1
|
#18/18
|
In-the-wild
|
550894 |
17.1053
0.735645 minDCF
|
|
VoxCeleb Short
VoxCeleb1
|
#17/18
|
Short-duration
|
394724 |
27.2344
0.922910 minDCF
|
|
3D-Speaker Device
3D-Speaker
|
#1/18
|
Channel/device
|
180000 |
6.8733
0.451193 minDCF
|
|
3D-Speaker Distance
3D-Speaker
|
#1/18
|
Distance
|
175163 |
10.3440
0.537387 minDCF
|
|
3D-Speaker Dialect
3D-Speaker
|
#1/18
|
Accent/dialect
|
180000 |
12.0000
0.664867 minDCF
|
|
FFSVC 2022 Cross-Channel
FFSVC 2022
|
#1/18
|
Channel/device
|
72000 |
5.7472
0.351306 minDCF
|
|
FFSVC 2022 Cross-Domain
FFSVC 2022
|
#1/18
|
Distance
|
66546 |
5.8125
0.357548 minDCF
|
|
Whisper40 Whisper
Whisper40
|
#13/18
|
Speaking style
|
17600 |
14.8750
0.868437 minDCF
|
|
Lombard Grid Lombard
Lombard Grid
|
#13/18
|
Speaking style
|
29524 |
1.7139
0.100335 minDCF
|
|
VOiCES Noise/Reverb
VOiCES
|
#12/18
|
Noise/reverb
|
55000 |
8.7040
0.382320 minDCF
|
|
CHiME-6 Domestic Far-Field
CHiME-6
|
#12/18
|
Distance
|
36487 |
17.2988
0.733554 minDCF
|
|
CHiME-6 Overlap
CHiME-6
|
#14/18
|
Overlap
|
39600 |
23.8333
0.823583 minDCF
|
|
ESD
ESD
|
#14/18
|
Speaking style
|
437408 |
9.4684
0.428828 minDCF
|
|
AliMeeting Near/Far
AliMeeting
|
#1/18
|
Distance
|
220000 |
3.6000
0.189475 minDCF
|
|
AliMeeting Overlap
AliMeeting
|
#1/18
|
Overlap
|
165000 |
4.9000
0.280680 minDCF
|
|
VoxKnesset
VoxKnesset
|
#15/18
|
Aging
|
158312 |
20.2049
0.837278 minDCF
|
|
VoxPopuli Aging
VoxPopuli
|
#18/18
|
Aging
|
146575 |
18.2139
0.992315 minDCF
|