Abstract
Speaker diarization in code-switched multilingual settings remains challenging, particularly for low-resource language pairs such as Kazakh-Russian. Existing pretrained pipelines (Pyannote 3.x, NeMo Sortformer) are predominantly trained on English-centric data, leading to suboptimal speaker discrimination in bilingual conversations. We present a lightweight (4.8M parameters) End-to-End Neural Diarization (EEND) model specifically optimized for Kazakh-Russian code-switching. The architecture combines a Transformer encoder with permutation-invariant training, regularized by SpecAugment and aggressive audio-domain augmentation. Evaluated on the public norwooodsystems bilingual dataset (140 files, 20.3 hours), the model achieves a pooled Diarization Error Rate (DER) of 18.98 percent, competitive with Pyannote 3.1 (19.37 percent) and substantially better than Pyannote 3.0 (20.90 percent) and NeMo Sortformer (24.28 percent). Paired statistical analysis (Wilcoxon signed-rank test on 140 files) identifies a robust and large-effect advantage in Missed Detection rate (6.34 percent versus 9.95 to 11.29 percent, p less than 0.001, Cliff's delta of 0.66) and the lowest Speaker Confusion at the aggregate level (4.04 percent versus 6.30 to 11.38 percent), at the cost of higher False Alarm. The low-Miss operating point is particularly suited to forensic transcription pipelines, where missed speech is unrecoverable while false alarms can be filtered downstream. The model runs approximately four times faster than Pyannote 3.1 with a six times smaller parameter footprint. A detailed per-file error analysis reveals systematic annotation issues in benchmark data and motivates further work on improved evaluation protocols for low-resource bilingual diarization.
| Original language | English |
|---|---|
| Title of host publication | 2026 International Conference on Cybersecurity, Digital Forensics, and AI Applications (ICCSDFAI) |
| Publisher | IEEE |
| Pages | 856-861 |
| Number of pages | 8 |
| ISBN (Electronic) | 9798319520487 |
| ISBN (Print) | 9798319520494 |
| DOIs | |
| Publication status | Published (VoR) - 17 Aug 2026 |
Keywords
- Modeling
- Automatic speech recognition
- Kazakh-Russian;bilingual speech processing
- Speech Training
- Density estimation robust algorithm
Fingerprint
Dive into the research topics of 'Bilingual End-to-End Neural Diarization for Kazakh-Russian Code-Switched Speech'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver