What happened

The paper introduces an iterative pseudo-labeling training approach for Mandarin-English code-switching automatic speech recognition, marking the first application of this method to CS-ASR.

The approach uses a large unlabeled corpus to generate pseudo-labels, creating a semi-supervised dataset for training.

Training proceeds in three phases: pseudo-label generation, two-stage bilingual model training, and iterative improvements.

Why it matters

Code-switching, where speakers alternate languages within a single utterance, is notoriously difficult for ASR systems because dedicated training data is scarce.

By leveraging unlabeled data through iterative pseudo-labeling, this approach offers a potential path to improving CS-ASR without relying on costly manually transcribed code-switched speech.

The iterative refinement process could help models progressively improve their own predictions, making better use of available bilingual resources.

Key facts

Code-switching involves alternating languages within the same utterance and poses significant challenges for ASR.

This paper applies iterative pseudo-labeling to CS-ASR for the first time.

The approach has three phases: pseudo-label generation, two-stage bilingual model training, and iterative improvements.

What to watch next

Whether the iterative pseudo-labeling approach can be extended to other code-switching language pairs beyond Mandarin-English.

How the quality of generated pseudo-labels evolves across iterations and whether it leads to sustained performance gains.

Potential integration of this semi-supervised method into practical ASR systems deployed in multilingual or code-mixed contexts.

Sources