Total: 1
Multi-talker automatic speech recognition (ASR) has attracted increasing attention for overlapping speech scenarios. Hypothesis Clustering and Merging (HCM) achieves strong performance by clustering hypotheses in transcript space, but it does not explicitly consider speaker identity during clustering. As a result, HCM may fail when multiple speakers utter identical content and may not fully utilize enrollment information in target-speaker settings. In this paper, we incorporate continuous speaker embeddings into the HCM framework by redefining the clustering distance in joint transcript-speaker space. The proposed method improves robustness in target-speaker-free scenarios under identical-content conditions and enables more effective use of enrollment information in target-speaker multi-talker ASR. Experimental results show consistent improvements over conventional HCM, achieving up to 46% relative WER reduction under identical-content conditions while maintaining competitive performance on LibriMix benchmarks. In addition, the proposed speaker-aware formulation improves target-speaker multi-talker ASR by enabling embedding-based speaker selection, achieving 23% WER reductions compared to discrete speaker-ID prompting.