SoulX-Transcriber
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
Python★ 286
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
Next-gen AI+IoT framework for T2/T3/T5AI/ESP32/and more – Fast IoT and AI Agent hardware integration
html5 js 录音 mp3 wav ogg webm amr g711a g711u 格式,支持pc和Android、iOS部分Web浏览器、Hybrid App(提供Android iOS App源码)、微信,提供ASR语音识别转文…
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kald…
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pi…
search projects, people, and tags