| |
How We Made a Text-to-Speech Model Respond in Sub-50 ms
Researchers at Alibaba's Qwen team developed a text-to-speech system using their 1.7B CustomVoice model that achieves sub-50 millisecond response times while handling 10 requests per second on a single NVIDIA H100 GPU, significantly outperforming competing implementations like vLLM-Omni and VoxServe. The system produces audio at approximately 630 characters per second with a cost of roughly $2 per million characters, substantially cheaper than commercial competitors like ElevenLabs ($100/1M characters). Key optimizations included removing leading silence from audio output and tuning frame accumulation parameters to balance low latency with continuous playback.
Read Full Article →
← More Tech news