OpenAI-compatible TTS endpoint using OmniVoice: 0.3s response time
A new OpenAI-compatible TTS server, utilizing OmniVoice, offers rapid speech generation with a response time of approximately 0.3 seconds on an RTX 3080. This server can clone voices from short reference clips and is optimized for quick generation, processing long texts paragraph by paragraph. The developer uses it to convert school books into audiobooks in their own voice, enabling listening while driving. More details are available on the project page.
This new OpenAI-compatible TTS server achieves a 0.3-second response time, a speed benchmark for voice cloning and text-to-speech generation unlike many other solutions.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月8日 16:31 UTC
收录当时偏移:UTC+02026年10月9日 02:00 UTC
- 发布
- 2026年10月8日 16:31
- 收录
- 2026年10月9日 02:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
I want to share a TTS server with an OpenAI-compatible API that generates speech really fast (about 0.3 seconds for a sentence on an RTX 3080) and can clone a voice from a short reference clip. I’ve optimized the server so generation starts quickly, and it processes long text in sequence, paragraph by paragraph. I use it to turn school books into audiobooks in my own voice, so I can listen to them while driving.
Out of the box, it’s already tuned for the best settings, but you can change them however you like, for example, the CFG (guidance) scale.
Here are the links to the repo and to a page that showcases it, where you can listen to all the voices. As always, it’s open source and free for anyone to use and modify.