OpenAI, the company behind ChatGPT, has developed a tool capable of mimicking a voice based on a 15-second audio snippet. Named Voice Engine, the model is described in a blog post by OpenAI. It functions by converting textual input into speech, leveraging a brief audio sample to replicate a voice, including nuances like intonation and emotion. According to OpenAI, the sample duration required is only fifteen seconds.
However, OpenAI has not disclosed detailed information about the tool, such as the data it’s trained on, and there is no whitepaper or technical description available. The company mentions that Voice Engine is trained on a combination of licensed and publicly available data, clarifying that user data is not used in training. Additionally, samples created by users are deleted after use.
While Voice Engine is not yet accessible to users, OpenAI hints that it may be monetized in the future, potentially charging $15 per million characters or approximately 160,000 words pronounced. Similar to Meta’s Voicebox, which generates text from short audio files, Voice Engine may also face restrictions on public availability due to potential misuse. OpenAI specifically highlights concerns regarding the upcoming US presidential elections and the risk of abuse during political campaigns.
To address potential misuse, OpenAI has limited access to Voice Engine and requires testers to sign statements agreeing not to generate texts without the subject’s permission. The generated audio also includes a watermark to indicate its origin, and OpenAI monitors usage proactively. Furthermore, if Voice Engine is released publicly, OpenAI plans to maintain a list of voices that should not be cloned. Examples of Voice Engine’s capabilities are showcased on OpenAI’s blog.