AI Text-to-Speech Tools Advance at Breakneck Speed, Enhancing Training Development

It’s remarkable how rapidly AI-powered text-to-speech technology has advanced over the past six months.

The current level of technology is so advanced that it can produce spoken words from a written script that sound almost natural, making it difficult for some listeners to discern the difference. Some high-end models can accurately generate speech in numerous languages, while also allowing for the manipulation of attributes such as intonation, inflection, intensity, and pauses. My Chinese colleagues have even confirmed the quality of a Chinese-language training video I produced using a text-to-speech tool that doesn’t support attribute management. Moreover, it’s possible to combine multiple languages seamlessly within a single script, transitioning from spoken Chinese to spoken English and back again.

These tools will be game-changing for training developers who produce content in multiple languages, as a single text-based script can be written in one language, translated into other languages, and then rendered by the AI text-to-speech tool. Below are two samples I produced from the same script, one in English and the other in Chinese. (Note – these videos were produced as tests. They have not been cleaned up for actual rollout to learners.)

English Language Version:

English Language Test

Chinese Language Version:

Chinese Version Test

In addition to examining the specific tool utilized to generate these examples, I have also explored several other tools, such as Amazon Poly, Speechify, Descript, and Google Cloud Text to Speech.

continue Reading