Mobile News

Microsoft Unveils VALL-E, Audio AI That Can Simulate Any Voice From 3-Second Prompts

January 10, 2023

Microsoft researchers recently announced VALL-E, a new text-to-speech AI model that can accurately mimic a person’s voice when given a three-second audio sample. Once it has learned a specific voice, VALL-E can synthesise audio of that person saying anything—while attempting to retain the speaker’s emotional tone. When combined with other generative AI models like GPT-3, VALL-E’s creators believe it can be used for high-quality text-to-speech applications, speech editing in which a recording of a person could be edited and altered from a text transcript (making them say something they did not actually say), and audio content creation.

According to Microsoft, VALL-E is primarily a “neural codec language model,” and is based on EnCodec, which Meta revealed in October 2022. VALL-E creates discrete audio codec codes from text and acoustic prompts, as opposed to other text-to-speech methods that typically synthesise speech by manipulating waveforms. It processes how a person sounds, breaks the relevant data down into discrete components (referred to as “tokens”) using EnCodec, and then uses training data to match what it “knows” about how that voice might sound if it spoke other phrases beyond the three-second sample.

Microsoft trained VALL-E’s speech synthesis functionalities using Meta’s LibriLight audio library. It includes 60,000 hours of English language speech from over 7,000 speakers, sourced primarily from LibriVox public domain audiobooks. The voice in the three-second sample should closely resemble a voice in the learning algorithm for VALL-E to produce a good result.

The American technology giant offers dozens of audio examples of the AI model in action on the VALL-E example website. The “Speaker Prompt” data set is the three-second audio given to VALL-E that it must try to emulate. The “Ground Truth” is a previously recorded version of that same speaker saying a specific phrase for comparative purposes (sort of like the “control” in the experiment). The “Baseline” sample is generated by a traditional text-to-speech synthesis method, and the “VALL-E” sample is generated by the VALL-E model.

A block diagram of VALL-E as shown in the example website by Microsoft researchers
Photo Credit: Microsoft

Researchers only supplied the three-second “Speaker Prompt” sample and a text string (what they would want the voice to say) into VALL-E to get those results. Some VALL-E results appear computer-generated, but others could be misunderstood for human speech, which is the model’s goal. Because of VALL-E’s potential to fuel wrongdoings and deceit, Microsoft has not made VALL-E code available for others to explore. The researchers appear to be aware of the potential social harm that this technology may cause.

They write in the paper’s conclusion: “Since VALL-E could synthesize speech that maintains speaker identity, it may carry potential risks in misuse of the model, such as spoofing voice identification or impersonating a specific speaker. To mitigate such risks, it is possible to build a detection model to discriminate whether an audio clip was synthesized by VALL-E. We will also put Microsoft AI Principles into practice when further developing the models.”

Affiliate links may be automatically generated – see our ethics statement for details.

Catch the latest from the Consumer Electronics Show on Gadgets 360, at our CES 2023 hub.

Poco C55 Tipped to Be a Rebranded Redmi 12C, Expected to Launch Soon

Microsoft Unveils VALL-E, Audio AI That Can Simulate Any Voice From 3-Second Prompts

LEAVE A REPLY Cancel reply

TECH NEWS

Gartner Predicts Legal, Risk and Compliance Functions to Double Technology Spend...

Microsoft to End Support for Windows Mail, Calendar and People Apps...

IDC Predicts: Asia/Pacific Business Leaders to Demand 80% Success Rate on...

The Cooling Conundrum: AI and Automation Push Data Centers Toward 3X...

Gartner Identifies Four Emerging Challenges to Delivering Value from AI Safely...

The Future of Data Protection: A Deep Dive into NAKIVO Backup...

TOP STORIES

Most banks and insurers adopt cloud solutions with the primary objective...

India’s Web3 Ecosystem Has Over 400 Firms, Karnataka Emerges as Industry...

Next-generation spirits innovation to be shaped by premiumization, convenience, generational shifts,...

Trump Triumph: What it Means for Big Tech, Tariffs, Semiconductors, Automotive...

High- cyber-maturity organizations expect to achieve their business outcomes by 27%...

AI Adoption in 2024: 74% of Companies Struggle to Achieve and...

Cyber Security

Tenable Forecasts Data Security in the Cloud to Take Centre Stage...

Blockchain-Enhanced Cybersecurity-Safeguarding Digital Identities and Data

New F5 Report Unveils Scary Truths About API Security in the...

SteelFox exploits Foxit PDF Editor and AutoCAD for banking data theft...

Kaspersky identifies new stealthy ransomware

Gartner Survey Shows AI Enhanced Malicious Attacks as Top Emerging Risk...