---
title: "New Microsoft AI Can Clone Your Voice From Three Seconds of Audio"
description: "Microsoft has a new text-to-speech AI that can clone your voice, tone and all, from just a quick three-second snippet of audio. It's called VALL-E."
date: "2023-01-15"
modified: "2023-01-15"
authors:
  - name: "Maggie Harrison Dupré"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/mharrison"
url: "https://futurism.com/the-byte/new-microsoft-ai-clone-your-voice"
categories:
  - "Artificial Intelligence"
tags:
  - "ai"
  - "generative ai"
  - "microsoft"
  - "the digest"
---

# New Microsoft AI Can Clone Your Voice From Three Seconds of Audio

![Microsoft has a new text-to-speech AI that can clone your voice, tone and all, from just a quick three-second snippet of audio. It's called VALL-E.](<https://futurism.com/wp-content/uploads/2023/01/voicecopy.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## VALL-E Parking

Microsoft [says its new text-to-speech AI](<https://arstechnica.com/information-technology/2023/01/microsofts-new-ai-can-simulate-anyones-voice-with-3-seconds-of-audio/?comments=1&comments-page=1>) can clone your voice, tone and all, from a three-second snippet of audio. It's called VALL-E, and we have mixed feelings.

The underlying tech behind the system, which Microsoft refers to in a [new paper](<https://valle-demo.github.io/>) as a "neural codec language model," is complex — but in practice, using the system appears to be wildly simple. Plug in an audio sample, then some text, and voilà: real-sounding speech.

Of course, many text-to-speech apps already exist. Most news sites, us included, for example offer machine-powered dictation services, while speaking assistants like Siri and Alexa are hugely popular.

Most existing speech-generating programs, however, require a large amount of input. They also haven't exactly figured out how to make AI voices sound particularly human, mostly due to the fact that emotional tone and tiny inflections are incredibly complex to convey.

If Microsoft's system [really can](<https://www.engadget.com/the-morning-after-microsofts-vall-e-ai-can-replicate-a-voice-from-a-three-second-sample-121605576.html>) deliver on the tone piece, with that little required on the input side? That's a big deal.

## Mixed Feelings

According to its creators, VALL-E has a number of applications, including "zero-shot TTS, speech editing, and content creation," adding that OpenAI's GPT-3 language modeling system — a technology that Microsoft, per its [absolutely massive investment](<https://futurism.com/the-byte/openai-billions-bad-ai>) into OpenAI, has put a ton of resources into and is [already working](<https://futurism.com/the-byte/microsoft-openai-office-deal>) into [several products](<https://www.theinformation.com/articles/microsoft-and-openai-working-on-chatgpt-powered-bing-in-challenge-to-google>) — would be a particularly useful piece of tech to combine with the new speech generator as a means of churning out content.

And if the latter is something you might be into, Microsoft does have a point. Theoretically, by combining VALL-E and GPT-3 — two powerful pieces of AI-driven tech — you could patch together a ton of real-sounding, believable content, *incredibly* quickly.

But that, of course, is where some ethically-tricky hypotheticals enter the picture.

Fake and misleading sound bytes are obviously a concern here — after all, if you only need three seconds of audio, you could theoretically use anything from a celebrity interview to a real person's Instagram story to impersonate someone.

That said, Microsoft was careful to address that concern, explaining that it's refraining — at least for now — from making the code open source due to "potential risks in misuse of the model." They also claim that they're working on incorporating some kind of system that detects whether audio was created using VALL-E, buuuuuut maybe [they should ask their friends](<https://futurism.com/the-byte/professors-alarmed-ai-undergrads>) over at OpenAI how easy that really is.

**READ MORE:** [*Microsoft's new AI can simulate anyone’s voice with 3 seconds of audio*](<https://arstechnica.com/information-technology/2023/01/microsofts-new-ai-can-simulate-anyones-voice-with-3-seconds-of-audio/?comments=1&comments-page=1>) \[*Ars Technica*\]

**More on Microsoft \<3 AI:** *[Microsoft Working on Deal to Add OpenAI's GPT into MS Word](<https://futurism.com/the-byte/microsoft-openai-office-deal>)*

## Author
At Futurism, I've reported extensively on the rise of AI as a cultural and business force shaping the media industry, and more broadly how those dynamics are changing how we all consume and share information and relate to one another. I'm also fascinated by public health policy and ethics, the role of emerging tech in politics and governance — and the powerful people and forces at those intersections — climate change, and the environment. My investigation on Sports Illustrated's use of AI-generated authors with fictional biographies won a 2024 Mirror Award for "Best Story on Media Coverage of Artificial Intelligence in Journalism and the Media" from Syracuse University's SI Newhouse School, I contributed to Niemen Lab's 2025 Predictions for Journalism series, and I've discussed my work for Futurism during appearances on NPR, CNN, the BBC, the CBC, and more. I grew up in rural Pennsylvania and attended the University of Massachusetts Amherst, where I played Division I field hockey for the Minutewomen as a midfielder. Since then, I've lived in New Orleans, Louisiana and Manhattan, New York. I spend my free time running, reading, perusing archival fashion, and searching for the world’s best negroni. I also have a debonair tuxedo cat, Westley, who's named after "The Princess Bride."

### Author social links  
[Bluesky](<https://bsky.app/profile/mharrisondupre.bsky.social>)