goJumboGPT

AI AI images, video and voice: how they are made

Voice cloning: how a few seconds of audio becomes a fake call

How AI voice cloning works, how little audio it needs, the scams it enables, and the family safe word that defeats almost all of them.

8 min read How we write

The short answer

  • A voice clone is built by turning a short recording into a numerical fingerprint of how someone sounds, then using that fingerprint to steer a speech model.
  • Seconds of clean audio are enough for a usable copy, and a voicemail greeting, a social video or a short phone call all supply it.
  • A clone copies timbre, accent and cadence, but it does not know what the person knows, which is the weakness every practical defense is built on.
  • The two dominant uses are the family emergency call and payment fraud at work, and both rely on urgency, secrecy and an unusual way to send money.
  • Hanging up and calling back on a number you already had defeats almost every version, because the attacker controls the call, not the line.
  • Agree a spoken phrase with your family now, since it costs nothing and works even when the voice is perfect.

Voice cloning works because a voice is far simpler to describe than it sounds. A model listens to a short recording, seconds rather than minutes, and reduces it to a list of a few hundred numbers that capture pitch range, timbre, accent and speaking rhythm. Feed that list into a speech model along with any text you like, and out comes that voice saying something it never said. No permission needed, no long sample, no special equipment, and on ordinary hardware it takes about as long as reading the sentence aloud would. That is the whole trick, and understanding it tells you exactly where the defenses have to go.

What a voice clone actually is

Modern speech systems separate two things that used to be tangled together: what is said and who says it. A speaker encoder handles the second part. It is trained on recordings of thousands of different people, with the one instruction that clips of the same person should land near each other in its output and clips of different people should land far apart. What it produces is called an embedding, a compact numeric summary of a voice rather than a copy of any particular recording.

The synthesis model is trained separately to turn text into audio, conditioned on one of those embeddings. Because the two were trained on many voices, a fresh embedding from a voice the system has never encountered still works. That is what zero shot cloning means: nothing is retrained for you, the new voice is just a new set of coordinates. Give the same model more material, typically several minutes of clean speech, and it can be tuned properly, which buys better handling of unusual names, laughter and emotional range.

There is a second method worth knowing about, because it sounds much more convincing. In speech to speech conversion, a human performs the line and the model keeps the timing, stress and emotion of that performance while swapping the timbre for the target's. Fraud that involves a live back and forth conversation usually works this way. Someone is really talking, in real time, wearing another person's voice.

How little audio it takes, and where it comes from

The published tools generally ask for a sample somewhere between a few seconds and a minute. Quality matters more than length: one clean sentence recorded close to a microphone beats five noisy minutes from across a room. Sources are everywhere and none of them require a breach.

  • An outgoing voicemail greeting, which anyone can hear by calling when you are busy.
  • Any video posted publicly, including ones where you appear in the background of someone else's.
  • Webinars, conference talks, podcast appearances, recorded meetings, customer service calls.
  • A short call from an unknown number in which you are encouraged to keep talking.

That last one deserves a correction, because a widely repeated version of it is wrong. The claim that answering "yes" to a recorded question lets someone authorize charges against you is not how payments or banks work, and no reliable case has ever shown it. The real reason a scammer keeps you talking is to collect audio and to see whether the number reaches a live human. The general pattern of these calls is in how spoofed numbers and robocalls operate.

Practically, you cannot keep your voice private if you have ever spoken in public. Plan on the assumption that a copy is available rather than trying to prevent one.

What a clone copies well, and what it does not

Feature of the real personDoes the clone get it?What that means for you
Timbre and pitchYes, this is the easy partRecognizing the voice proves nothing at all
Accent and pronunciation habitsMostly, if the sample carried themRegional speech is no longer a check
Emotion under stressPoorly from text, well when a human performs itCrying and panic are easy to fake and also easy to justify
Word choice and running jokesOnly if the attacker researched themUnusual phrasing is worth noticing
Shared private knowledgeNoThis is where every reliable defense lives
Natural interruption and overlapAwkward in fully synthetic callsTalking over them can expose a scripted system
Unusual names and numbersOften mangled or oddly stressedAsk them to repeat a name only they would say often

Read the right hand column as a single idea. The clone reproduces the surface of a person and none of the contents. Every defense below is a way of asking for the contents.

The two scams that actually use it

The family emergency call is the common one. A grandparent, parent or partner picks up. The voice is distressed and says there has been an accident, an arrest, a hospital, a problem abroad. Often a second person takes over as a lawyer, officer or doctor, which conveniently explains why the first voice has gone quiet. The request is money now, by a method that cannot be reversed, and above all do not tell anyone else in the family because of the embarrassment, the bail conditions, the investigation.

The workplace version targets whoever can move money. A finance employee gets a call, sometimes a voicemail, sometimes a live video meeting with familiar faces, from an executive who needs a payment made quietly for an acquisition or a supplier problem. The email trail, if there is one, looks right because the mailbox was compromised weeks earlier or the domain is one character different. The audio exists to make a request feel authorized, not to do the whole job.

Notice that the technology sits inside an old structure. Urgency, secrecy and an irreversible payment channel were the shape of these frauds long before any of this was possible. What changed, and what did not, is set out in what AI actually changed about scams. The equivalent pretext by message, where the bank asks you to move money to a safe account, is in how bank impersonation works.

Why detection will not save you

People ask for a tell in the audio. There used to be several: flat delivery, breaths in the wrong places, a slight metallic ring, no room echo. They are mostly gone, and what survives does not survive a phone line. Calls are compressed heavily, which strips the fine spectral detail any subtle tell would live in, and the caller has decided what you hear: background noise, a poor signal, a borrowed phone, sobbing. Every artifact you might notice has a ready explanation supplied in advance.

Automated detectors exist and are used by banks and platforms, but they are graded on being mostly right across millions of calls, not on being right about yours. Treated as a verdict for one call, the score means little, for the reasons set out in why detector scores are not evidence. The same limitation applies to images, which is why spotting a generated picture has moved toward checking sources rather than pixels. Do not build your safety on catching the fake. Build it on a check the attacker cannot pass.

The defenses that actually work

  1. Hang up and call back on a number you already have. Not a number they give you, not a returned call to the one that just rang. This defeats nearly every version, because the attacker controls their call and not your address book.
  2. Change channel. Text the person, message them on an app you both use, call a sibling or a colleague. A caller who objects to you verifying through a second route has told you what you needed to know.
  3. Ask something only they would know and that is not online. Not a birthday or a pet name. What we ate on Sunday, the name of the person who fixed the boiler, which seat you sat in.
  4. Agree a family phrase in advance. Any short, memorable, unguessable phrase, said out loud in a real emergency. Tell it to people face to face and never write it in a message.
  5. At work, require a second approval for any payment change, made on a phone number from your own records. Make it a written rule so nobody has to feel awkward about applying it to a senior person.
  6. Slow down deliberately. Say you will call back in five minutes. Real emergencies survive five minutes, and scripted pressure usually does not.

Set this up this week

Pick a phrase and tell three people. Shorten your voicemail greeting or replace it with the standard automated one, which removes the easiest clean sample of you. Say plainly to older relatives that a call in your voice asking for money is now something anyone can fake, and that you will never mind being called back to check. That conversation lands better when it is framed as your rule rather than as their vulnerability, and helping a relative without taking over covers how to have it.

If money has already gone, act in the next hour rather than the next day. Call your bank and use the words fraud and recall, keep every message and transaction reference, and report it to your national fraud reporting service. The order of operations, and what is realistically recoverable by payment method, is in what to do after sending money to a scammer.

Common questions

Is a short voicemail greeting enough for a scammer to work with?

Seconds of clean speech is enough for a copy that sounds like you on a phone call. A minute or more of good quality audio produces a better one that handles unusual words and emotion. Since a voicemail greeting or any posted video supplies this, treat the sample as already obtainable and put your effort into verification habits instead.

Can I hear the difference between a cloned voice and a real one?

Usually not, and phone compression removes most of what you might otherwise catch. The remaining clues are behavioral rather than acoustic: odd handling of names, awkwardness when you interrupt, resistance to being called back, and a refusal to answer a question that depends on shared history.

Is voice cloning illegal?

Using a cloned voice for fraud or impersonation is already a crime nearly everywhere, and several jurisdictions have added specific rules covering synthetic voices in robocalls, campaigning and commercial use. Cloning your own voice, or someone's with their consent, is generally lawful. Rules vary by country and this is general information rather than advice for your situation.

What should a family safe word be?

Something short, memorable and impossible to guess or research: a nonsense pairing of two unrelated words works well. Avoid pet names, street names, birthdays and anything that has ever appeared on social media. Share it in person, never in writing, and agree that any request for money by phone requires it.

Are voice passwords for banking still safe?

Voice alone is a weak authenticator now, and providers that once relied on a spoken phrase have been adding checks behind it. If your bank offers voice identification, treat it as a convenience rather than a lock, and make sure a genuine second factor such as an app prompt or a device check also stands between a caller and your money.