XDRIPACADEMY
Sign in

Curriculum·F111 AI-Era Threats and Verification·60 min

Deepfakes and voice clones

By the end of this lesson you can

  • Explain what a voice model actually learns, and why a short public clip is now sufficient input
  • Compute why per-call detection accuracy is the wrong thing to rely on across a lifetime of calls
  • Deploy the three out-of-band defences that work against a flawless fake: code word, callback, second channel
  • Identify why the perceptual tells in a fake are a bonus rather than a defence
AutopsyThe Ferrari executive and the book questionnothing, and that is why it is here

July 2024. A Ferrari executive gets WhatsApp messages from an unfamiliar number, apparently from chief executive Benedetto Vigna. There is a confidential acquisition. There is an NDA to sign. The market regulator and the stock exchange have supposedly already been told.

Then a call. The voice carries Vigna's southern Italian accent. The executive notices something slightly off in the tone, but "slightly off" is not something you can act on.

So he stopped trying to act on it. He asked the caller to name the book Vigna had recommended to him a few days earlier.

The caller could not, and hung up.

The decision point is that he did not try to detect the fake. He reached for a fact that existed between exactly two people and had never been published anywhere, and required the caller to produce it. That works against a crude clone and against a perfect one.

Hold this next to the autopsy for this whole course. At Arup, the finance worker was suspicious too, and there was no such question. The difference between $25.6M gone and nothing at all was not perception. It was procedure.

Primary source

A parent gets a call. It is their child's voice, crying, saying there has been an accident and money has to be sent right now. The voice is right. The panic is right. The story never happened, and the child is asleep two rooms away.

This is not rare and it is not speculative. Voice cloning that once needed hours of studio audio now works from a short clip pulled off social media. Understanding how that is possible, and precisely where it stops, brings the fear down to a size you can act on.

What a voice model actually learns

A modern voice model does not record and replay you. It learns the pattern: your pitch, your rhythm, the way you land on particular sounds, the shape of your pauses. Given enough of that pattern it can generate sentences you never said, in your voice.

The unsettling part is how little pattern it needs. Around thirty seconds of clear speech is often enough for a convincing result, and the clip does not have to be of anything sensitive. A birthday toast works fine. A podcast appearance is luxurious.

Video deepfakes apply the same idea to a face. Real-time video is harder to do well than audio, which is why voice-only calls remain the most common vector, but the Arup case is the proof that live multi-person video is now inside the budget of an organised group.

Where the fakes still fail, and why that does not save you

Clones are good, not magic. They break in fairly predictable places.

Unscripted back-and-forth. A clone reading a script sounds excellent. A clone forced to answer a surprising question in real time often stumbles, repeats itself, or flattens out. Listen to how an answer arrives, not only to what it says.

Specific shared memories. "What did we argue about at dinner last Sunday?" is a wall the attacker cannot climb, because the information was never online to train on.

Emotional logic. Real people react to what you say. A scam steamrolls, pushing the same request regardless of your replies, because the operator's goal is the transfer and not the conversation.

Those tells are real. Treat them as a bonus, and here is the arithmetic for why.

Worked example
Why 73 percent is not good enough

A 2023 UCL study published in PLOS ONE played genuine and generated speech to 529 participants. They correctly identified the deepfakes 73 percent of the time. Training them to listen for known artefacts improved it only slightly.

Two things make 73 percent a generous number. The generators were 2023-era. And the participants knew they were being tested, which you will not.

Now put it against a lifetime rather than a single call. Suppose you face four of these across your life, and you perform at the study's rate every time:

P(catching all four) = 0.73 x 0.73 x 0.73 x 0.73 = 0.73^4 = 0.284

So the probability that at least one gets through is:

1 - 0.284 = about 72 percent

Read that as the design constraint it is. A per-call accuracy of 73 percent compounds into a roughly three-in-four chance of being fooled at least once, on the calls that matter, over a normal life.

Now do the same for a code word. The attacker's success rate is not 27 percent. It is zero, on every call, because there is no listening involved. Detection is a probability. A shared secret is a gate.

The emergency call is the highest-pressure attack there is

It is engineered to flood you with fear so that you act before you think, and that is the entire design. If a call about a loved one in danger demands money or gift cards immediately, the urgency itself is the strongest evidence that something is wrong. Real emergencies survive a five-minute callback. Scams do not.

The three defences that do not use your ears

You cannot out-listen a technology that improves monthly, so do not try. Use methods that hold even if the fake is flawless.

A family code word. Pick a word or short phrase with the people closest to you. If someone calls claiming to be a relative in trouble, you ask for it. A perfect clone still does not have it. This single habit defeats the emergency-call scam outright and you can set it up over dinner tonight.

Choose something that has never been posted anywhere: not a pet's name, not a birthplace, not a school. The Ferrari question worked precisely because a private book recommendation had no public existence to train on.

The callback rule. Whatever the call claims, you hang up and dial the number you already have saved. If they are fine, you have just learned the call was fake. If you cannot reach them, contact someone else close to them before you do anything with money. Note the direction: you dial out, on a number you already had. A number they gave you is their number.

A second channel for money at work. Any request to move funds, arriving by voice or video, gets confirmed on a different established channel that you initiate. No genuine senior person will be offended. The ones who pressure you to skip it are telling you exactly what they are.

None of these ask you to detect anything. They route around the question. The clone can be perfect and the code word still stops it.

Common misconception

A video call is stronger verification than a phone call, because you can see them.

This was true until recently and the Arup case is where it stopped being true. The finance worker was already suspicious of a phishing email, which was the correct instinct. He was then put on a video call where every other participant, including senior colleagues he recognised, was generated from public footage. The video did not verify anything. It overrode a doubt that was already correct.

Seeing a face is now evidence of exactly the same weight as hearing a voice, which is to say almost none. Escalating from chat to a call to video feels like increasing verification. It is increasing production value.

Do these three things this week

  1. Set a family code word. It costs nothing and it is the highest-value item in this lesson.
  2. Decide your callback rule and follow it the next time something feels urgent, including when you are fairly sure the call is real. A habit only protects you if it runs without a decision.
  3. Lower your panic in advance by knowing this exists. The attack runs on surprise, and you are no longer surprised.

The lab for this course, F111-L, is exactly this made concrete: establish an out-of-band protocol with one real person who could plausibly be impersonated to you, and then test it.

Key takeaway

A short public clip is enough to clone a voice, and people identified deepfake speech only 73 percent of the time under test conditions kinder than real life, which compounds to roughly a three-in-four chance of being fooled at least once across a handful of calls. So do not defend by listening. Defend with a code word, a callback on a number you already had, and a second channel for money, all of which reduce the attacker's success rate to zero regardless of how good the fake gets. Set it up before the frightening call arrives, because fear is what the attack is counting on.

These come back later

How much audio does a usable voice clone need?
Around thirty seconds of clear ordinary speech, of the kind in any video you have posted. It does not need to be private or high quality. Assume any voice that exists online can be imitated.
In the UCL study, how often did people correctly identify deepfake speech?
73 percent, across 529 participants, and training barely improved it. They also knew in advance they were being tested. Roughly one in four got through under ideal conditions.
What is the single highest-value thing in this lesson?
Agree a code word with close family this week. A perfect voice clone still does not have it, and it costs nothing to set up.

Sources and review

Confidence high·Volatility high·Reviewed 2026-08-05·Owner unassigned

Contested

The 73 percent detection figure comes from 2023-era generators and from participants who had been told they were being tested. Both caveats push real-world performance lower, not higher. Do not present it as a current measurement; present it as a generous upper bound that was already inadequate.

Ferrari has not published a detailed account. Reporting is consistent across outlets on the book question and the outcome, but the internal timeline is second-hand. Attribute it, do not embellish it.

Track your progress

Create a free account to mark lessons complete and pick up where you left off.