July 2024. A Ferrari executive gets WhatsApp messages from an unfamiliar number, apparently from chief executive Benedetto Vigna. There is a confidential acquisition. There is an NDA to sign. The market regulator and the stock exchange have supposedly already been told.
Then a call. The voice carries Vigna's southern Italian accent. The executive notices something slightly off in the tone, but "slightly off" is not something you can act on.
So he stopped trying to act on it. He asked the caller to name the book Vigna had recommended to him a few days earlier.
The caller could not, and hung up.
The decision point is that he did not try to detect the fake. He reached for a fact that existed between exactly two people and had never been published anywhere, and required the caller to produce it. That works against a crude clone and against a perfect one.
Hold this next to the autopsy for this whole course. At Arup, the finance worker was suspicious too, and there was no such question. The difference between $25.6M gone and nothing at all was not perception. It was procedure.
A parent gets a call. It is their child's voice, crying, saying there has been an accident and money has to be sent right now. The voice is right. The panic is right. The story never happened, and the child is asleep two rooms away.
This is not rare and it is not speculative. Voice cloning that once needed hours of studio audio now works from a short clip pulled off social media. Understanding how that is possible, and precisely where it stops, brings the fear down to a size you can act on.
What a voice model actually learns
A modern voice model does not record and replay you. It learns the pattern: your pitch, your rhythm, the way you land on particular sounds, the shape of your pauses. Given enough of that pattern it can generate sentences you never said, in your voice.
The unsettling part is how little pattern it needs. Around thirty seconds of clear speech is often enough for a convincing result, and the clip does not have to be of anything sensitive. A birthday toast works fine. A podcast appearance is luxurious.
Video deepfakes apply the same idea to a face. Real-time video is harder to do well than audio, which is why voice-only calls remain the most common vector, but the Arup case is the proof that live multi-person video is now inside the budget of an organised group.
Where the fakes still fail, and why that does not save you
Clones are good, not magic. They break in fairly predictable places.
Unscripted back-and-forth. A clone reading a script sounds excellent. A clone forced to answer a surprising question in real time often stumbles, repeats itself, or flattens out. Listen to how an answer arrives, not only to what it says.
Specific shared memories. "What did we argue about at dinner last Sunday?" is a wall the attacker cannot climb, because the information was never online to train on.
Emotional logic. Real people react to what you say. A scam steamrolls, pushing the same request regardless of your replies, because the operator's goal is the transfer and not the conversation.
Those tells are real. Treat them as a bonus, and here is the arithmetic for why.
A 2023 UCL study published in PLOS ONE played genuine and generated speech to 529 participants. They correctly identified the deepfakes 73 percent of the time. Training them to listen for known artefacts improved it only slightly.
Two things make 73 percent a generous number. The generators were 2023-era. And the participants knew they were being tested, which you will not.
Now put it against a lifetime rather than a single call. Suppose you face four of these across your life, and you perform at the study's rate every time:
P(catching all four) = 0.73 x 0.73 x 0.73 x 0.73 = 0.73^4 = 0.284
So the probability that at least one gets through is:
1 - 0.284 = about 72 percent
Read that as the design constraint it is. A per-call accuracy of 73 percent compounds into a roughly three-in-four chance of being fooled at least once, on the calls that matter, over a normal life.
Now do the same for a code word. The attacker's success rate is not 27 percent. It is zero, on every call, because there is no listening involved. Detection is a probability. A shared secret is a gate.
It is engineered to flood you with fear so that you act before you think, and that is the entire design. If a call about a loved one in danger demands money or gift cards immediately, the urgency itself is the strongest evidence that something is wrong. Real emergencies survive a five-minute callback. Scams do not.
The three defences that do not use your ears
You cannot out-listen a technology that improves monthly, so do not try. Use methods that hold even if the fake is flawless.
A family code word. Pick a word or short phrase with the people closest to you. If someone calls claiming to be a relative in trouble, you ask for it. A perfect clone still does not have it. This single habit defeats the emergency-call scam outright and you can set it up over dinner tonight.
Choose something that has never been posted anywhere: not a pet's name, not a birthplace, not a school. The Ferrari question worked precisely because a private book recommendation had no public existence to train on.
The callback rule. Whatever the call claims, you hang up and dial the number you already have saved. If they are fine, you have just learned the call was fake. If you cannot reach them, contact someone else close to them before you do anything with money. Note the direction: you dial out, on a number you already had. A number they gave you is their number.
A second channel for money at work. Any request to move funds, arriving by voice or video, gets confirmed on a different established channel that you initiate. No genuine senior person will be offended. The ones who pressure you to skip it are telling you exactly what they are.
None of these ask you to detect anything. They route around the question. The clone can be perfect and the code word still stops it.
A video call is stronger verification than a phone call, because you can see them.
This was true until recently and the Arup case is where it stopped being true. The finance worker was already suspicious of a phishing email, which was the correct instinct. He was then put on a video call where every other participant, including senior colleagues he recognised, was generated from public footage. The video did not verify anything. It overrode a doubt that was already correct.
Seeing a face is now evidence of exactly the same weight as hearing a voice, which is to say almost none. Escalating from chat to a call to video feels like increasing verification. It is increasing production value.
Do these three things this week
- Set a family code word. It costs nothing and it is the highest-value item in this lesson.
- Decide your callback rule and follow it the next time something feels urgent, including when you are fairly sure the call is real. A habit only protects you if it runs without a decision.
- Lower your panic in advance by knowing this exists. The attack runs on surprise, and you are no longer surprised.
The lab for this course, F111-L, is exactly this made concrete: establish an out-of-band protocol with one real person who could plausibly be impersonated to you, and then test it.
A short public clip is enough to clone a voice, and people identified deepfake speech only 73 percent of the time under test conditions kinder than real life, which compounds to roughly a three-in-four chance of being fooled at least once across a handful of calls. So do not defend by listening. Defend with a code word, a callback on a number you already had, and a second channel for money, all of which reduce the attacker's success rate to zero regardless of how good the fake gets. Set it up before the frightening call arrives, because fear is what the attack is counting on.