Deepfake Voice Attacks on Finance Teams: A Defense Guide
Three seconds of audio can clone a voice. How deepfake calls target finance teams, what the Arup and Ferrari cases teach, and the playbook that stops them.

The call comes in at 4:50pm on a Friday. It is the CFO's voice — cadence, accent, the slight impatience — asking you to push through a confidential payment before close of business. Everything about the voice is right. The only thing wrong is that the CFO never made the call.
That scenario stopped being hypothetical years ago. A finance employee at engineering firm Arup joined a video conference with what looked and sounded like the company's CFO and several colleagues, then made 15 transfers totaling about $25.6 million. Every other participant on that call was a deepfake. The technology has only improved since, while the economics have gotten worse: CrowdStrike's 2025 Global Threat Report recorded a 442% increase in voice-based phishing between the first and second half of 2024, and the Verizon 2026 Data Breach Investigations Report still attributes 62% of breaches to the human element.
This guide is about the specific collision of that technology with one function: the people authorized to move money. It covers why finance teams are the target, what three well-documented cases reveal about what works and what fails, and a verification playbook that holds even when the voice is perfect.
Why the calls go to finance
Deepfake voice is not a new attack so much as a force multiplier for an old one. CEO fraud — an urgent payment request impersonating a senior executive — predates AI entirely, and business email compromise remains one of the costliest crime categories the FBI tracks, with IC3 reporting roughly $3 billion in BEC losses in its 2025 annual report against $20.9 billion in total cybercrime losses. What a cloned voice adds is the trust layer email could never fake: a familiar voice answering your questions in real time.
Regulators have flagged the shift explicitly. FinCEN's November 2024 alert warned financial institutions of a rise in suspicious activity involving deepfake media, from AI-generated identity documents to executive impersonation. Deloitte's Center for Financial Services projects generative AI-enabled fraud losses could reach $40 billion in the US by 2027, up from $12.3 billion in 2023.
The targeting logic is simple. Finance teams combine payment authority with a culture of executive responsiveness, and their approval workflows — who can release what, above which threshold — are often reconstructable from the outside. An attacker who clones the CEO does not need to compromise a single system. They need one person to treat a voice as authentication. The same logic drives invoice fraud: the ask is routine, the pressure is social, and the money moves through legitimate channels.
Three calls, three lessons
| Case | Channel | The ask | Outcome |
|---|---|---|---|
| Arup (2024) | Video conference with deepfaked CFO and colleagues | 15 "confidential" transfers | ~$25.6M lost |
| Ferrari (2024) | WhatsApp messages, then a live voice-cloned call from the "CEO" | Urgent secret acquisition requiring a currency transaction | Foiled — the executive asked a question only the real CEO could answer |
| LastPass (2024) | WhatsApp calls, texts, and deepfake voicemail of the CEO | Establish urgent contact outside normal channels | Foiled — the employee ignored it and reported to security |
Read together, the cases are unusually instructive, because the difference between the eight-figure loss and the two near-misses was never detection technology. Nobody's ears saved them.
At Ferrari, the voice impersonating CEO Benedetto Vigna was convincing — right down to the southern Italian accent — but the executive on the call did something no clone could survive: he asked about a book Vigna had recently recommended. The impostor hung up. At LastPass, the employee flagged exactly the signals the company trains on: contact outside established business channels and forced urgency. At Arup, by contrast, the employee's initial suspicion of a phishing email was dissolved by the video call itself — the fake meeting was the reassurance, and no procedural checkpoint stood between that reassurance and the transfers.
A voice clone does not need to fool your ears forever — only long enough for urgency to override procedure. Verification has to live in the process, not in the listener.
How a cloned call comes together
Understanding the production pipeline clarifies where defenses bite. First comes reconnaissance: attackers use OSINT — earnings calls, keynote videos, LinkedIn org charts, press releases — to pick the voice, the victim, and a plausible pretext. Second, the clone: McAfee's voice cloning research found freely available tools producing an 85% voice match from as little as three seconds of audio. Third, the channel: attacks favor personal mobile numbers and WhatsApp, where email security and call recording never look — the same blind spot exploited by the broader wave of vishing and smishing attacks. Some campaigns invert the flow as callback phishing: an email plants a phone number, and the victim initiates the call, arriving pre-trusting.
Then comes the script: a confidential acquisition, a supplier emergency, a regulator deadline — always urgent, always secret, always ending in a payment or a credential. The FBI's December 2024 public service announcement on generative AI fraud documents this full stack, from cloned audio to AI-generated documents backing up the pretext.
The verification playbook for finance teams
1. Make out-of-band callback non-negotiable. Any voice or video request to move money, change payment details, or share credentials gets verified on a separate channel: hang up and call back on the number in the corporate directory — never the number that called, never a number provided during the call. The FBI additionally suggests agreeing a family-style secret word or phrase for high-risk approvals; treat it as one layer, not the layer, since anything shareable can leak.
2. Build controls that assume identity is unverifiable. Dual authorization above a defined threshold, separation of duties between initiating and releasing payments, and a mandatory cooling-off period for first-time or changed beneficiaries. These controls work because they are indifferent to how convincing the requester sounds — the Arup loss required one person able to execute fifteen transfers on the strength of one meeting.
3. Protect the vendor-change pathway. Bank-detail changes are the quiet cousin of the dramatic CEO call and where much BEC money actually moves. Every change to supplier payment information gets verified against a known contact at the supplier, on a previously established number, regardless of how legitimate the request looks.
4. Make refusing the CEO safe. The Ferrari save happened because an executive felt entitled to challenge his CEO. Put it in writing: no one will ever be penalized for delaying a payment to verify it, and real executives expect verification. An urgency-plus-secrecy request that discourages verification is itself the strongest indicator of fraud you will ever get.
5. Rehearse it, then measure it. A policy nobody has practiced loses to a voice everybody recognizes. Run deepfake voice simulations against finance and treasury roles specifically, testing whether the callback actually happens under time pressure. Feed the results into a human risk score so the people with payment authority — your highest-consequence users — get targeted coaching rather than generic annual training. And when an attempt occurs, report it: internally first, then to IC3; regulated institutions should reference FinCEN's deepfake alert in suspicious activity filings.
The bottom line
Voice used to be a biometric your organization got for free — recognize the CFO, trust the CFO. Cheap cloning has revoked that assumption, and it is not coming back. The finance teams that handle this well are not the ones buying deepfake detectors; they are the ones that moved trust out of human perception and into process, then rehearsed that process until it survives a convincing voice on a Friday afternoon. The technology behind the calls will keep improving. A callback to a known number works anyway.
Frequently asked questions
Can you reliably tell a deepfake voice from a real one?
Not anymore, and defense plans should assume you cannot. Early voice clones had telltale artifacts — flat intonation, odd pauses, latency when answering questions. Modern real-time cloning has eroded most of these cues, and a stressed employee on a bad phone line was never a reliable detector to begin with. The controls that work are procedural: out-of-band callback on a known number, dual authorization for payments, and verification steps that do not depend on recognizing anyone's voice.
How much audio does an attacker need to clone someone's voice?
Very little. McAfee's research found that some freely available tools can produce an 85% voice match from just three seconds of audio, and better matches with more sample material. For executives, sample audio is effectively public — earnings calls, conference keynotes, podcasts, and social media video provide hours of clean recordings. Assume every senior leader's voice can be cloned, and design payment controls accordingly.
What should a finance team do immediately after a suspected deepfake call?
Freeze any payment or account change the call requested, then verify the request through a separate, known-good channel — call the person back on the number in your directory, not the one that called you. Preserve evidence: call recordings, phone numbers, messages, and payment details. Report internally to security, and externally to the FBI's IC3 in the US. Regulated financial institutions should also reference FinCEN's deepfake alert when filing a suspicious activity report. Speed matters most on the payment side — recalls are only realistic in the first hours.
Does a call from a known number or personal WhatsApp account prove identity?
No. Caller ID is trivially spoofed, phone numbers can be hijacked through SIM swapping, and messaging accounts can be taken over or simply imitated with a new account carrying the executive's photo. Both the Ferrari and LastPass incidents arrived through WhatsApp — a channel where corporate security tooling has no visibility. Treat the channel and the number as unverified metadata; only an independent callback or an established approval workflow verifies identity.
NOUSEC simulates attacks across 8 channels and turns the results into one number your board can read.
Book a demo