How AI Voice Cloning Is Targeting Banks and Call Centers

September 15, 2026
Scott Weinberg, CEO and Pat Shaw, SVP

How AI Voice Cloning Is Targeting Banks and Call Centers

For years, the call center was considered the safer channel. Phishing emails could be filtered. Malicious links could be blocked. But a phone call was a conversation between two people, and people are good at recognizing other people. That assumption no longer holds. The voice on the line may not be a person at all, and the attack behind it may already know more about your customer than your own agent does.

This is not a distant risk. It is happening on ordinary weekday afternoons, inside institutions that believed their fraud controls and their people were both well prepared.

The call is no longer the weak link. It’s the front door.

Attackers now combine breached data, dark web research, and IVR mining, quietly probing automated phone systems to learn how an account is structured, before a human ever gets on the line. By the time a live agent answers, the caller already knows the account balance, the last transaction, and the exact questions a legitimate customer would be able to answer without hesitation.

That preparation pays off. In a TransUnion survey of call center organizations, nearly two-thirds of financial industry respondents said the majority of account takeovers they see originate in the call center, not online. The channel we built to help customers has become the channel attackers prefer.

And increasingly, the attack does not remain in a single channel. An attacker may gather information through the IVR, manipulate a call center agent, change account or authentication details, and then complete the fraud through digital banking or a payment channel. The vulnerability may not be one failed control, but the gaps between otherwise effective controls.

The voice on the line might not be a person at all

What has changed most in the last two years is not the social engineering script. It is what a criminal can fabricate before making the call.

A usable voice clone can now be produced from as little as three seconds of audio, often lifted from an earnings call, a webinar, a voicemail greeting, or a video posted to social media. The most cited case in the industry remains the 2024 incident at engineering firm Arup, where a finance employee joined a video call with people who appeared to be the company’s CFO and several colleagues, authorized fifteen transfers totaling $25.6 million, and later learned that every participant on that call had been AI-generated.

The financial scale of this is not a rounding error, and the FBI has started counting it separately. For the first time in the Internet Crime Complaint Center’s twenty-five year history, the 2025 report broke out artificial intelligence as its own category: 22,364 complaints and roughly $893 million in reported losses. Business email compromise overall reached $3.05 billion across 24,768 complaints. The FBI’s own caveat matters more than either number: the AI figure is understated, because many victims never find out an AI was on the other end of the line.

Verizon’s 2026 Data Breach Investigations Report found that 62 percent of confirmed breaches involve a non-malicious human element, up from 60 percent the year before. The same report began tracking pretexting as an initial access vector for the first time, and found that phone and text lures drew click rates roughly 40 percent higher than email in phishing simulations, with 41 percent of social engineering breaches now arriving through non-email channels. That statistic should reframe how we think about this problem. It is not primarily a technology gap. It is a trust gap, and AI has found the fastest way through it.

Why call centers are exposed in particular

Call center agents are trained, appropriately, to be helpful, efficient, and reassuring. That is the job. It is also precisely the instinct an attacker is counting on.

Pindrop, which analyzed more than 1.2 billion calls for its 2025 Voice Intelligence and Security Report, measured a 1,300 percent increase in deepfake fraud attempts across contact centers in a single year, moving from roughly one attempt a month to seven a day. Synthetic voice fraud rose 149 percent in banking contact centers and 475 percent in insurance. The number that should concern us most: 53 percent of fraudsters successfully passed knowledge-based authentication. Voice authentication and security questions were built on two reasonable assumptions: a voice was hard to fake, and only the real customer knew the answers. Neither has aged well.

The challenge extends beyond voice authentication and security questions. Caller ID, one-time passcodes, agent-assisted password resets, device enrollment, profile changes, and escalation procedures can all become part of the attack path. An attacker does not necessarily need to defeat every control. The attacker needs to find a sequence of actions that produces the desired outcome.
Contact centers that rely on video verification are not automatically safer. Video verification should not automatically be assumed to solve the problem either. Real-time face manipulation, virtual cameras, and injection techniques create another verification surface that institutions need to evaluate and test.

What strengthens our defense

None of this means the phone call is beyond saving as a trusted channel. It means we need to stop asking our people to do a job that a voice alone can no longer support, and give them controls that do not depend on “it sounded right.”

  • Out-of-band verification for high-stakes requests. Any request to move money, reset credentials, or change account details should be confirmed through a second, independent channel that the attacker does not control, not through the same call in which the request was made.
  • Behavioral and device signals, not just voice. Device fingerprinting, call metadata, and behavioral patterns can catch what a human ear cannot: a caller who sounds right but is not authenticating the way the real customer normally does.
  • Retrain the workforce around today’s threat, not yesterday’s. Security awareness training built around spotting typos and suspicious links will not help an agent facing a fluent, well-informed, AI-assisted caller. Training needs to teach process discipline, not pattern recognition.
  • Recovery readiness, because some attempts will succeed. No control catches everything. A tested incident response plan, ready before the first fraudulent call comes in, is what determines whether a successful attack becomes a contained loss or an escalating one.
  • Adversarial testing against realistic attacks. Institutions should test their defenses using synthetic and cloned voices, replay attacks, and AI-assisted social engineering. The goal is not simply to determine whether a control detects a deepfake, but whether the institution can stop an attacker from authenticating, changing account information, or moving money.

 

Staying ahead of a fast-moving threat

I do not think the right response to this is alarm. I think it is clarity. The phone call has been a trusted channel for a long time, and it can remain one, but only if our verification processes assume that a convincing voice is no longer proof of anything on its own.

The institutions that adapt their call center defenses now, deliberately and calmly, will be the ones whose customers keep trusting that channel. The ones that wait for a headline incident to force the issue will be making these changes under far more difficult circumstances.

The institutions asking the right questions today are not simply asking whether they have voice-security technology in place. They are asking whether those controls work against the attacks they are likely to face and whether the surrounding processes still hold when an individual control fails.

A convincing voice is no longer proof of identity. The controls around it have to provide the proof.

Onward.