AI can clone a voice from 3 seconds of audio. Here's how these scams actually work, the real tells, and the one habit that stops them.

Casey Rowland

TL;DR
Just 3 seconds of audio is enough to clone a voice convincingly, usually scraped from social media videos or a voicemail greeting.
The average loss from a family-emergency voice scam is around $11,000, and the FTC logged $2.7 billion in imposter-scam losses in 2024 alone.
You probably can't reliably detect this by ear. The auditory tells exist, but researchers are consistent that they're not something most people can catch under pressure.
The real defense is a process, not a skill: a family code word, agreed on in advance, that a scammer can't know.
Hang up and call back on a number you already have before sending money or information, every time, no exceptions.
AI voice is one of the wildest things that we've seen with the development of AI. The fact that it only takes three seconds of an audio recording to create a highly accurate AI voice agent can be scary given what people can do with it.
The "grandchild in trouble" call used to rely on a stranger's voice being close enough to fool someone in a panic. Now it doesn't have to be close, it can be an actual clone of the specific person's voice, built from a few seconds of audio anyone posts publicly without thinking twice about it.
How this actually works
A few seconds of audio, pulled from a public social media video, an Instagram Reel, a TikTok, or even a voicemail greeting, is enough for AI tools to generate new speech in that person's voice, saying anything the scammer types. Places that you wouldn't think twice about having your voice recorded can be used against you.
The most common version is a family emergency call. Someone claiming to be a grandchild, child, or other relative, distressed saying they've been in an accident, arrested, or are in some kind of trouble, and need money sent immediately.
The tactic works because it's designed to override the normal instinct to slow down and verify. A panicked, tearful voice that sounds exactly like your child or grandchild doesn't feel like something to fact-check, it feels like an emergency.
What a call like this actually sounds like
These scams follow a recognizable shape, even though the details change. A composite, illustrative version, not a specific real case:
The phone rings from an unfamiliar or spoofed number. The voice on the other end sounds exactly like a grandchild, upset, talking fast.
"I'm in trouble, please don't tell Mom and Dad." The request to keep it secret is deliberate, it removes the chance that someone else in the family says "wait, let's check."
A second voice often joins, claiming to be a lawyer, police officer, or bail bondsman, adding authority and urgency, and providing payment instructions.
The ask is specific and time-pressured: a certain amount, sent a certain way (gift cards, wire transfer, sometimes cash picked up by a courier), within the hour.
Nothing about that sequence is unusual for this scam type. The secrecy request and the second "authority figure" voice are both common enough to be worth recognizing on their own, independent of whether the voice itself sounds cloned.
The scale of this
Stat | Figure |
|---|---|
Audio needed to clone a voice | As little as 3 seconds |
Average loss, family-emergency scam | ~$11,000 |
FTC imposter-scam losses, 2024 | $2.7 billion |
FBI-reported AI scam losses targeting adults 60+ | $352 million |
The tells (and why you shouldn't rely on them alone)
There are auditory signs researchers point to. Most people cannot reliably catch these in the moment, especially while panicked unless they are somewhat educated. Don't treat this list as a foolproof detector.
A slight pause before responses. AI needs a beat to process what you've said, which can come across as an odd hesitation before each reply.
Flat or mismatched emotion. The words might be distressed, but the tone can sound strangely even or robotic underneath it.
A shift in pacing or personality mid-call. This can happen when a call moves from an AI-generated opening to a live scammer taking over.
Audio artifacts. Faint buzzing, echo, or an unnatural smoothness that a real phone call typically doesn't have.
Ear-only detection isn't something to bet money on. Defense against this type of scam comes from follow a process that you've put into place, not your perception of the situation.
The one habit that actually stops this
The FTC's own recommendation is a family code word, a word or phrase agreed on in advance with close family members. Ask for the word whenever someone calls with an urgent request for money. A scammer working from a cloned voice and a script has no way to know it.
Pair that with a second check: hang up and call the person back directly, on a number you already have saved, not a number the caller gives you. If it's really your relative, they'll still be reachable two minutes later. If it wasn't them, you've just avoided sending money to a stranger.
A few other things worth doing:
Ask a specific personal question a stranger couldn't answer, a recent shared event, not something guessable from social media
Never pay by gift card or wire transfer based on a phone call alone, both are close to impossible to reverse
Treat urgency itself as a signal to slow down, not speed up, real emergencies rarely require an instant, untraceable payment
Reducing your own exposure
Detection and verification both deal with a call that's already happening. There's an earlier layer worth thinking about too: none of this works without a sample of someone's real voice to clone in the first place. A few habits reduce how much raw material is out there:
Set social media accounts with video or voice content (public Reels, TikToks, long unlisted YouTube videos) to private where that's realistic for your family
Be cautious about answering calls from unknown numbers with extended conversation, robocalls and "can you hear me" scam calls are sometimes specifically fishing for recordable audio, not just a yes
Consider a generic voicemail greeting rather than one recorded in your own voice, especially for elderly relatives who are more frequently targeted
None of this eliminates the risk, a public figure or anyone with any video online at all has enough audio out there already, but it's a real lever for reducing how easy a specific person is to target, on top of the verification habits above.
The bottom line
AI voice cloning has made an old scam harder to catch by ear, so the fix isn't sharper listening, it's a habit: agree on a family code word now, before anyone needs it, and make hanging up and calling back the automatic response to any urgent money request by phone. This hits families with elderly parents hardest, which is exactly the gap EverGuard's family plan and recovery support are built to cover if a scam like this does get through.
FAQ
How can you spot a fake AI voice?
Listen for a slight pause before responses, flat or mismatched emotional tone, sudden shifts in pacing, or faint audio artifacts like buzzing or echo. These help, but they're not reliable enough to depend on alone, especially in an urgent, emotional call.
How can I tell if someone is using an AI voice on a call?
You often can't with confidence by ear alone. The reliable method is verification, not detection: ask for a pre-agreed family code word, or hang up and call the person back on a number you already have.
Is there a real AI voice message scam happening right now?
Yes. The FTC and FBI have both issued warnings specifically about AI-enhanced family emergency and grandparent scams, and losses tied to them are a growing share of reported imposter-scam fraud.
How much audio does a scammer need to clone someone's voice?
As little as 3 seconds, often pulled directly from a public social media video or a voicemail greeting, no special access required.

