Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

How Deepfakes Work — and How Detection Is Trying to Keep Up

A technical walkthrough of the neural networks behind deepfakes and the detection methods trying to keep pace with them.

How Deepfakes Work — and How Detection Is Trying to Keep Up — Woyce Technologies

A deepfake video of a CEO announcing a fake acquisition can move a stock price before anyone confirms it's fake. A cloned voice can authorize a wire transfer over the phone. None of this requires a Hollywood budget anymore — it requires a laptop, a few minutes of source footage, and software that's freely downloadable. Understanding how that software actually works is the first step to knowing where it breaks down, and where detection still has a fighting chance.

For most organizations the problem is practical, not academic. Finance teams approve payments over calls, help desks reset passwords for voices they recognise, and executives appear in hours of public video. Each of those is an opening once faces and voices can be synthesized on demand. Knowing how deepfakes work tells you which controls are worth investing in and which detection claims to treat with caution.

This guide explains what counts as a deepfake, how the original autoencoder and GAN face-swap pipeline works, how diffusion models and modern voice cloning changed the threat, the main families of detection and why they struggle, why provenance standards such as C2PA matter, and the process controls businesses should put in place now.

What a Deepfake Actually Is

"Deepfake" is a portmanteau of "deep learning" and "fake" — it refers to synthetic media, usually video or audio, generated or altered by neural networks to depict someone doing or saying something they didn't. That covers a wide range of techniques:

  • Face swaps — replacing one person's face with another's in existing video footage.
  • Face reenactment (puppeteering) — mapping one person's expressions and mouth movements onto a target face, often used to make someone appear to say new words.
  • Voice cloning — synthesizing a person's speech patterns and timbre to generate new audio in their voice.
  • Full synthesis — generating an entirely artificial person or scene that never existed, rather than manipulating real footage.

These are distinct problems with overlapping techniques, and conflating them is a common source of confusion. A tool that's excellent at swapping faces in a well-lit, front-facing video may fail completely on voice cloning, and vice versa. Detection systems built for one category often don't transfer to another.

It's also worth separating the technology from the intent. The same face-reenactment technique that powers a fraudulent video call also powers legitimate dubbing for foreign-language films, de-aging effects in movies, and accessibility tools that let people who've lost their voice continue to communicate in something close to their original tone. The generation methods are neutral; what makes something a harmful "deepfake" versus a disclosed synthetic-media product is consent and transparency about what the audience is seeing.

The Core Technology: Autoencoders and GANs

The first wave of consumer deepfake tools, popularized around 2017-2019, relied on a neural network architecture called an autoencoder.

How autoencoders learn faces

An autoencoder has two parts: an encoder that compresses an image down into a small set of numbers (a "latent representation"), and a decoder that reconstructs the image from those numbers. Train an encoder-decoder pair on thousands of photos of Person A, and the network learns a compact, generalized representation of Person A's face — the underlying structure that stays consistent across lighting, angle, and expression.

The trick behind classic face-swap deepfakes is to train one shared encoder on both Person A and Person B, but two separate decoders — one specialized for reconstructing A, one for B. Because the encoder is shared, it learns a general "face" representation that captures expression, pose, and lighting independent of identity. Feed a frame of Person A's face through the shared encoder, then run the resulting latent code through Person B's decoder, and out comes Person B's face — but wearing Person A's expression and pose. Do this frame by frame across a video and you get a face swap.

GANs sharpen the output

Raw autoencoder output tends to look blurry and waxy. Generative Adversarial Networks (GANs) fix this by pitting two networks against each other during training: a generator that produces fake images, and a discriminator that tries to tell real from fake. The generator improves by trying to fool the discriminator; the discriminator improves by getting harder to fool. Run this arms race for enough iterations and the generator learns to produce highly realistic textures — skin pores, hair strands, reflections in eyes — that plain autoencoders miss, a dynamic documented extensively in adversarial-training research on arXiv.

Most well-known deepfake tools from this era (including the original open-source projects that popularized the term) combine an autoencoder backbone with GAN-style refinement to sharpen the final frames.

This generation of tools has well-documented weak points that show up under scrutiny: profile views and extreme angles the training data didn't cover, occlusions like hands or hair passing in front of the face, fast head movement that outruns the frame-by-frame reconstruction, and jewelry or glasses that flicker in and out of consistency. Anyone doing manual review of suspect footage still starts by looking at these failure points, because autoencoder-based swaps remain common even as newer techniques emerge — they're cheap to run and don't require the heavier compute that diffusion-based methods need.

Face-swap pipeline with a shared encoder and a decoder per person: a frame of Person A becomes a latent code that decoder B turns into B's face with A's expression and pose.

The Newer Wave: Diffusion Models

Since roughly 2022, diffusion models have overtaken GANs as the dominant technology behind the most convincing synthetic media, including tools built for full face and body synthesis rather than just swaps.

Diffusion models work by a curious inversion of the generation process:

  1. During training, the model is shown images with progressively more random noise added, until the image is pure static.
  2. It learns to reverse this process — predicting, at each noise level, what the slightly-less-noisy version should look like.
  3. To generate a new image, the model starts from pure random noise and repeatedly applies its learned denoising step, gradually resolving structure, then detail, then texture.

Because this process is guided by text or image conditioning, diffusion models can generate faces, scenes, and motion that were never filmed at all — not just remixes of existing footage. This is what powers modern text-to-video and image-to-video tools, and it's a meaningfully different threat model than classic face-swapping: there's no "original video" to compare against, because nothing was ever recorded.

Voice cloning follows a similar arc

Early voice cloning needed hours of clean training audio per target voice. Modern text-to-speech systems, many built on diffusion or transformer architectures similar to those used in large language models, can produce a passable clone from a few seconds to a few minutes of reference audio, then generate arbitrary new sentences in that voice with correct prosody and emotional inflection — the same capability now driving a wave of voice cloning fraud.

The pipeline typically separates two problems: capturing the speaker's timbre and vocal identity (what makes a voice recognizably theirs), and modeling prosody — pacing, stress, intonation, the way someone naturally pauses mid-sentence. Older systems handled these jointly and produced flat, robotic cadence even when the timbre matched. Newer systems condition generation on both a short reference clip and, in many cases, an emotional or stylistic prompt, which is why recent cloned audio can carry urgency, hesitation, or anger convincingly — a capability that maps directly onto how these systems get used in social-engineering scams.

Why This Matters Right Now

The technical barrier to producing convincing synthetic media has collapsed. What once required machine learning expertise, a capable GPU, and days of training per target is increasingly available as a hosted product: upload a photo or a short audio clip, type what you want said, and download the result in minutes. The generation side of this problem has industrialized faster than the detection side, which is the structural reason deepfakes keep showing up in fraud, disinformation, and impersonation incidents across industries — from fabricated executive video calls to cloned-voice phone scams targeting employees and family members.

This asymmetry — cheap, fast, accessible generation versus slow, resource-intensive detection — is the central dynamic anyone building security, trust, or compliance systems needs to internalize. It is not a problem that gets solved once; it's a moving target.

How Detection Tries to Keep Up

Detection approaches generally fall into a few families, each with different strengths and blind spots.

Detection approachWhat it looks forStrengthsWeaknesses
Artifact-based forensicsBlending boundaries, inconsistent lighting/shadows, unnatural blinking, texture anomaliesWorks without prior knowledge of the source modelNewer generators produce fewer detectable artifacts each generation
Physiological signal analysisSubtle signals like blood-flow-driven color changes in skin (remote photoplethysmography), consistent with a real pulseHard to fake without deliberately modeling itCompression and low video quality can destroy the signal in real and fake video alike
Learned classifiersA neural network trained on large datasets of real vs. fake media to spot statistical patternsCan catch patterns humans can't seeProne to overfitting on the specific generators in its training set; degrades against new tools
Provenance and watermarkingCryptographic signing or invisible watermarks embedded at capture or generation timeDoesn't rely on spotting flaws at allOnly works if the capture device or generator participates; strippable by re-encoding in some schemes
Metadata and consistency checksFile metadata, compression history, audio-video sync, background consistencyCheap to run, catches careless fakesEasily defeated by anyone who bothers to clean metadata

The arms-race problem

Every detection method that relies on spotting flaws in generated content has the same structural weakness: once a flaw is publicly known to be detectable, it becomes a target for the next generation of generative models to eliminate. Academic papers on deepfake detection routinely get cited by developers of the next generative model as a to-do list. This is why the field has shifted emphasis toward provenance — proving what's real — rather than only forensics — proving what's fake, after the fact.

Provenance standards are gaining traction

Groups like the Coalition for Content Provenance and Authenticity (C2PA) have proposed standards for cryptographically signing content at the point of capture or edit, creating a tamper-evident record of an image or video's history. Cameras, editing software, and generative AI tools that adopt this standard can attach a verifiable trail showing what created or modified a piece of content. The catch is adoption: provenance only helps if it's embedded broadly enough that its absence becomes meaningful, and that requires cooperation across camera manufacturers, social platforms, and AI vendors that don't share commercial incentives to move at the same pace.

Two-column contrast of forensics, which proves content is fake but loses to the next generator, and provenance, which signs content at capture but depends on broad adoption.

Benefits of Deepfake Detection and Provenance

No detector catches everything, but layered detection and provenance still change the economics of deception. Here is what they buy an organisation, provided they sit alongside process controls rather than replacing them.

Raising the cost of a successful attack

A careless fake with stale metadata or obvious blending artefacts gets caught by cheap automated checks. Attackers then need better tools, more source material and more effort for each attempt. Detection rarely stops a determined, well-resourced adversary, but it filters out the large volume of low-effort attempts that make up most fraud and spam, leaving human reviewers with fewer, harder cases.

Faster triage of suspicious media

Newsrooms, trust-and-safety teams and investigators receive far more content than people can examine frame by frame. Forensic and classifier tools act as a first pass, ranking which clips deserve close manual review. The output is a probability, not a verdict, but even a rough ranking makes expert time go further. It also creates a record of which checks were run on each item, which helps when a decision is later questioned.

Proof that something is real

Provenance flips the question. Instead of hunting for flaws, a signed capture or edit history lets a publisher or platform show where content came from and what changed. For organisations whose own videos and statements might be impersonated, publishing provenance-signed media gives audiences a way to check authenticity, and makes unsigned fakes stand out more as adoption grows.

Evidence for investigations and disputes

When a fraudulent payment or a reputational incident involves synthetic media, forensic analysis and provenance records help reconstruct what happened. Metadata, compression history and detector scores can support internal investigations, insurance claims and legal proceedings, even if none of them is conclusive alone. Preserving original files and call recordings promptly matters, because re-shared copies lose much of the forensic signal.

Support for legitimate synthetic media

Disclosure and provenance tools also help the legitimate side of the industry. Dubbing studios, accessibility tools and marketing teams using consented synthetic voices or faces can label their output verifiably, which separates disclosed synthesis from deceptive impersonation and makes responsible use easier to defend.

Deepfake Detection Use Cases

Detection and verification are applied wherever a decision depends on trusting a face or a voice. These are the settings where organisations most often deploy them, usually combining technical checks with changes to the process itself.

Payment approvals and finance teams

Problem: Attackers impersonate executives on calls or video meetings to request urgent transfers. How it's applied: Out-of-band callbacks, second approvers and, in some organisations, voice or video analysis on high-risk requests. Outcome: The fraud depends on a single person acting on a convincing voice; adding verification steps breaks that chain even when the media itself is flawless.

Identity verification and onboarding

Problem: Remote account opening and KYC checks can be targeted with synthetic faces, replayed video or real-time face swaps. How it's applied: Liveness checks, injection-attack detection and consistency tests between documents and live capture. Outcome: Fewer synthetic identities pass onboarding, although real-time swaps keep raising the bar for liveness detection.

Call centres and help desks

Problem: Cloned voices are used to pass voice-based verification or to talk agents into password resets. How it's applied: Removing voice recognition as a sole factor, adding knowledge or device-based checks, and in some cases screening calls for synthetic speech. Outcome: Account takeover through social engineering becomes harder, with agents following a scripted verification path rather than trusting what they hear.

Newsrooms and fact-checkers

Problem: Viral clips of public figures need to be assessed quickly before they are reported or debunked. How it's applied: Forensic tools, reverse searches, metadata analysis and provenance checks, combined with traditional sourcing. Outcome: Faster, better-documented verification decisions, and published explanations of why a clip was judged authentic or not. Tools inform the judgment; editors still make the call.

Remote hiring

Problem: Some organisations report candidates using face or voice swaps in video interviews. How it's applied: Live identity checks at key stages, in-person or verified final interviews, and consistency checks across the hiring process. Outcome: Lower risk of hiring someone who is not who they claim to be, especially for roles with access to sensitive systems. The checks add little friction for genuine candidates, who expect identity verification at some point anyway.

Common Deepfake Defense Mistakes

Organisations often respond to deepfakes in ways that feel decisive but leave the real gaps open.

Buying a detector and calling it done

A single detection product, however good its benchmark, will miss output from generators it was not trained on and degrade on compressed or live media. Teams that treat its verdict as final end up trusting a fake that scored "real". Detection belongs inside a layered process, not in place of one. Ask any vendor how accuracy holds up on new generators and live calls before relying on it.

Keeping voice or video as a sole authentication factor

Help desks that reset passwords for a familiar voice, and finance teams that act on a video call from "the CEO", are relying on exactly what deepfakes forge. Changing the process is cheaper and more reliable than trying to detect every clone. A callback to a number already on file defeats a perfect voice clone for almost no cost.

Training on the concept instead of the scenario

Telling staff that "deepfakes exist" changes little. Without concrete scenarios, such as a call from a cloned CFO asking for an urgent transfer and secrecy, employees do not recognise the pattern when it happens to them. Pressure and urgency are deliberate parts of the attack, so staff need explicit permission to slow down and verify.

Treating missing provenance as proof of a fake

Most legitimate content still carries no provenance signature. Flagging everything unsigned as suspicious overwhelms reviewers and wrongly discredits real media. Use provenance as a positive signal where present, not as a gate, and revisit that stance as adoption grows.

Ignoring the organisation's own exposure

Executives' voices and faces are often available in hours of earnings calls, conference talks and interviews. Organisations that never consider this public footprint are surprised when it is used against them. Know which people are most exposed and protect the processes they can authorise.

Deepfake Defense Best Practices

Organizations don't need to become deepfake forensics experts, but a few practical postures reduce exposure meaningfully:

  1. Treat voice and video as insufficient authentication on their own. Any process that authorizes a payment, a password reset, or a data release based solely on a phone call or video call "sounding like" the right person is exploitable today. Add an out-of-band verification step — a callback to a known number, a code word, a second approver, or behavioral biometrics — for anything above a defined risk threshold.
  2. Establish a verification workflow for executive communications. If your CEO's face and voice are public (earnings calls, conference talks, interviews), assume there's enough footage in the wild to clone both. Build a process for verifying urgent, unusual, or high-stakes requests that doesn't rely on recognizing a voice or face.
  3. Watch for the metadata and behavioral cues that are still hard to fake well. Sudden requests for secrecy or urgency, slightly-off phrasing, requests to switch communication channels mid-conversation — these social-engineering tells often remain even when the media itself is convincing.
  4. Don't over-rely on any single detection tool. Given how quickly generators outpace individual detectors, a defense-in-depth approach — process controls plus technical detection plus provenance where available — holds up better than betting on one detection product catching everything.
  5. Train employees on the specific scenario, not just the general concept. "AI voice cloning exists" is abstract; "a caller who sounds exactly like your CFO may ask you to wire money urgently and tell you not to confirm by email" is actionable.
  6. Agree verification code words for the highest-risk roles. A pre-agreed phrase or challenge known only to a small group gives executives and finance staff a quick way to confirm identity on unexpected calls.
  7. Rehearse the response. Run tabletop exercises for a deepfake-driven payment request or a fake executive video circulating online, so teams know who verifies, who communicates and how fast.
  8. Publish authentic channels and signed media. Tell customers and partners where official announcements appear, and sign your own media with provenance where tools support it, so impersonations are easier to spot.

Layered business defence against deepfakes: never treat voice or video alone as authentication, verify executive requests out of band, train staff on concrete scenarios, then add detection and provenance.

Limitations and Open Questions

Detection is not a solved problem, and it's worth being honest about where it's genuinely stuck rather than overselling current tools.

  • Compression degrades everything. Social media platforms re-encode video aggressively, which destroys many of the subtle artifacts forensic detectors rely on — in both fake and real content, making the signal noisier for everyone.
  • Cross-generator generalization is weak. A classifier trained to detect output from one popular generation tool often performs poorly against a different tool's output, because each generator has its own subtle statistical fingerprint. Detectors need constant retraining as new generation methods emerge.
  • Live, real-time deepfakes raise the stakes further. Real-time face and voice swapping during live video calls — rather than pre-rendered video — is now feasible on consumer hardware, which removes the option of offline forensic analysis before a decision gets made — the same real-time liveness problem now facing online age verification systems.
  • Legitimate uses complicate blanket policies. Voice cloning and synthetic media have real, non-deceptive applications in film dubbing, accessibility tools, and content localization. Detection and provenance systems need to distinguish disclosed, consensual synthesis from deceptive impersonation — a distinction that's about intent, disclosure, and consent over one's own likeness, not the technology itself.
  • Provenance adoption remains partial. Until capture devices, editing tools, and platforms broadly support content provenance standards, the absence of a provenance signature can't reliably be treated as a red flag — too much legitimate content still lacks one.

What to Watch Next

The trajectory worth tracking isn't any single detection breakthrough — it's whether provenance infrastructure gets embedded widely enough to shift the burden of proof. Camera and phone manufacturers building capture-time signing into hardware, platforms displaying provenance information natively, and regulatory requirements around AI-generated content labeling are all moving pieces that matter more, long-term, than any individual forensic technique. In the near term, expect real-time generation to keep improving faster than real-time detection, which keeps the practical burden on process controls — verification steps, out-of-band confirmation, employee training — rather than on any tool catching every fake automatically.

Teams building verification workflows, fraud controls, or AI-security processes around this threat can get hands-on help from Woyce Technologies.

FAQ

How can I tell if a video is a deepfake?

Look for inconsistent lighting or shadows between the face and background, unnatural blinking patterns, blurring or warping at the edges of the face, and audio that doesn't quite sync with lip movements. These cues are getting harder to spot as generation quality improves, so for high-stakes situations, don't rely on visual inspection alone — use out-of-band verification instead.

What's the difference between a deepfake and a cheapfake?

A deepfake uses machine learning to generate or alter media, such as swapping a face or cloning a voice. A "cheapfake" achieves a similarly deceptive effect through simple, non-AI editing — slowing down footage, splicing clips out of context, or basic photo editing. Cheapfakes are often just as effective at deceiving people despite requiring far less technical sophistication.

Can deepfake detection software be trusted?

Detection tools are useful as one signal among several, but no detector catches everything reliably, especially against generation methods it wasn't trained on. Treat detection software output as a probability estimate that informs a decision, not as a definitive verdict on its own. Ask vendors how their tool performs on generators released after its training data, on compressed social media video, and on live calls, since those are the conditions where accuracy usually drops. Pair any detector with process controls such as callbacks and second approvers.

How much footage or audio does it take to make a deepfake of someone?

For face swaps, tools historically needed hundreds to thousands of images per identity, though modern few-shot methods can work from far less. Voice cloning has become especially efficient — some modern systems produce a usable clone from just seconds of clean reference audio, which is why publicly available interviews and earnings calls are a meaningful exposure point for executives.

Are deepfakes illegal?

Laws vary widely by jurisdiction and by use case. Several countries and US states have passed laws specifically targeting non-consensual deepfake pornography, election-related synthetic media, and fraud committed using synthetic media, but general-purpose deepfake creation isn't uniformly illegal — legality often hinges on intent, consent, and how the content is used rather than the technology itself.

Will watermarking solve the deepfake problem?

Watermarking and provenance standards help, but they're not a complete solution. They only work if adopted broadly across capture devices and generation tools, and some watermarking schemes can be stripped by re-encoding or cropping content. Provenance is best understood as raising the cost of deception, not eliminating it. It works best alongside detection tools and process controls, such as callbacks to a known number, that don't depend on spotting a fake.

Do deepfakes only affect video and images?

No — audio-only voice cloning is one of the fastest-growing categories, largely because it requires less source material and less compute than convincing video synthesis. Phone-based scams using cloned voices to impersonate executives, family members, or officials have become a distinct and fast-moving fraud vector separate from video deepfakes. The defense is the same as for video: never treat a familiar voice as authentication, and verify urgent requests out of band.

Conclusion

Deepfakes have moved from research curiosity to a commodity capability. Autoencoders and GANs made face swaps possible, diffusion models made full synthesis convincing, and modern voice cloning needs only seconds of audio. Generation is now cheap and fast, while detection remains slow, resource-intensive, and always one generator behind.

That asymmetry shapes the right response. Forensic and learned detectors are useful signals, but they degrade under compression, generalize poorly to new tools, and struggle with real-time fakes on live calls. Provenance standards like C2PA shift the question from proving something is fake to proving something is real, but they only help once adoption is broad. Legitimate uses such as dubbing and accessibility also mean policy has to focus on consent and disclosure, not the technology itself.

For businesses, the most reliable defence is process: never treat a voice or face as sufficient authentication for payments, resets, or data release, require out-of-band verification for high-risk requests, and train staff on concrete scenarios. If you are designing verification flows or AI-driven fraud controls, talk to our AI and machine learning team.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.