Deepfake is one of the most serious digital threats of the AI era — where anyone's face,
voice, and likeness can be fabricated with just a few seconds of source data.
From videos of politicians saying things they never said, to video calls impersonating
relatives asking for urgent money transfers — Deepfake is changing how we perceive
visual and audio evidence.
This article explains what Deepfake is, how it works, the most common forms it takes,
its real-world harms in Vietnam, and how to protect yourself effectively in the age of AI.
If you are new to foundational AI concepts, read What is AI Agent?
first to understand the broader AI landscape before diving into Deepfake —
a specific application of deep learning with direct consequences for security and trust
in the digital society.
What Is Deepfake?
Deepfake is AI-synthesized content — including video, audio, or images —
in which a real person's face, voice, or actions are replaced or fabricated so
convincingly that they are difficult to distinguish from reality with the naked eye.
The term "Deepfake" combines two English words: "deep learning" and "fake,"
reflecting both the machine learning technology that underpins it and the deceptive
nature of the content it produces.
The history of Deepfake begins in 2017, when an anonymous user with the handle
"deepfakes" posted face-swap videos of celebrities on the Reddit community r/deepfakes.
The technique at the time relied primarily on GANs — Generative Adversarial Networks —
to automatically learn how to swap one person's face onto another.
Although the quality was crude and easy to spot, the sudden appearance of this technology
immediately triggered waves of concern from security researchers, legislators,
and the broader internet community.
Within just a few years, Deepfake advanced dramatically with the emergence of diffusion
models in the early 2020s.
Unlike GANs, which required large training datasets and often left characteristic artifacts,
diffusion models produce smoother, more realistic content and require significantly less
input data — sometimes just a handful of photos or a few seconds of audio are enough to
create a frighteningly convincing Deepfake.
This pace of progress genuinely alarms cybersecurity experts, because the gap between
high-quality Deepfakes and public access to the technology is narrowing rapidly.
The scope of Deepfake extends far beyond face-swap video.
Synthetic audio (voice cloning), fully synthesized portrait images, text impersonating
someone's writing style, and even real-time video calls with a fake face all belong to
the modern Deepfake ecosystem.
The boundary between real content and AI-synthesized content is blurring at an alarming
rate, posing unprecedented challenges for individual users, organizations,
and government regulators alike.
How Deepfake Works
At a conceptual level, most first-generation video Deepfakes relied on the GAN
(Generative Adversarial Network) architecture.
Two neural networks operate in adversarial fashion: the Generator creates fake content
that tries to look real, while the Discriminator tries to tell real from fake.
These two networks train in parallel and continuously — the Generator improves to fool
the Discriminator, and the Discriminator improves to detect the Generator —
until the generated content reaches a convincing enough threshold.
A typical face-swap pipeline operates through three main conceptual stages:
- Face detection and alignment: the system locates the face in each video frame
and normalizes size and angle to ensure consistency across frames.
- Swapping: the source face (the person being impersonated) is mapped and
transformed onto the target face through a specialized encoder-decoder network
trained to learn this mapping in fine detail.
- Blending: the result is integrated into the original video so that lighting,
skin tone, and movement match the surrounding environment as naturally as possible.
Voice cloning operates on a different principle but also relies on deep learning.
The model learns to analyze and encode the unique acoustic features of a specific voice —
including rhythm, fundamental frequency, phoneme articulation, and timbre.
The system then synthesizes any text using those features, reconstructing the original voice.
Modern systems need as little as 3 to 10 seconds of source audio to produce a voice copy
usable in real time, opening frightening possibilities for impersonation over phone calls
and video calls.
Beyond GANs and diffusion models, other techniques also contribute to the modern Deepfake
ecosystem. Neural Radiance Fields (NeRF) allow 3D scene synthesis from 2D images,
enabling video generation of real people from multiple angles without any real footage.
Transformer-based generative models — the foundation of systems like GPT and DALL-E —
are being applied to video and audio synthesis, delivering better temporal consistency
and finer content control than earlier architectures.
An important note: this article does not provide specific instructions on workflows
or tools for creating Deepfakes.
The purpose of this section is to help you understand the technology at a conceptual level —
so you can recognize the genuine threat and apply more effective countermeasures in
everyday life.
Three Common Types of Deepfake
Video Deepfake is the most common and widely discussed form in mainstream media.
The two primary techniques are face-swap — fully replacing one person's face with
another person's on a video — and lip-sync manipulation — keeping the face intact
but altering mouth movement so the person appears to "say" what the attacker wants.
Video Deepfakes appear in both entertainment and malicious contexts, from humorous
celebrity face-swap clips to fabricated videos of politicians allegedly endorsing
extreme positions they never actually expressed.
Audio Deepfake (voice cloning) is increasingly dangerous precisely because of its
invisibility. Without an image to observe, listeners rely solely on hearing and the
familiarity of a voice to judge authenticity — which is exactly the most exploitable
vulnerability.
Fraudsters can clone a CEO's voice to call the accounting department and request an urgent
wire transfer, or impersonate a child calling from abroad to ask parents for money.
Particularly dangerous is the scenario where the attacker combines a fake voice with a
spoofed phone number (SIM swap) to create a comprehensive impersonation attack that is
very difficult to detect in the moment of a manufactured emergency.
Additionally, real-time voice changers — software that transforms the speaker's voice
live during a video call — also belong to this category, allowing fraudsters to completely
alter their voice while speaking directly to a victim.
Text-based AI impersonation is the least discussed form but equally dangerous.
AI can learn a person's writing style, word choices, emoji habits, and even characteristic
typos, then impersonate them in emails, social media messages, or online forums.
Combined with personal data harvested from data breaches, an attacker can build a highly
convincing fake digital identity capable of executing sophisticated social engineering
attacks targeting the victim's colleagues, friends, or clients.
The Harm Deepfake Causes
Political misinformation is one of the most far-reaching and hardest-to-control harms
of Deepfake.
Fabricated videos of politicians making extreme statements, confessing to corruption,
or announcing shocking policies can spread virally on social media — especially during
the sensitive period before elections or during national crises.
In a world where information travels faster than ever through sharing platforms, a
cleverly crafted Deepfake video can reach millions of views before fact-checking
organizations have time to analyze it.
Financial fraud through Deepfake is causing serious economic damage at the enterprise
scale.
The most prominent example is CEO fraud — attackers create a Deepfake video of a CEO
or CFO and conduct a video call with the head accountant, pretending to request an urgent
wire transfer for a confidential deal.
A well-known 2024 attack in Hong Kong caused a financial firm to lose $25 million USD
when attackers used Deepfake to simultaneously impersonate multiple senior executives in
a single fake video call, creating the illusion of a completely real board meeting.
Harassment and non-consensual intimate imagery (NCII) is the harm with the most
direct and profound impact on individuals.
Non-consensual Deepfake images and videos cause severe psychological damage, destroy
personal relationships, and ruin the careers of victims.
Women, celebrities, and public figures are the most frequently targeted groups,
though in reality anyone with enough face photos or public videos online is a
potential victim of this type of attack.
Reputation damage through Deepfake can devastate an individual's or organization's
career and life within hours of viral spread.
Even when a fabricated video is later exposed and confirmed as a Deepfake, reputational
damage is often impossible to fully recover — particularly when the fake content
has had time to spread widely.
The "where there's smoke there's fire" effect means suspicion lingers in the minds of
some members of the public, and cognitive psychology research shows that false information
received first influences how people process later corrections — even when those
corrections are clear and credible.
Shocking or emotionally charged fake content spreads far faster than factual corrections —
this is the information asymmetry that Deepfake victims face on deeply unequal terms.
How to Identify a Deepfake
Although increasingly sophisticated, Deepfakes still often leave characteristic telltale
signs that attentive viewers can detect.
Abnormal blinking is one of the most reliable indicators for first-generation
Deepfakes — synthetic models frequently struggle to accurately simulate the frequency,
duration, and natural opening-and-closing of human eyes.
Eye movement that is too sparse, too frequent, irregular, or mismatched with facial
expression are all warning signs worth noting when watching a suspicious video.
Blurry or uneven face edges most commonly appear at the boundary between the swapped
face and the rest of the head — particularly in the hairline, ears, chin, and neck area.
When zooming in on these regions in a suspicious video, you can often see mismatched pixels,
abnormal blurring, or unnatural color transitions between the face and the background.
Inconsistent lighting is another reliable indicator: light hitting the face does not
match the actual light sources visible in the scene, or shadows are inconsistent with
the direction of ambient light.
Stiff or unnatural neck and shoulder movement typically reflects the limitations of
Deepfake models in synthesizing full-body motion outside the face region they focus on.
Audio-to-lip-movement lag — even just a few milliseconds — is one of the most
recognizable signs in Deepfakes that combine both audio and video.
Uneven background blur around the face outline — especially visible in the hair or
ear region — also reveals an imperfect artificial compositing process.
However, it is important to be honest about reality: modern Deepfakes based on current
diffusion models are increasingly effective at eliminating the artifacts listed above.
Detecting Deepfake with the naked eye is becoming less reliable each year,
and anyone who is confident they can spot high-quality Deepfakes through pure observation
may be underestimating the capability of current technology.
Combining manual observation with specialized detection tools is the most sound strategy
available today.
Current Deepfake detection tools fall into two main groups by audience:
tools for individuals (free, easy to use, no installation required) and platforms for
enterprises (paid, with API access, high-volume processing, system integration).
Each group has its own strengths, and the choice depends on scale of need and technical
resources available.
Microsoft Video Authenticator analyzes each video frame to detect artificial blending
at the pixel level, producing a confidence score for each frame.
Microsoft developed this tool as a direct response to concerns about Deepfakes in election
campaigns and political information, targeting fact-checking organizations and
investigative journalists who need to verify the authenticity of video evidence.
Deepware Scanner (deepware.ai) is a free solution for individuals, allowing video
upload directly through the browser and returning analysis results within minutes —
no software installation or deep technical knowledge required.
The process is simple: visit deepware.ai, upload the video you want to check,
wait for the system to process it, and read the confidence score — the higher the
Deepfake probability score, the more suspicious the content, warranting further verification.
Intel's FakeCatcher uses a unique detection method based on photoplethysmography
(PPG) — a technique that analyzes subtle circulatory signals hidden in videos of real
people, specifically the very slight color changes in skin that correspond to heartbeat
and breathing rhythm.
Since current AI synthesis models cannot perfectly recreate this physiological signal,
FakeCatcher can detect Deepfakes even when typical visual artifacts have been well handled.
Hive Moderation is a platform for enterprises and large organizations, providing
an API for high-volume processing of video, image, and audio content with the ability
to integrate into automated moderation pipelines.
The general workflow for most Deepfake detection tools is: upload content → system
extracts features → ML algorithm classifies → returns a probability score of being
a Deepfake.
An important principle: never fully trust a single result from a single tool —
verify using multiple tools and combine with contextual judgment for a comprehensive
assessment.
Deepfake in the Vietnamese Context
Vietnam is among the countries significantly affected by the wave of Deepfake-based fraud
targeting everyday users in their daily lives.
The most common form is fake video calls impersonating relatives — fraudsters use
real-time Deepfake to pose as children, siblings, or friends abroad,
constructing a fictional emergency such as an accident, kidnapping, sudden debt,
or urgent need for money to resolve a legal matter, and demanding an immediate transfer.
Because the video call includes both images and audio matching the face and voice of
someone familiar, victims typically believe the caller immediately without checking further
or consulting others.
Bank transfer fraud combining AI voice and Deepfake video is rising sharply.
According to information from the Department of Cybersecurity and High-Tech Crime
Prevention (A05) of the Ministry of Public Security, the number of online fraud cases
involving AI and Deepfake technology increased significantly from 2023,
with estimated total damages reaching hundreds of billions of dong annually.
Many victims only realize they have been defrauded after completing the transaction,
and recovering money lost in high-tech fraud cases is nearly impossible in practice.
Deepfake KOL and celebrity videos in fabricated advertisements are also a growing
concern.
Many singers, actors, and KOLs with large followings have had their faces and voices
used without consent to promote financial products, high-return investment schemes,
unverified health supplements, or Ponzi-style investment scams — without their knowledge
or agreement.
Consumers trust familiar faces and are easily persuaded to participate without careful
verification, leading to financial losses and eroding trust in the digital environment.
Legal Framework and Self-Protection
Regarding the legal framework in Vietnam, Article 16 of the Cybersecurity Law 2018
prohibits the posting of false, distorted, or defamatory information intended to harm the
legitimate rights and interests of organizations and individuals in cyberspace.
Decree 13/2023/ND-CP on personal data protection includes provisions on the
unauthorized use of biometric information — an important legal foundation directly
relevant to Deepfakes that use someone's face and voice.
However, there is currently no specific Deepfake regulation in Vietnam comparable to the
EU AI Act; this gap is being studied for supplementation as technology evolves and the
number of violations continues to rise.
At the international level, the EU AI Act 2024 is the most comprehensive and pioneering
legal framework for AI governance currently in existence.
The Act requires mandatory disclosure for all realistic AI-synthesized content,
including Deepfakes created for legitimate purposes such as art or education.
Major platforms including YouTube, Meta, and TikTok have also deployed their own policies
requiring users to self-declare and label AI content, while proactively removing
non-consensual Deepfakes when reported by victims or moderation teams.
To protect yourself effectively, apply the following practical rules in everyday life:
- Verify through an independent channel: when you receive an urgent video call requesting
a money transfer — even from someone you know, even with clear images and voice —
end the call and dial back using a phone number you already have saved, or contact
another family member to confirm.
- Set a family codeword: agree with family members on a special word or phrase that
only your family knows, to use as an authentication code in emergency situations
requiring quick identity confirmation.
- Limit public data: think carefully before sharing clear face photos, videos,
or voice recordings publicly — these are the raw materials fraudsters need to create
a Deepfake targeting you.
- Be suspicious of urgency pressure: any situation that creates pressure to "transfer
immediately, don't tell anyone" is a red flag requiring you to stop and verify —
regardless of how convincing the person looks and sounds.
- Enable two-factor authentication: activate 2FA or MFA
on all important accounts to minimize the risk of account takeover being used as a
launchpad for subsequent impersonation attacks targeting your relatives and colleagues.
Legitimate Deepfake: Ethical and Creative Boundaries
Not every Deepfake application is negative — and this is important context for understanding
the complete picture.
In film and media, de-aging techniques and digital resurrection have been used in many major
Hollywood productions, saving significant production costs and enabling stories that could
not be told otherwise.
In education, Deepfake has been used to recreate important historical figures, making
history lessons more vivid and accessible for students.
In the field of accessibility, Deepfake is opening new doors.
Automatic language dubbing technology — syncing lip movement to a new language rather than
simply overdubbing — can break down language barriers for educational and health information
content.
People who have lost their voice to illness can also use voice cloning to preserve and reuse
their own voice — a deeply humane application of technology that is often unfairly
stigmatized.
The boundary between legitimate Deepfake and ethical or legal violation is defined by
two factors: consent from the person whose image or voice is used, and disclosure
that the content is AI-synthesized.
When both conditions are met, Deepfake becomes a valuable creative tool.
When either is violated — particularly when consent is absent — that is when technology
becomes a harmful weapon.
Deepfake and the Future of Digital Trust
The explosion of Deepfake raises a philosophical question deeper than any purely technical
one: can we continue to trust what we see and hear in the digital world?
Trust in visual evidence — "seeing is believing" — has long been the foundation of
journalism, legal systems, and personal communication in human society.
Deepfake is systematically eroding that foundation, creating what researchers call
the "liar's dividend": even genuine content can be denied by simply claiming it is
a Deepfake.
Cybersecurity researchers are racing against Deepfake creators in an adversarial loop
without end.
Every time a new detection method is published and deployed, generative models are
retrained to evade that specific detection approach.
The 2024–2025 generation of diffusion models has already begun producing content that
bypasses many detectors built for earlier GAN-based methods.
This trend shows that technical detection will become progressively harder, and long-term
solutions must include raising public awareness, strengthening legal frameworks,
and developing content provenance authentication standards.
One promising approach is Content Provenance — proving the origin of content.
The Coalition for Content Provenance and Authenticity (C2PA) — a coalition including
Adobe, Microsoft, Google, Intel, and many major organizations — is building technical
standards to embed digital signatures into content at the point of creation (camera,
microphone, editing software), enabling chain-of-custody verification from original
source to distribution.
When this standard is widely deployed, content lacking provenance proof will by default
be treated with greater suspicion — this may be the most important paradigm shift in
the fight against Deepfake in the coming decade.
At the individual level, the most important thing is not to become a technical expert
in Deepfake detection — that is increasingly beyond the reach of ordinary users and
impractical.
What matters is building a habit of verification: not making important decisions
(money transfers, sharing sensitive information, changing behavior) based on a single
source in a digital environment, no matter how convincing that source looks and sounds.
Critical thinking applied to digital content — similar to the critical thinking about
news and advertising we learned in the analog world — is the most durable shield
against the Deepfake threat in the long run.
At the organizational level, businesses need to build additional identity verification
processes for high-value or sensitive transactions.
No video call alone — regardless of how the caller looks and sounds — should be the sole
evidence sufficient to authorize a large wire transfer or disclose critical confidential
information.
Multi-step verification procedures and out-of-band confirmation codes (through a
completely different channel from the one making the request) must become the minimum
standard in risk management for any modern organization.
Training employees on Deepfake and running simulated attack scenarios is a worthwhile
defensive investment — far cheaper than the damage caused by a successful CEO fraud attack.
Key Takeaways
Deepfake is a dual-use technology: it has legitimate and beneficial applications in
entertainment, education, and accessibility, but it is simultaneously being exploited for
fraud, propaganda, and harassment at a scale unprecedented in the history of media.
Understanding what Deepfake is, how to recognize it, and how to protect against it not
only protects yourself but also helps protect the people around you — especially older
family members who have less exposure to AI technology information and are more vulnerable
to new forms of fraud.
The AI world is changing rapidly and Deepfake is just one of many challenges that
the development of artificial intelligence poses for society.
Learning more about the broader AI landscape — from beneficial applications like
AI Agent to security and ethical concerns — helps you build a
solid foundation of awareness to navigate the digital world with greater confidence
and safety in the years ahead.
Alongside that, strengthening personal account security through 2FA
and MFA is the most practical and immediately actionable step today
to reduce the risk of becoming a victim of impersonation and digital fraud attacks.
Remember three golden principles when facing suspicious digital content:
- Stop: do not act immediately when you feel time pressure or strong emotional agitation.
Urgency pressure is precisely what fraudsters want to create to prevent you from thinking
and verifying.
- Verify: call back through an independent channel using a phone number you already
have saved, ask another person, research through multiple sources before making any
decision with financial consequences or that affects others.
- Report: if you discover a fraudulent Deepfake, report it to the authorities
(A05 — Ministry of Public Security), to the social media platform where the content
appeared, and warn people around you to prevent additional victims in the community.
AI technology, including Deepfake, will continue to evolve without pause in the years
ahead.
But with the right awareness, systematic verification habits, and the support of
increasingly refined detection tools, we are fully capable of navigating the AI
information environment safely and responsibly — protecting ourselves, our families,
and our communities against the threats that this technology is and will continue to create.
What is AI Agent?
What is MFA?
What is 2FA?