A Taipei Fraud Ring Trained an AI on 22 of Its Own Employees So Its Men Could Phone Victims as Women — and 57 People Are Now Charged

The 22 cloned voices belonged to staff on the payroll, not to anyone famous. That one detail means no reference recording of the speaker exists anywhere — so the detection test everyone reaches for cannot be run at all.
By Sukrit Bhatia
l
31
 min read
What are deepfakes — business risk overview article
Table of Content
No items found.

The Taipei District Prosecutors' Office indicted 57 people on 2 September 2026 over a romance-fraud operation that ran from 2022, took at least NT$900 million and reached more than 20,000 people. The ring employed its own software engineer to build an AI voice-changing program, and trained it on the voices of 22 female employees from its own public-relations team so that male operators could hold live phone conversations with victims as young women who did not exist.

That detail inverts the detection problem. Almost every synthetic-voice case in the news involves impersonating a real, recorded person, which means a defender can compare the suspect audio against a reference. Here the personas were manufactured from insider voices that were never public, so no reference recording exists anywhere — and speaker comparison cannot be run at all, rather than merely returning a wrong answer.

What follows is the case as charged, an accounting of which figures in it are reported and which are calculated, why voice conversion is the harder detection problem, and which Taiwanese instrument reached this ring while the one written about deepfakes missed it entirely.

  • 57 people were indicted in Taipei on 2 September 2026 over a romance-fraud ring running since 2022, charged under the Fraud Crime Hazard Prevention Act, the Money Laundering Control Act and Criminal Code provisions on group internet fraud. At least NT$900 million was taken from more than 20,000 people.
  • The voices were the ring's own staff. An in-house engineer trained a voice-changing model on 22 female employees from the public-relations team, so operators could speak as invented women on live calls.
  • No reference recording exists, so speaker comparison is unavailable — not inaccurate, unavailable. Only synthetic-speech detection on the signal itself remains, and callback rules have no genuine number to call.
  • Both headline figures are floors, so the NT$45,000 per-victim average derived from them is bounded in neither direction. Only the per-defendant and per-voice ratios hold: at least 351 victims per person charged, at least 909 per donated voice.
  • Detection of converted speech is weaker than product claims suggest. In VoxENES 2026 the best of eight detectors managed 28.98% EER, five were at or below chance, and one fell to 48.4% EER under 64 kbps MP3.
  • Taiwan's deepfake provision did not apply. Article 31(4) of the Act attaches to advertisements; this ring used private calls, so the AI features as method rather than as an element or aggravation.

What did Taipei prosecutors charge on 2 September 2026?

The Taipei District Prosecutors' Office indicted 57 people over a romance-fraud operation that had been running since 2022, took at least NT$900 million (about US$28.3 million) and reached more than 20,000 people. The charges are brought under the Fraud Crime Hazard Prevention Act, the Money Laundering Control Act and Criminal Code provisions covering internet fraud committed by three or more people. Prosecutors are seeking at least 25 years for the leader, Huang Chien-hao, and 18 years for his wife, Hsu Hsiang-chi.

The detail that separates this case from every other voice-fraud story of the past two years is where the voices came from. The ring employed its own software engineer, surnamed Chung, who built it an AI voice-changing program. That model was trained on the voices of 22 female employees from the ring's own public relations team — staff on the payroll, not victims and not public figures — and it let male operators, and by some accounts older female operators too, hold phone conversations with targets as young women who did not exist.

That inversion is the whole reason this case matters to anyone building detection. Almost every synthetic-voice incident that makes the news involves impersonating a real, identifiable person: a chief executive, a head of state, a parent. Here the ring impersonated nobody. It manufactured personas from consenting insider voices, which means there is no reference recording of the “real” speaker anywhere in the world — and the detection methods that work by comparing a suspect call against a known voice have nothing to compare it to.

ElementAs reportedSource
Defendants57 people indictedFocus Taiwan
Prosecuting officeTaipei District Prosecutors' OfficeTaipei Times
ProceedsAt least NT$900 million (about US$28.3 million)Focus Taiwan
VictimsMore than 20,000 peopleTaipei Times
PeriodOperating since 2022Liberty Times
StatutesFraud Crime Hazard Prevention Act; Money Laundering Control Act; Criminal Code internet fraud by three or more personsFocus Taiwan
RingleadersHuang Chien-hao (25+ years sought); Hsu Hsiang-chi (18 years sought)Newtalk
In-house engineerSurnamed Chung — built the voice-changing software, a chat assistant and a performance-management systemNewtalk
Voice training data22 female employees' voices from the ring's public-relations teamLiberty Times
SeizuresFerrari and McLaren supercars, 47 luxury watches, NT$300 million in real estate, bank funds, 67 computers holding the AI softwareTaipei Times

Table 1: What the Taipei District Prosecutors' Office is reported to have charged, 2 September 2026.

Which figures in this case are reported, and which are calculated?

NT$900 million and 20,000 victims will be divided into each other within a day of this indictment being reported, and the resulting NT$45,000 average will travel further than any other number in the case. It is worth being precise about what that figure is, because it is not an estimate of the average loss.

Prosecutors said at least NT$900 million and more than 20,000 people. Both are lower bounds. Dividing one lower bound by another produces a quotient that is bounded in neither direction: if the true total is higher, the average rises; if the true victim count is higher, it falls. Nothing in the reporting tells you which effect dominates. The number is arithmetically correct and epistemically empty, and it is about to be quoted as a finding.

Two derived figures do survive, and they survive for a structural reason: their denominators are exact counts rather than floors. There were exactly 57 people indicted and exactly 22 voices in the training set. A floor divided by an exact number is still a floor, so both results can only move upward as the true totals come in.

Derived figures · which way can each one move?

Two of the reported totals are floors, so most figures computed from them cannot be pinned down — including the per-victim average.

Prosecutors reported at least NT$900 million and more than 20,000 people. Divide one lower bound by another and the result is not an estimate of anything: it can move either way. Two derived figures survive, because their denominators are exact counts.

≥ NT$900m taken · floor > 20,000 victims · floor 57 indicted · exact 22 donor voices · exact from 2022 · approximate
Derived figureArithmeticDirection the true value can travel
NT$45,000
Mean loss per victim — the figure most likely to be quoted from this case
900,000,000 ÷ 20,000
Indeterminate

A larger true total pushes it up; more true victims push it down. Floor ÷ floor bounds nothing.

≥ 351
Victims per person indicted
20,000 ÷ 57
Genuine floor

Numerator is a floor, denominator an exact charge count. The true figure can only be higher.

≥ 909
Victims reached per donated voice — the industrialisation number
20,000 ÷ 22
Genuine floor

Same structure: floor over an exact count. 22 voices carried at least 20,000 approaches.

~NT$18.8m
Mean take per month across the run
900,000,000 ÷ ~48 months
Indeterminate

A floor total pushes it up; a longer true run pushes it down. Only the start year is reported.

22×–45×
Lifetime loss against the first purchase (NT$1,000–2,000)
45,000 ÷ 1,000 … 2,000
Indeterminate

Built on row one, so it inherits its indeterminacy. Derived-from-derived compounds, never cancels.

Why the two survivors are the interesting ones

The only figures that hold their direction are per-defendant and per-voice — and those are exactly the industrialisation numbers. At least 351 victims for every person charged, and at least 909 for every voice the model was given, are claims that can only strengthen as the true totals come in. The headline average cannot be defended in either direction, which is why this analysis states it with its arithmetic attached rather than as a finding.

Inputs are as reported from the Taipei District Prosecutors' Office indictment by Focus Taiwan (CNA) and the Taipei Times, 2–3 September 2026: at least NT$900 million, more than 20,000 people, 57 indicted, 22 employee voices used to train the model, operating since 2022. Every figure in the left column is calculated by DuckDuckGoose from those inputs and was not reported by prosecutors or by any outlet. The ~48-month duration assumes a mid-2022 start and an indictment in September 2026.

Derived figureArithmeticStatus of inputsDirection it can moveWhat may honestly be said
NT$45,000900,000,000 ÷ 20,000Floor ÷ floorIndeterminate — either wayQuote only with both bounds attached. It is not an average loss.
≥ 351 victims per person indicted20,000 ÷ 57Floor ÷ exact countUpward onlyA genuine floor. Safe to state as “at least”.
≥ 909 victims per donated voice20,000 ÷ 22Floor ÷ exact countUpward onlyA genuine floor, and the clearest measure of what the model bought them.
About NT$18.8 million per month900,000,000 ÷ ~48 monthsFloor ÷ approximate durationIndeterminateOnly the start year is reported. Treat as an order of magnitude.
22× to 45× escalation45,000 ÷ 1,000 to 2,000Derived ÷ reported rangeIndeterminateInherits the indeterminacy of row one. Derived-from-derived compounds.
About 33% of the floor recovered300,000,000 ÷ 900,000,000Partial seizure ÷ floorIndeterminateReal estate only; cars, watches and bank funds are unquantified in the reporting.

Table 2: Every figure this article calculates rather than quotes, with its arithmetic and the direction the true value can travel. None of these appears in the indictment reporting.

The two survivors are the interesting ones, and not by accident. At least 351 victims for every person charged, and at least 909 for every voice the model was given, are the industrialisation numbers — and they are the two figures in this case that can only get worse for the defence as more victims come forward. The headline average, which is the one that will be repeated, is the one that cannot be defended in either direction.

This is a general habit worth adopting rather than a quirk of this case. When an enforcement figure is published as a floor — and enforcement figures usually are, because prosecutors charge what they can prove — anything computed from it inherits that floor, and the direction of the inheritance depends on where in the fraction it sits. The same trap sat in the numbers around ASIC’s deepfake investment-scam takedowns and in every “losses topped X” headline since.

Why 22 insider voices break the detection method most people reach for

Ask how you would test whether a suspicious call was synthetic and most answers arrive at the same operation: get a recording of the person the caller claims to be, compare the two, and report how close they are. That operation takes two inputs. In this case the second one does not exist.

The women on those calls were never real. Their photographs were taken from elsewhere, their names were invented, and their voices were composites derived from 22 employees whose own voices were never published and are not in any enrolment database a bank or a platform could query. There is no “real Ms Lin” to record. A speaker-comparison test does not fail here in the sense of returning a wrong answer — it cannot be run at all, because one of its operands is missing.

Speaker comparison · operand availability

Voice-comparison detection needs a second recording. In the Taipei case, no such recording exists.

Every method that decides whether a call is fake by comparing it to a known voice takes two inputs. The Taipei ring's model was trained on 22 of its own employees, whose voices were never public and whose personas were invented — so the second input is missing, and the operation cannot be run at all.

Compare suspect call  →  reference of the claimed speaker  →  verdict

Case typePublic-figure cloneFabricated video of a head of state or celebrity
Reference available?Yes — abundantYears of broadcast and archive audio of the real person
Comparison runsBoth operands present
Case typeExecutive impersonationA cloned CFO on a payment call
Reference available?Yes — enrollableThe real executive can be recorded on demand
Comparison runsBoth operands present
Case typeInsider voice conversionTaipei: 22 staff voices, worn by operators as invented women
Reference available?None in existenceThe persona was never a person; the 22 donors were never public
UndefinedNothing to compare against

What is left when the reference is gone

One question survives: not “is this the person they claim to be?” but “was this audio generated at all?” That verdict is reached from the signal on its own — the artefacts a conversion model leaves in the waveform — and it needs no second recording, which is precisely why it is the only route open in a case built on manufactured identities.

Facts from the Taipei District Prosecutors' Office indictment as reported by Focus Taiwan (CNA) and the Taipei Times, 2–3 September 2026. The three case types are DuckDuckGoose's framing of the detection problem, not a categorisation used by prosecutors. “Reference available?” describes whether a recording of the claimed speaker can be obtained — not whether any investigator sought one.

What remains is a different question, and it is worth stating the two side by side because they get conflated constantly. Speaker verification asks whether this audio is the person it claims to be. Synthetic-speech detection asks whether this audio was generated at all. The first needs a reference; the second needs only the signal in front of it. Against manufactured identities the first is unavailable by construction, and the second is the only route open.

Attack shapeWho is impersonatedReference recording available?What detection can rest on
Public-figure cloneA real, widely recorded personYes — abundant archive audioSpeaker comparison and synthetic-speech detection both available
Executive or family impersonationA real person known to the victimYes — the real person can be recorded on demandSpeaker comparison, plus out-of-band callback to the known number
Insider voice conversion (this case)Nobody — the persona is manufactured from staff voicesNo — none exists anywhereSynthetic-speech detection on the signal alone

Table 3: Three ways a synthetic voice reaches a victim, and what each leaves available to a defender.

The practical consequence for anyone running controls: a callback rule is useless against a manufactured persona, because there is no genuine number to call back. The advice that works against voice cloning aimed at a real colleague — hang up and dial the number you already have — has no object here. That asymmetry is the same one that makes training-based defences fail against synthetic media: the instruction assumes a real referent exists.

Voice conversion is not the voice cloning in every other headline

Two different technologies get filed under “AI voice” and they leave different evidence. Text-to-speech synthesis builds an utterance from written text in a target voice; the operator types, and the model speaks. Voice conversion takes an existing human utterance and re-timbres it, so a live speaker's words, timing, breathing and emotional delivery survive while the apparent identity changes.

The reporting on this case describes conversion, not synthesis. The program is consistently called a voice-changer, its purpose was to let staff speak as someone else on live calls, and Liberty Times describes 22 employees' voice samples being fed into a chat system so that male employees and older women could talk to victims in young female voices. Nobody was typing sentences for a model to read aloud; people were holding conversations while wearing a voice.

That distinction matters operationally. Conversion preserves the parts of speech that human listeners use to judge sincerity — hesitation, laughter, the catch in a voice when a story lands — because a real person is generating them in real time. It is a far better fit for romance fraud than synthesis, which has to fake spontaneity. The ring did not pick the harder technology; it picked the one suited to a four-year emotional relationship. We have written before about how voice cloning works in practice and how little audio a usable model needs; the 22-voice detail here is a reminder that the constraint is rarely data volume.

It also matters for detection, and not in the defender's favour.

How hard is converted speech to detect on a phone call?

Harder than synthesis, and harder still once telephony has compressed it. This is measured rather than argued. VoxENES 2026, a benchmark published on 13 July 2026 by Aastha Sharma and Guangjing Wang of the University of South Florida, ran eight spoofing detectors against 53,628 audio samples covering ten synthesis methods — seven text-to-speech and three voice conversion — across two languages and ten post-processing conditions.

Three of its findings bear directly on this case. Voice conversion was the harder class overall. Speaker-embedding approaches, which performed respectably on some text-to-speech systems, “degrade substantially on VC” — the measured version of the argument above. And the headline result is sobering for anyone assuming detection is a solved product: the best detector managed 28.98% equal error rate overall, and five of the eight performed at or below chance, at 47% EER or worse.

MeasurementFigureWhy it matters hereSource
Best detector, all conditions28.98% EERThe ceiling, not the floor, on off-the-shelf performancearXiv
Detectors at or below chance5 of 8 (≥47% EER)Most legacy detectors do not generalise to current generatorsarXiv
Hardest voice-conversion system testedNo detector below 41% EER on Seed-VCConversion, the class used in this case, is the weak spotarXiv
Speaker-embedding methods on conversion“Degrade substantially on VC”The identity-comparison family is the wrong tool twice overarXiv
MP3 at 64 kbps, one detector26.7% → 48.4% EERLossy compression at telephone-grade bitrates can erase the cues a detector relies onarXiv
White noise at 10 dB SNROne detector improved to 17.4% EERDegradation is not monotonic; noise can perturb artefacts helpfullyarXiv

Table 4: Selected measurements from VoxENES 2026 (53,628 samples, ten synthesis methods, eight detectors). Equal error rate: lower is better, 50% is a coin flip.

The compression result deserves a moment. A detector that fell from 26.7% to 48.4% EER under 64 kbps MP3 is a detector rendered useless by the ordinary act of carrying audio over a consumer channel, before any adversary does anything clever. The authors attribute it to reliance on fine-grained high-frequency spectral structure, which lossy codecs discard. Any claim that a voice detector “works” needs to name the channel it was measured on, which is the same argument we have made about quoted detection accuracy figures generally and about why some generators are harder than others.

None of this means detection is hopeless, and the counter-intuitive white-noise result is a hint as to why: the artefacts are there, and what varies is whether a given model has learned cues that survive the channel. It does mean that a defender who plans to catch this attack class with a single detector, chosen on a benchmark number, is planning badly. Explainability matters more than a headline score when the operating channel differs from the test set, which is the case we have made for explainable detection outputs.

The escalation ladder started at NT$1,000 of fish oil

Operators opened fake female profiles on dating apps and waited to be contacted. What followed was not an immediate request for money but a purchase — health supplements, fish oil, lingzhi mushroom products, a portable charger, at NT$1,000 to NT$2,000. The Liberty Times account calls this a boiling-frog approach, and its function was selection: a target who buys a small unnecessary thing for a stranger has demonstrated both willingness and a working payment method.

From there the requests climbed to hair dryers, air purifiers and mobile phones, and then left goods behind entirely for money: living expenses, medical expenses, help with moving house, lingerie, jewellery, aesthetic procedures. Some accounts describe real-world meetings being arranged to reinforce the relationship's plausibility. The four teams named in the reporting map onto that ladder rather than onto seniority — a name-list team building profiles, a new-customer team making first contact and running the small purchase, a cultivation team holding the relationship and raising the asks, and a returning-customer team working victims who had already paid.

Read alongside the derived figures, the ladder explains the numbers. A floor of 20,000 targets against a floor of NT$900 million is consistent with an operation making most of its money from a small tail of heavily cultivated victims while processing a very large number of people through the NT$1,000 filter. That is the same funnel shape we described in how deepfakes fuel romance and investment fraud, and it is why the per-victim average is not just epistemically weak but conceptually misleading: there is no typical victim in a funnel like this one.

The voice model removed the last headcount limit

Romance fraud has always been labour-bound. The product being sold is sustained attention from a specific person, and until a voice call is involved, one operator can fake that in text at scale. The phone call was the ceiling: it required somebody whose voice matched the photograph, and it required them to be available whenever a particular victim rang.

A conversion model removes exactly that constraint and nothing else. It does not write better scripts or find better targets. It decouples the voice a victim hears from the person speaking, so any operator on shift can take any relationship at any hour, and the roster stops needing to match the personas. At least 909 victims for every voice donated is what that decoupling bought, and it explains a detail that would otherwise look strange: why a fraud ring paid an engineer to build software rather than buying a subscription.

The seizure list supports reading this as a manufacturing operation rather than a crime with a gadget. Sixty-seven computers held the software. The same engineer built a chat assistant and, according to Newtalk and CNEWS, an ERP-style performance-management system that tracked operator output — the tooling of a sales floor, applied to fraud, with monthly review meetings whose punishments for underperformance were violent enough that prosecutors describe them in the indictment. This is the same industrialisation we traced in the synthetic-fraud supply chain and in the Sapphire Network ad operation: the synthetic-media component is one module in a stack that otherwise looks like a badly run company.

Worth stating plainly, because it cuts against how these stories are usually framed: the AI did not make this ring's deception more convincing. Its operators were already convincing, over years, in text and in person. The model made the convincing cheaper to supply.

What the 67 computers say about where the model ran

Sixty-seven machines seized with the AI software on them is an unusual detail to publish, and it carries an inference worth drawing carefully. On-premises hardware at that count is consistent with conversion running locally on operator workstations rather than through a hosted service.

There are good operational reasons a ring would choose that. Real-time conversion on a live call is latency-sensitive, and a round trip to a third-party API adds delay a victim can hear. A commercial provider is also a witness: it holds logs, it can be served with a production order, and it can cut off an account. Local inference leaves the evidence on hardware the ring controls, which is precisely why 67 computers ended up in an evidence room rather than 67 API keys in a subpoena return.

The inference is ours and it is not certain — the reporting says the machines contained the software, not that inference ran on them, and nothing rules out a hybrid arrangement. It is stated here as a reading of the seizure inventory rather than as a finding, and it is testable: a prosecution that details the software's architecture will confirm or kill it. It matters because a defender's leverage differs enormously between the two worlds. If synthetic-voice tooling runs locally on commodity machines, then platform-side and vendor-side controls — the levers most regulation reaches for — have nothing to grip, and the only remaining control points are the channel where the audio arrives and the payment rail where the money leaves.

Where does the reporting disagree, and what should you not repeat?

Coverage of this indictment is unusually consistent on the load-bearing figures and inconsistent on detail. Every disagreement found while reporting this piece is set out below rather than resolved silently, because a reader repeating one of these should know it is contested.

PointOne readingThe otherWhat this article uses
Value of seized real estate and frozen fundsNT$300 million in real estate plus bank funds (Taipei Times)A figure that renders in translation as NT$3 billion (Liberty Times)NT$300 million. The higher figure would exceed the total proceeds and is most likely a units artefact, but it is flagged rather than dismissed.
Whose voices were convertedMale operators speaking as the public-relations women (Focus Taiwan)Male employees and older women converted into young female voices (Liberty Times)Both, attributed. The second is the more specific claim and appears only in Chinese-language reporting.
What the model didImpersonated the public-relations women on calls (Focus Taiwan)Learned the voiceprints of women and took over some conversations (Taipei Times)Described as voice conversion on live calls, which both readings support.
The entry productsFish oil and portable chargers (Focus Taiwan)Lingzhi mushroom products and health supplements (Liberty Times)Both, as examples rather than an inventory. The NT$1,000–2,000 price band is common to all accounts.
Team namesLead generation, new customers, public relations, returning customers (Focus Taiwan)Name-list, new-client, cultivation, repeat-client teams (Newtalk)The Newtalk naming, which is a translation of the Chinese terms, with the roles described rather than the labels relied on.
The engineer's identitySurnamed Chung (Focus Taiwan)Rendered Zhong in Chinese-language reports (Newtalk)Surname only. The two are the same surname under different romanisations; no full name is confirmed across sources.
Duration“Running since 2022” (Liberty Times)No start date given (Taipei Times)From 2022, with the month unreported — which is why the monthly average in Table 2 is marked indeterminate.

Table 5: Points where otherwise-careful reporting differs, and what this article does with each.

One absence is worth naming too. No press release on this case could be found on the Taipei District Prosecutors' Office website while researching this piece, and the office's published press-release listing does not carry it. Every fact above therefore rests on reporting of the indictment rather than on the indictment itself or on a prosecutorial statement read directly. That is a real limitation, and it is the reason the legal analysis below is built on statutes that can be read in full rather than on the charge sheet, which cannot.

Which Taiwanese law reached this ring, and which one missed it entirely?

Taiwan has a deepfake provision in its principal anti-fraud statute, and it would not have touched this operation. That is not a criticism of the drafting so much as a demonstration of where legislatures have assumed synthetic media lives.

The Fraud Crime Hazard Prevention Act, as amended on 21 January 2026, states its purpose in Article 1 as being “to prevent and combat fraud hazards, prevent and stop fraud crimes through the inappropriate use of finance, telecom, and Internet, and protect the rights and interests of victims and citizens”. Its synthetic-media hook sits in Article 31(4), which requires disclosure where an advertisement “uses deep fake technologies or AI-generated individual images”. That is a transparency duty attached to advertising, and this ring did not advertise. It made one-to-one contact on dating apps and moved the relationship to private calls.

The aggravating provisions do not reach it either. Article 44 raises penalties by half for circumstances including using equipment outside the Republic of China to commit crimes against people inside it, and for offences involving minors, people aged 80 or over, or foreign nationals. Committing fraud through a synthetic voice is not among them. The result is visible in the charge sheet: 57 people charged with fraud, money laundering and group internet fraud, with the AI appearing as method rather than as an element or an aggravation.

The same pattern shows in the implementing machinery. The Ministry of Digital Affairs brought four subsidiary regulations into force on 30 November 2024 requiring internet advertising platforms of a certain scale to verify advertisers' and investors' identities, publish transparency reports, run fraud-prevention plans and remove fraudulent advertisements within 24 hours. Every one of those duties attaches to advertising. And when the Executive Yuan approved draft amendments to the Act on 13 November 2025, the four announced priorities were cooperation with financial institutions and virtual-asset service providers, heavier penalties for high-value property crime, better victim compensation, and treating a perpetrator's extravagant spending as a sentencing factor — with no synthetic-media provision among them.

InstrumentWhat it attaches toDoes it reach this ring?Source
Fraud Crime Hazard Prevention Act, Art. 31(4)Disclosure of deepfake or AI-generated imagery in advertisementsNo — no advertising was used, and the provision covers images, not voiceLaws & Regulations Database (Taiwan)
Fraud Crime Hazard Prevention Act, Art. 44Aggravation for offshore equipment, minors, people 80+, foreign nationalsNot on the synthetic-media ground — synthetic voice is not a listed aggravating circumstanceLaws & Regulations Database (Taiwan)
Fraud Crime Hazard Prevention Act, general fraud provisionsFraud committed through finance, telecoms and the internetYes — one of the three statutes chargedFocus Taiwan
Money Laundering Control ActMovement and layering of proceedsYes — charged, and the route by which the supercars and watches were reachedNewtalk
Criminal Code, internet fraud by three or more personsOrganised commission of fraud onlineYes — charged, and the provision carrying the group's scaleTaipei Times
MODA advertising-platform regulations, in force 30 Nov 2024Advertiser identity verification, transparency reports, 24-hour takedownNo — duties attach to advertising platforms; first contact here was on dating appsMinistry of Digital Affairs
Draft amendments approved 13 Nov 2025VASP cooperation, high-value property crime, victim compensation, extravagant spendingPartly — on the money side only; no synthetic-media provisionExecutive Yuan

Table 6: Which Taiwanese instrument reaches which part of this case, read against the statute and the implementing regulations themselves.

The lesson generalises well beyond Taiwan. Deepfake law written in the last three years has mostly been written about broadcast harms — election content, scam advertisements, non-consensual imagery — because those are the harms that were visible when the drafting started. A synthetic voice on a private call between two people is not a publication, has no platform intermediary to place a duty on, and leaves no artefact for a regulator to order removed. We traced the same gap in the 2026 regulatory picture and in the Warsaw litigation over deepfake scam ads, where the entire fight was about who must take an advertisement down.

How does NT$900 million compare with Taiwan's national fraud picture?

Against the national trend, this ring was operating in a country that was getting rapidly better at stopping fraud. The National Police Agency's 165 anti-fraud dashboard, as reported on 25 August 2026, shows monthly telecom and internet fraud cases down 43% and monthly financial losses down 71% against August 2024, when the anti-fraud centre was established. Fake-investment fraud fell hardest — cases down 83%, losses down 89%, with daily losses falling from NT$200–300 million to about NT$30 million.

Romance-investment fraud, the category closest to this case, improved least of the four that improved: cases down 40% and losses down 46%, roughly half the rate of the fake-investment collapse. And one category moved the other way, with social-media-post and single-page shopping fraud up 56% in cases and 733% in losses, now about 9% of all telecom fraud losses.

Two readings follow, and they are compatible. The controls Taiwan built after 2024 were aimed at the money — account opening, transfer interdiction, virtual-asset service providers — and they worked spectacularly against fraud whose conversion event is a large transfer to an investment platform. They work much less well against fraud whose conversion event is a sequence of small consumer purchases and personal gifts spread over years, which is exactly what the ladder in this case was built from. NT$900 million accumulated over four years is a rounding error against national monthly losses in the billions, and that is the point: this operation was not big enough to trip a systemic control, 20,000 times over.

What should a platform hosting first contact change?

The dating app is where this attack begins and where it is cheapest to interrupt, and the control that matters is not identity verification at signup. A ring that can staff four teams can pass a document check; several of the 57 were real people with real documents. The signal is behavioural and it is specific to voice conversion.

  • Treat voice as a channel with its own risk posture, not as a trust upgrade. Many platforms treat a completed voice call as evidence a match is genuine, which is precisely the inference this ring monetised. A voice call should raise confidence in liveness, not in identity.
  • Look for many personas behind few operators. The decoupling that makes conversion valuable also makes it visible: shared devices, shared network paths, shift-shaped activity patterns and handovers mid-relationship are all detectable without touching audio content.
  • Score the first small purchase as the qualifying event it is. An off-platform request to buy a low-value physical good early in a conversation is not an incidental oddity; in this case it was the selection filter for the entire funnel.
  • If you analyse call audio, measure your detector on your own codec. The VoxENES compression result makes a benchmark figure obtained on clean audio close to meaningless for a platform running voice over a lossy consumer channel.

The broader point is one we have argued about where detection belongs in a verification stack and about fraud that happens after onboarding rather than during it: a control placed only at the front door cannot see an attack whose whole mechanism unfolds over the following months.

What should a bank or wallet operator change?

The money in this case did not move in one suspicious transfer. It moved as consumer purchases, gifts and small personal remittances, which is the shape most likely to clear every rule an institution has, because each transaction is individually unremarkable and individually authorised.

  • Look for escalation across a relationship, not anomalies within a transaction. A customer whose payments to new counterparties climb from NT$1,500 to NT$150,000 over eighteen months is the pattern; no single payment in that sequence is odd.
  • Do not require a synthetic-media finding before acting. Nobody will ever establish that a particular call in this case was converted, and waiting for that finding is waiting forever. The same conclusion applied to the Bacolod mayor impersonation case, where no recording of the call exists at all.
  • Treat repeat victimisation as a supervised state. A returning-customer team existed because previously defrauded people are the best remaining prospect. An institution that has already reimbursed a romance-fraud loss is holding a high-risk customer, not a closed case.
  • Ask where the counterparty goods are going. Health supplements and consumer electronics bought for a person the customer has never met are a recognisable category, and one that sits below every investment-fraud rule written since 2024.

Institutions running voice channels of their own face the narrower version of the same problem. The relevant question is not whether a caller sounds like the account holder, which converted audio can satisfy, but whether the audio was generated — a distinction we have set out in liveness detection versus deepfake detection and in our survey of detection tooling.

Frequently Asked Questions

How many people were indicted in the Taipei voice-conversion romance fraud case?

The Taipei District Prosecutors' Office indicted 57 people on 2 September 2026. Prosecutors are seeking at least 25 years' imprisonment for the leader, Huang Chien-hao, and 18 years for his wife, Hsu Hsiang-chi.

How much money did the ring take, and from how many victims?

At least NT$900 million, about US$28.3 million, from more than 20,000 people, over a period running from 2022. Both figures are lower bounds as reported, so any average computed from them is not bounded in either direction.

Was the average loss really NT$45,000 per victim?

That figure is NT$900 million divided by 20,000, and it should not be treated as an average loss. Both inputs were reported as floors, so the quotient can move either way as the true totals emerge. The reliable derived figures are the per-defendant and per-voice ratios, whose denominators are exact counts.

What is the difference between voice conversion and voice cloning?

Text-to-speech cloning generates speech from written text in a target voice. Voice conversion re-timbres a live human utterance, so a real speaker's words, timing and emotion survive while the apparent identity changes. This case describes conversion, which suits a sustained relationship better than synthesis.

Why can't investigators simply compare the calls to the real person's voice?

Because there is no real person. The personas were invented and the model was trained on 22 of the ring's own employees, whose voices were never public. No reference recording of the claimed speaker exists, so speaker comparison cannot be run; only synthetic-speech detection on the signal itself remains available.

How reliable is synthetic-speech detection on telephone audio?

Less reliable than product claims imply. In the VoxENES 2026 benchmark the best of eight detectors reached 28.98% equal error rate overall, five performed at or below chance, and one fell from 26.7% to 48.4% EER under 64 kbps MP3 compression — a bitrate ordinary consumer channels use.

Did Taiwan's deepfake law apply to this case?

Not to the synthetic voice. The Fraud Crime Hazard Prevention Act's synthetic-media provision, Article 31(4), requires disclosure of deepfake or AI-generated imagery in advertisements, and this ring used private calls rather than advertising. The charges rest on general fraud, money laundering and group internet fraud provisions.

What did the AI actually buy the operation?

Capacity, not credibility. The model decoupled the voice a victim heard from the person speaking, so any operator could take any relationship at any hour. At least 909 victims were reached for every voice in the training set.

What should organisations take from the seizure of 67 computers?

It is consistent with conversion running locally rather than through a hosted service, which would suit a latency-sensitive live call and leave no third-party logs. That reading is an inference from the inventory rather than a confirmed finding, and it matters because local tooling puts the attack beyond platform-side and vendor-side controls.

Methodology and limits of this analysis

Every source cited was opened and read in full. Facts about the indictment rest on five news reports — Focus Taiwan (CNA), the Taipei Times, the Liberty Times, Newtalk and CNEWS — with the Chinese-language reports carrying detail the English-language ones do not, including the team names, the engineer's internal software and the conversion of older female operators' voices as well as male ones. A second Taipei Times report of 2 September was also read and was consistent with the front-page report cited here; it is not listed separately because every claim it supports is carried by a source that is.

Legal analysis rests on the statute and the implementing instruments themselves rather than on commentary: the Fraud Crime Hazard Prevention Act as published in the Laws & Regulations Database of the Republic of China, the Ministry of Digital Affairs release setting out the four advertising-platform regulations, and the Executive Yuan release on the November 2025 draft amendments. Detection figures are taken from VoxENES 2026 as published.

Four limits are worth stating explicitly. First, no press release on this case could be located on the Taipei District Prosecutors' Office website, and the charge sheet has not been read; everything about the prosecution is therefore reported at one remove. Second, every figure in Table 2 is calculated by DuckDuckGoose from reported inputs and appears in no source — the arithmetic is shown so it can be checked rather than trusted. Third, the reading of the 67 seized computers as evidence of local inference is an inference, labelled as such where it appears. Fourth, DuckDuckGoose has not analysed any audio from this case; no media from it is public, and nothing here should be read as a detection finding about these calls.

Where sources conflict, both readings are carried in Table 5 with the outlets holding them. The NT$300 million versus NT$3 billion seizure discrepancy is unresolved and is flagged rather than silently normalised.

Last update: Q3 2026.

Sources

By Sukrit Bhatia
DuckDuckGoose AI

About the author

By Sukrit Bhatia
DuckDuckGoose AI

Discover the Power of Explainable AI (XAI) Deepfake Detection

Schedule a free demo today to experience how our solutions can safeguard your organization from fraud, identity theft, misinformation & more