Voice AI · Consent, Security & Fraud

AI Voice Cloning for Swiss Business Calls: Consent, Security and the Fastest Way to Damage a Brand

The same four seconds of audio, two different outcomes

A voicemail greeting, a conference panel recording, a LinkedIn video — a few seconds of someone talking is now enough raw material to build a passable clone of their voice. Used on purpose, with that person's knowledge, it can make an AI phone assistant sound like the founder answering calls after hours, or give a brand a consistent voice across every recorded touchpoint. Used by someone else without asking, it is the opening move of a fraud playbook that has already moved millions of dollars out of real companies. The technology behind both outcomes is identical. What separates them is not a technical safeguard but a handful of questions almost nobody asks before pressing “clone”: whose voice is this, who agreed to it, who is allowed to generate speech with it, and who switches it off the day that person leaves.

This is not a hypothetical for anyone shopping for a Swiss AI phone assistant. Some providers already sell voice cloning as a feature. Weissmann AI's own Enterprise tier, for demanding businesses that want a specific person's voice rather than a standard one, lists it among its capabilities, offered on request rather than switched on by default. That is worth stating plainly rather than glossing over, because it means the questions in this article are not academic for anyone evaluating that option — from Weissmann or from any other Swiss provider offering the same thing.

What it actually takes to clone a voice today

Modern voice-cloning systems need surprisingly little source audio. Production-grade services can produce a convincing replica from well under a minute of clean speech, and some produce a rough approximation from a handful of seconds. The source material doesn't need to be secretly recorded — a voicemail greeting, a podcast appearance, an all-hands meeting uploaded to a shared drive, a talk on YouTube are all usable inputs, and none of them were ever meant to be a security asset in the first place.

This matters because “keep your voice private” is not a realistic defence for anyone whose job involves talking to people in public. The practical response has to be about controlling what a clone is allowed to do once it exists, not about hiding a voice that, for most business leaders, was never hideable to begin with.

What $25 million in fifteen transfers actually looked like

In January 2024, a finance employee at the Hong Kong office of the UK-headquartered engineering firm Arup received an email that appeared to come from the company's UK-based chief financial officer, referencing a confidential transaction. He was suspicious enough not to act on the email alone. What changed his mind was a video call: on it, he saw and heard people who looked and sounded exactly like the CFO and several colleagues he recognised. Every one of them was a deepfake, built from publicly available video and audio of Arup's own executives. Over the days that followed, he authorised fifteen transfers totalling roughly US$25 million (about HK$200 million), before the fraud came to light when he called Arup's actual head office to follow up on the “secret transaction.” Hong Kong police confirmed the case in February 2024; Arup itself confirmed it publicly in May 2024. As of the most recent public reporting, the funds have not been recovered and no perpetrator has been identified — this account is drawn from that public reporting, not from a firsthand review of the transaction records.

The lesson isn't “the deepfake was unconvincing, watch more closely next time.” It worked precisely because everything about that call looked and sounded right, voices included. What was missing was a callback to a phone number the employee already had on file, or a second approver who hadn't been on the call. Eyes and ears are no longer reliable evidence on their own for a decision of that size. Voice and video should never be the entire authorisation for moving money or changing access — only a starting point that a separate, independent check confirms.

Why the FBI's warning applies well below CFO level

In December 2024, the FBI's Internet Crime Complaint Center (IC3) issued a public warning describing exactly this category of fraud at smaller scale: criminals generating short AI audio clips of a family member's voice to demand urgent money during a manufactured crisis, using cloned audio to talk their way into someone's bank account, or impersonating a recognisable public figure to solicit payment. None of that requires a boardroom or a CFO — it requires an owner who has ever posted a video greeting on the company website, spoken at a Chamber of Commerce event, or left a warm voicemail message, which describes most Swiss business owners with any public presence at all.

The FBI's advice to individuals — hang up, call back on a number you already have, don't act on a single call alone — is the consumer version of the same callback-verification principle the Arup case illustrates at corporate scale. It is also the cheapest, most reliable defence available to a business of any size, and it costs nothing to write into a procedure this afternoon.

Consent is not a checkbox someone signs once

Cloning an employee's or founder's voice on purpose starts with consent, and real consent is narrower and more ongoing than most businesses assume when they first ask.

  • A specific, limited purpose stated up front — “your voice will be used for the after-hours phone greeting and nothing else,” not a blanket permission to use it however the business later decides.
  • The chance to hear or review actual output before it goes live, not just a description of the plan in the abstract.
  • A genuine ability to withdraw consent later, not only at the moment of signing — a clone, once approved, is not supposed to become permanent by default.
  • For employees specifically, an honest acknowledgement of the power imbalance in the request: a real ability to decline without professional consequence, recorded in writing, not assumed during a busy first week.

Access control: who is allowed to make that voice say something new

Once a clone exists, the safest mental model is to treat it like a credential, not a creative asset sitting in a shared folder.

  • A named, short list of people authorised to generate new speech with a specific person's cloned voice — not “the marketing team,” a list of individuals.
  • Every generation logged — who requested it, when, and for what stated purpose — so an unusual request stands out rather than blending into routine use.
  • Separation between everyday text or chat tools staff already use and the specific ability to produce cloned audio, so a compromised general account doesn't automatically grant voice-generation rights.
  • A default of the smallest set of people who can still do the job, reviewed periodically — not the largest set that seemed convenient when the feature was switched on.

Revocation: what actually happens the day someone leaves

This is the step most policies skip, because it's the one nobody wants to think about while the relationship is still good.

  • Consent ends on the departure date by default, not “whenever someone gets around to updating the system.”
  • The underlying voice model or embedding is deleted or deactivated — not merely left unused, which leaves it recoverable.
  • Any API keys or generation permissions tied to that person's clone are revoked the same day, on the same checklist as returning a laptop or badge.
  • For an ongoing use — a phone greeting that's still live — ask the departing person, while still employed and on good terms, whether they consent to its continued use after they leave. That conversation is far easier to have before a dispute than after one.

Watermarking helps, until it doesn't

Machine-readable marking of synthetic audio — the kind Article 50(2) of the EU AI Act will require from providers of generative AI from 2 August 2026, implemented in practice mostly through the C2PA / Content Credentials standard — is a genuinely useful idea with one hard limit: it only appears in audio actually created through a platform that adds it. A clone built with an open-source model, a self-hosted tool, or any platform that simply hasn't implemented the standard carries no watermark at all, and there is currently no reliable way to detect that absence after the fact from the audio itself.

That makes watermarking useful for a business that generates its own synthetic audio responsibly — it lets your own content prove it's yours — but it is not a fraud-detection tool you can rely on to unmask someone else's unmarked clone. The obligation is also extraterritorial in the same way the rest of the EU AI Act is: it reaches a Swiss business whose synthetic audio is used by, or offered to, people in the EU, regardless of where the business is registered. For the full regulatory picture beyond voice specifically, see our guide to AI transparency and disclosure.

What can go wrong even with a written policy

A policy on paper doesn't enforce itself. The realistic failure modes are quieter than a dramatic breach.

  • Consent scoped to “the phone greeting” quietly gets reused in a marketing video months later, without anyone asking again.
  • A departed contractor's access is forgotten because the IT offboarding checklist was written before voice cloning existed as a category to include on it.
  • Staff apply less scrutiny to a familiar-sounding voice on the phone than they would to a suspicious email, because “hearing is believing” is a hard instinct to override even when you know better.
  • A vendor's watermarking claim gets treated as a complete guarantee rather than the partial signal it actually is.

When cloning a voice is not the right answer at all

None of this is an argument that every business needs a policy binder before it can safely use an AI phone assistant. Most don't need voice cloning specifically, and it's worth saying so.

  • A business with occasional, low-stakes customer contact gains little from the governance overhead of consent, access review and offboarding that a cloned voice requires — a well-configured standard voice does the same job.
  • A business without anyone able to own the access review and offboarding work honestly shouldn't add a capability it can't administer safely, however appealing a signature voice sounds in a sales conversation.
  • An executive who is personally uneasy about being cloned should be allowed to say no without being talked out of it — that instinct is not paranoia, it may be the most accurate risk assessment in the room.

The decision, in one sentence

Before cloning anyone's voice on purpose, write down who consented, who is allowed to use it, and how it gets switched off — and before trusting a familiar voice on the phone, remember that the same few seconds of audio a business might reasonably want to use are exactly what a fraudster needs too, so verify anything involving money or access through a second channel regardless of how certain the voice sounds. The concrete next step isn't a bigger security budget. It's a one-page written answer to those four questions, agreed before any recording is made — the checklist above, filled in with real names.

Frequently asked questions

Is voice cloning illegal in Switzerland?

There is no single Swiss law that specifically bans or permits voice cloning. Using someone's voice without their consent is likely to engage personality-rights protections and the revised Data Protection Act (revDSG), which treats voice as personal data because it can identify a specific individual. Using a clone to commit fraud is criminal fraud regardless of the AI involved. This is general information, not legal advice — get advice specific to your situation before acting on it.

Does Weissmann's AI phone assistant clone voices by default?

No. The Starter and Premium phone-assistant packages use a standard synthetic voice, not a real person's. Voice cloning appears only as an Enterprise-tier option, offered on request rather than switched on automatically. Anyone considering it — with Weissmann or any other provider — should apply the consent, access-control and revocation questions in this article before agreeing to it.

Can watermarking prove a recording is fake?

It can only prove a recording is genuine, and only if it was made on a platform that added the watermark in the first place. It cannot prove the opposite — that the absence of a watermark means a recording is fake — because plenty of authentic audio, and plenty of unmarked cloned audio, carries no watermark at all. Treat it as a partial signal, not a verdict.

What is the single most useful defence against a fraud like the Arup case?

Never let a voice or video call alone authorise moving money, changing bank details or granting access. Confirm it through a second, independent channel — such as calling back a number you already had on file before the call, not a number given to you during it.

What should happen to a cloned voice when the person leaves the company?

Their consent should be treated as ending on their departure date unless they explicitly agree in writing, while still employed, to any continued use. The underlying voice model and any related access should be deactivated the same day, alongside the rest of the standard offboarding checklist.

Does the EU AI Act's Article 50 apply to a small Swiss business that never sells into the EU?

Only if its synthetic audio reaches people in the EU or is otherwise offered to them — the obligation follows where the output is used or offered, not where the business is registered. A purely domestic Swiss use case sits outside its direct reach, though the underlying practices remain good practice regardless.

Is a generic AI phone-assistant voice the same risk as a cloned one?

No. A generic synthetic voice doesn't represent an identifiable real person, so it doesn't raise the same consent, impersonation or personality-rights questions a clone of a specific individual's voice does. The fraud risk described in this article is specifically about voices built to replicate someone real.

Key terms in the glossary

← Back to overview

Practical AI for your business

From idea to implementation – we show you what is concretely possible in your case.

Request a demo
Call us Request a demo