WhatsApp has begun testing a new machine-learning security feature that can identify potentially fraudulent messages directly on a user’s phone, as the Meta-owned messaging service seeks to tackle increasingly sophisticated scams without weakening its end-to-end encryption protections.
The optional feature, calledScam Alert, is initially being released to a limited group of beta users and security researchers. Rather than sending conversations to Meta’s servers for inspection, it downloads a compact machine-learning model to the user’s device and analyses incoming messages locally.
When the model determines that a conversation from an unknown sender resembles a known scam, WhatsApp displays a private warning inside the chat. The recipient can then block the sender, report the account or continue the conversation.
The warning is not visible to the person who sent the message.
WhatsApp said no message content leaves the device during classification and that neither the content nor the detection result is automatically reported to WhatsApp, Meta or another organisation. According to the company’stechnical explanation of Scam Alert, information can be shared only when the user deliberately chooses to submit a report.
The distinction is significant because WhatsApp’s default end-to-end encryption prevents the company from reading the contents of ordinary private messages and calls. Introducing server-side scanning would risk creating new privacy concerns, particularly if the system could inspect conversations without the participants’ consent.
Scam Alert instead attempts to place the detection capability at the endpoint—the user’s phone—where a decrypted message is already available to the legitimate recipient.
How the new warning system works
Once Scam Alert is enabled, WhatsApp downloads the detection model and runs it locally against incoming messages from people who are not saved in the user’s contacts.
The model was trained using patterns found in scam conversations that users previously reported to WhatsApp. It uses probabilistic classification to examine linguistic signals and the structure of a developing conversation.
This means the system is not necessarily looking for one prohibited word, suspicious link or predefined phrase. It can consider how the conversation progresses and whether the language resembles commonly observed fraud techniques.
A scammer may, for example, begin with a deliberately vague greeting, claim to have contacted the wrong person, impersonate a relative, advertise an unusually lucrative job or present an investment opportunity. The criminal may then attempt to build trust, create urgency, move the victim to another platform or demand an advance payment.
Scam Alert is intended to recognise combinations of signals associated with those approaches. However, WhatsApp describes the system as probabilistic rather than definitive. A warning means that the conversation resembles patterns seen in known scams; it does not prove that the sender is a criminal.
If the user believes the warning is incorrect, the conversation can be marked as trusted. WhatsApp will remove the notice and prevent Scam Alert from flagging that particular chat again.
The user can also voluntarily send the five most recent incoming messages to WhatsApp when marking a conversation as trusted. That information may help the company study false positives and improve future versions of the model, but sharing it is not required.
Scam Alert can be disabled at any time.
WhatsApp builds safeguards around the machine-learning model
WhatsApp is introducing more than a locally installed classifier. It has also designed a verification system intended to make the model’s delivery and operation open to scrutiny.
The company said each model version, including experimental variants, will be registered on a public transparency ledger before deployment. This is designed to prevent WhatsApp or Meta from secretly delivering a special model to one specific person.
Model variants used during experiments are selected on the device using a locally generated random value. According to WhatsApp, the server supplies the available configurations but cannot decide which particular model a named user receives.
WhatsApp will also publish the model weights, allowing qualified independent researchers to inspect what the system is designed to detect and investigate whether it could be used for purposes beyond scam classification.
Each model’s cryptographic hash is entered into a signed manifest on the transparency ledger. Researchers can therefore compare a model obtained from a device with the public record and determine whether it is an authorised version.
Users will be able to inspect some of the model’s activity through WhatsApp’s information-request facility. Under Account, Request Info and Scam Alert Activity, transparency records can show which messages were analysed, the result of the analysis, whether a warning appeared and which model version performed the classification.
Those records remain particularly important because on-device artificial intelligence is otherwise largely invisible to the person carrying the phone. A system may technically operate locally while still leaving the user with little understanding of what it examined or why it acted. WhatsApp’s proposed logs are intended to make that activity more observable.
Anonymous analytics without collecting conversations
Although the message classifier operates entirely on the phone, WhatsApp still needs a way to measure whether it is producing useful results or overwhelming users with inaccurate warnings.
The company plans to collect two narrowly defined categories of statistics: approximate numbers of warnings generated and aggregated counts of the actions people take after receiving those warnings. Those actions can include trusting the sender or blocking and reporting the account.
WhatsApp says raw message content is not included in this telemetry. The device first reduces the relevant events to counts, which are transmitted at randomised times without device identifiers and with only coarse timing information.
The resulting metrics pass through a confidential federated analytics system built on confidential virtual machines. These are Trusted Execution Environments intended to isolate the processing from the surrounding infrastructure.
Before sending the metrics, the WhatsApp client checks an attestation proving which software is running inside the protected environment. It compares that result with a third-party record of approved software and encrypts the data between the phone and the protected system.
Individual submissions are then combined into larger groups. WhatsApp says it receives only results that exceed a minimum cohort size and have been altered using differential privacy, a mathematical technique that adds controlled statistical noise to make it difficult to infer anything about an individual participant.
The company would consequently see broad trends—such as whether a new model is generating far more warnings or whether users frequently reject its classifications—without receiving the underlying messages or a readable record tied to one device.
These protections will still require independent testing. An architecture can be privacy-preserving in design while implementation defects, insecure update processes or unexpected interactions with other application components introduce risks. WhatsApp said it has expanded its bug-bounty work around Scam Alert and provided researchers with early Android application builds so they can test whether messages remain on the device and whether reporting is genuinely user-controlled.
The company is publishing the technical design before a wider release partly to invite that scrutiny.
Why encrypted platforms face a difficult detection problem
Scam prevention creates an unusual challenge for encrypted messaging services. Centralised platforms can analyse content on their servers, but end-to-end encrypted systems are deliberately designed so that the service provider cannot access private conversations.
Encryption protects users from surveillance, account-provider compromise and unauthorised interception. The same technical boundary, however, limits the provider’s ability to inspect an ongoing conversation and determine whether one participant is manipulating the other.
On-device classification offers a possible compromise. The provider supplies a model, but the examination occurs on hardware controlled by the recipient. No central service needs to decrypt the conversation.
That approach does not eliminate every concern. A model could produce false positives, fail to recognise unfamiliar scams or perform differently across languages and cultural contexts. Criminals may deliberately alter their phrasing after learning how the classifier works. Publishing model information also provides transparency for defenders while potentially helping fraudsters test ways to evade it.
Scam Alert is therefore best understood as an additional warning mechanism, not an automated judgment about a sender’s intentions.
Its initial focus on non-contacts should reduce unnecessary inspection of established conversations and concentrate the system on the point where many fraud attempts begin. It may nevertheless miss attacks launched from compromised accounts belonging to real friends, relatives or colleagues.
Those compromises are especially dangerous because they inherit the credibility and conversation history of a known contact. A local model restricted to unknown senders would not necessarily intervene, meaning users must still independently verify unusual requests for money, passwords, payment codes or confidential information.
Industrial-scale scam operations drive the need for new controls
The rollout comes as online fraud has developed into a global, cross-platform industry. Large criminal organisations operate scam centres that can run multiple campaigns simultaneously, including fraudulent investment schemes, fake employment offers, romance scams, impersonation fraud and task-based payment schemes.
Many of these operations do not remain on one service. A target may initially receive a text message, encounter an advertisement or make contact through a dating platform. The conversation can then move to WhatsApp or another private messenger before the victim is directed to a cryptocurrency exchange or payment service.
This fragmentation gives each technology company only a partial view of the operation. A message that appears harmless in isolation may be one stage in a longer manipulation process.
In August 2025, WhatsApp said it haddisabled more than 6.8 million accounts associated with criminal scam centresduring the first half of that year. Meta said some of the accounts were identified and removed before the operators could fully deploy them.
The company described one disrupted campaign in which criminals used multiple services to move potential victims through a sequence of fake online tasks. Targets were shown fictional earnings and were eventually instructed to deposit cryptocurrency to continue participating.
Some scam compounds also rely on human trafficking. The United Nations Office on Drugs and Crime has documented how people recruited with deceptive job advertisements can be transported to guarded compounds and forced to commit online fraud. UNODC warned in its2025 assessment of cyberfraud in the Mekong regionthat organised crime groups were expanding their operations and diversifying beyond their traditional geographical bases.
In June 2026, Meta, Microsoft, Coinbase, Starlink and law-enforcement agencies announced a coordinated operation against Southeast Asian scam networks. The participants said the action led to the disruption of more than one million online assets, the arrest of 63 suspected participants and the freezing of more than $3 million in cryptocurrency. Meta alone reported disabling 1.4 million Facebook and Instagram accounts, pages and groups, while Microsoft suspended approximately 20,000 accounts. Details were published inMeta’s account of the joint operation.
Those figures illustrate why filtering individual messages is only one part of a wider response that also requires account enforcement, financial tracing, telecommunications controls and international policing.
Scam Alert joins a growing collection of WhatsApp protections
The beta expands a series of controls intended to intervene at different stages of a scam or account takeover.
In March 2026, WhatsApp introduced warnings for suspicious device-linking requests. Fraudsters sometimes persuade a target to disclose a linking code or scan a QR code, falsely claiming that it is required to vote in a competition, access a service or verify an account.
If successful, the process adds the attacker’s computer or phone as a linked WhatsApp device. The criminal may then gain access to the victim’s account and use the trusted identity to target contacts.
WhatsApp’s linking protection uses behavioural signals to identify potentially fraudulent requests and presents contextual information about where the request originated. Meta described the change as part of abroader anti-scam update across WhatsApp, Facebook and Messenger.
The company has also introduced safety summaries when an unknown person adds a user to an unfamiliar group. Notifications from the group remain muted until the recipient decides to stay, allowing the person to leave without first opening the conversation.
Warnings shown when users attempt to share their screens with unknown contacts address another common technique. Criminals may ask victims to share their display and then watch as bank details, one-time passcodes or other sensitive information appear.
For individuals exposed to more sophisticated threats, WhatsApp launched Strict Account Settings in January 2026. The lockdown-style option automatically applies restrictive protections, including blocking attachments and media from unknown senders and silencing calls from unfamiliar numbers. WhatsApp said the setting was intended particularly for people such as journalists and public figures who may face rare but advanced attacks. It can be enabled through Privacy and Advanced in the application’s settings, according toMeta’s announcement.
Scam Alert addresses a different layer of the problem. Instead of concentrating on account configuration, device linking or dangerous attachments, it attempts to identify the social-engineering conversation itself.
A potentially important safeguard for three billion users
The scale of any WhatsApp security change is considerable. The service hasmore than three billion users across over 180 countries, making it one of the world’s largest private communications platforms.
That reach also creates a substantial classification challenge. Scam language varies by country, language, age group and type of fraud. Legitimate job offers, discussions about money or unexpected messages from new acquaintances could resemble malicious approaches, particularly during the opening stages of a conversation.
WhatsApp has not announced when Scam Alert will become generally available, which mobile platforms will receive it first or whether the beta will cover every language supported by the application. The limited rollout is intended to establish whether the system can provide useful warnings without generating excessive false positives, consuming too much battery or creating unintended privacy risks.
Even if the model performs well, the company’s design leaves the final decision with the user. That is necessary because social engineering depends on context, and software cannot reliably establish a stranger’s intentions from a handful of messages.
A warning can interrupt the sense of urgency on which many fraud attempts depend, but it cannot replace verification. Users should contact a supposed friend, relative, employer, bank or government organisation through a separately obtained and trusted channel before transferring money or disclosing credentials. They should never share registration codes, two-step verification PINs or device-linking codes, and should be particularly cautious when an unexpected contact promises guaranteed returns or demands an advance payment.
The most consequential aspect of Scam Alert may therefore be its timing. By placing a warning inside the conversation before a user has fully engaged, WhatsApp is attempting to disrupt the psychological progression of the scam while preserving the encryption that protects legitimate private communication.
Whether that balance works in practice will depend on the accuracy of the model, the integrity of its update and analytics systems, and the results of independent security testing during the beta.