Researchers identified 49% of users' real friendships from posting behavior alone, mapping the kind of relationships attackers impersonate in healthcare email attacks.

 

What happened

An attacker studying public review activity could correctly identify 49% of a user's social relationships while wrongly labeling only 10% of unconnected pairs, according to research from the McCombs School of Business at the University of Texas at Austin, published in Information Systems Research. Accepting a higher false-alarm rate of 20% pushed correct identification up to 63%. The team examined 4,299 Yelp reviewers in Louisiana and Pennsylvania during 2020, chosen because that platform publishes both review text and friend lists, which let the researchers check inferred connections against real ones. Most review platforms do not publish friend lists at all, which is what makes the finding matter, as the university explained. Knowing which colleagues an employee trusts is what makes an impersonated email convincing, and executive and vendor impersonation sits behind nearly every email-related healthcare breach recorded in 2025, according to Paubox's report on the top three healthcare email attacks.

 

Going deeper

Review length turned out to be the strongest signal as friends' posting habits move together, with one person writing longer reviews after a connection does, or writing a short one that complements a friend's longer post, none of which is visible to the person doing it. Knowing which relationships exist tells an attacker which sender name to use, which is the structure of spear phishing, and in a hospital it means an email appearing to come from a colleague in billing, a supervising physician, or a known supplier contact. Using the Pennsylvania data, the researchers calculated returns rising from 109% at 500 attempts to 1,098% at 10,000, because the cost of targeting and sending falls per message while the value of each successful impersonation does not.

 

What was said

"Social influence is a good thing. The problem is that network data could be leaked," said Yan Leng, assistant professor of information, risk, and operations management at McCombs, in the university's account of the research. On what platforms owe their users, Leng was direct, "The platforms need to protect not only what users disclose, but also what others can infer." She worked with co-authors at four other universities on the study.

 

In the know

Regulations including the European Union's General Data Protection Regulation govern what users disclose, while an attacker deriving connections from behavioral data has obtained something the user never published. Leng's proposed remedy is adding noise, meaning small deliberate random changes to data before release, such as varying review text length without altering meaning. In simulation, that approach cut the economic returns available to an attacker and in some cases turned them negative. Each platform would set its own privacy budget, balancing the analytical value it needs from the data against the social privacy risk it accepts. The researchers note that e-commerce marketplaces and media-sharing platforms carry the same exposure as review sites.

 

The big picture

Business email compromise produced $3.046 billion in reported losses during 2025, second only to investment fraud across every category tracked by the FBI Internet Crime Complaint Center. Healthcare staff use the same consumer platforms as everyone else, and a hospital's informal structure of who works alongside whom can be assembled from activity nobody considers sensitive. Organizations cannot control what employees post on review sites, which puts the practical response on verification habits rather than on restricting personal accounts. Staff should confirm any request involving credentials, payments, or access through a channel other than the one carrying the request, regardless of how well the sender appears to know their working relationships.

 

FAQs

Does deleting old reviews remove the exposure?

Partly, though platforms and third parties frequently retain copies, and archived versions circulate independently. The inference depends on patterns across a body of activity rather than any single post, so removing a few reviews does less than reducing ongoing activity.

 

What is a privacy budget in this context?

A limit an organization sets on how much information about individuals its released data may reveal, balanced against how useful the data remains for analysis. Adding more noise increases privacy and reduces analytical precision, and the budget defines where that trade sits.

 

Why do privacy regulations miss this kind of risk?

Most frameworks govern collection, disclosure, and consent for information a person provides. A relationship deduced from posting patterns was never provided by anyone, which places it outside rules written around disclosure.

 

Should healthcare organizations restrict what staff post online?

Restrictions on personal accounts are difficult to enforce and tend to breed workarounds. The more effective approach treats trusted-sender impersonation as an assumed capability and builds verification into the workflows an attacker would target, particularly credential resets, payment changes, and remote access requests.