A hospital IT director picturing a ransomware attack usually imagines a locked screen and a ransom note. That image misses what actually happens first. Encryption is still part of the attack, but it has slipped into a supporting role behind what criminals are really after, the data itself. The sensitive information has typically already left the building before a single file locks.

Data exfiltration is the unauthorized movement of data out of an organization's systems, and it has become the central objective of most ransomware operations rather than one possible outcome among several. The Episource breach of early 2025 shows why the distinction matters. Attackers reached the medical coding vendor's cloud environment, removed files containing the protected health information of roughly 5.4 million individuals across multiple US provider clients, and the organization's backup architecture was irrelevant to the outcome because the data was gone before any encryption took place.

 

What data exfiltration actually means

Data exfiltration is the transfer of data out of an organization without authorization. It can be copied from a device, pulled from a server, or lifted straight out of a cloud environment. Sometimes an outside attacker is behind it. Sometimes it is an employee, careless or otherwise.

The distinction worth drawing is between a breach and exfiltration, because the two are often treated as the same thing when they are not. A breach is unauthorized access to a system, whereas exfiltration is the act of getting data out of it. Healthcare breach data from 2025 shows how common the second has become, with the HHS OCR breach portal documenting that hacking and IT incidents, the category that includes ransomware and exfiltration, accounted for 59% of all reported healthcare breaches in the first half of the year. For a healthcare organization assessing its HIPAA obligations after an incident, the difference between an attacker who accessed a system and an attacker who removed data from it shapes the entire response.

Read more: What is ransomware?

 

Why exfiltration became the whole point

For years, ransomware worked on a simple premise, attackers encrypted an organization's data and demanded payment for the key, and the defense against it was equally simple in principle, which was to keep good backups so systems could be restored without paying.

That defense worked well enough that attackers had to change their approach. If a healthcare organization could recover from backups and simply refuse to pay, encryption alone gave attackers nothing. So they added a second threat. Steal the data first. Encrypt second. Threaten to publish or sell what was taken if the ransom goes unpaid. This is double extortion, and it is now the standard model, not the exception. The Verizon 2025 Data Breach Investigations Report found that 44% of breaches now involve ransomware, with data theft before encryption more commonly becoming the norm across those incidents.

The change matters enormously for healthcare as an organization that has invested in solid backups can restore its systems after an encryption attack and continue operating, but those backups do nothing about the copy of patient data now sitting on an attacker's server. The PIH Health ransomware attack showed how this plays out over time, with attackers exfiltrating roughly two terabytes of data, including Social Security numbers, financial account information, and medical records, and the organization taking more than a year to determine what had been taken and begin notifying the nearly three million people affected. Under HIPAA, exfiltrated protected health information triggers breach notification obligations regardless of whether systems were restored, which means an organization can recover operationally and still face patient notifications, an HHS OCR investigation, and the reputational consequences that follow.

 

Where the data actually comes from

The assumption that hospitals are the primary target no longer holds, the American Hospital Association reported that over 80% of stolen protected health information records in recent years were taken not from hospitals but from third-party vendors, software services, business associates, and nonhospital providers and health plans.

The Episource and Change Healthcare breaches both fit this pattern, and both carried consequences that reached far beyond the company actually attacked. When a billing vendor or a coding service is breached, every healthcare provider that relied on that vendor inherits the notification obligations, the regulatory scrutiny, and the reputational damage, even though their own systems were never touched. For a covered entity, this means the exfiltration risk extends to every business associate that handles its patient data, and a business associate agreement does not eliminate that exposure so much as define who is responsible for it.

 

How exfiltration happens

The methods attackers use to move data out have been cataloged in detail. The MITRE ATT&CK framework, maintained as a public knowledge base of adversary tactics and techniques, documents the primary exfiltration methods under a single category, and a handful of them account for most of what healthcare organizations encounter.

The most common path is exfiltration over the same channel an attacker uses to control compromised systems, where stolen data flows out alongside the commands coming in. Closely related is exfiltration to cloud storage services, where attackers upload stolen files to legitimate cloud platforms that blend in with ordinary business traffic. A more evasive technique is exfiltration over DNS, where data is broken into pieces and embedded inside DNS queries, the routine lookups that translate domain names into network addresses, which allows the data to slip past tools that are not watching that particular channel. Physical media such as USB drives remains a method as well, particularly where an insider is involved. The entry point that precedes all of this is usually mundane, with phishing and compromised credentials providing the initial access that everything else depends on.

 

Why is detection so difficult

The reason exfiltration is hard to stop is that attackers have learned to make it look like ordinary activity. Data leaving for a cloud storage service looks like an employee using a cloud storage service, and data tunneled through DNS queries looks like a computer doing what computers constantly do. A 2025 academic analysis published on arXiv examined the data exfiltration phase of double extortion attacks and found that it often originates from high-privilege accounts and backup servers, precisely the systems that generate large volumes of legitimate traffic and rarely trigger alarms when they send data outward.

Attackers routinely finish exfiltrating everything they want well before encryption or any ransom demand even enters the picture. Most organizations only realize they have been attacked once that theft is already complete. The Episource intrusion illustrated this clearly, with attackers maintaining access to the cloud environment for roughly ten days and removing millions of records during that window, well before the organization understood what was happening. Recovery efforts then focus on restoring systems while the actual loss, the stolen data, has already happened and cannot be undone.

 

What actually reduces the risk

Because exfiltration depends on an attacker first getting inside, the most effective interventions sit at the entry point rather than at the exit. The email that delivers initial access through a phishing link or credential harvest is the most controllable moment in the entire sequence, and it is where a healthcare organization has the clearest opportunity to break the chain before any data is at risk.

Paubox's 2025 Healthcare Email Security Report puts the employee reporting rate at just 5% of known phishing attacks in healthcare, which means the messages that most often begin an exfiltration chain go unreported in the overwhelming majority of cases. Pre-delivery filtering closes that gap by removing phishing attempts before clinical or administrative staff ever encounter them, and Paubox Inbound Email Security uses AI to analyze sender behavior, message intent, and contextual signals across every inbound message, catching the phishing attempts that signature-based filtering misses before they reach an inbox. On the outbound side, Paubox DLP provides a second layer directed at the exfiltration itself, monitoring outgoing email for protected health information and alerting administrators or blocking a message when PHI is being sent outside the organization, whether the cause is a careless employee or an attacker using a compromised account.

Learn more: Paubox Inbound Email Security

 

In the news

The Change Healthcare attack of February 2024 remains the clearest illustration of why exfiltration has become the dominant concern in healthcare cybersecurity. Attackers entered through a remote access portal that had no multi-factor authentication, moved through the network, and exfiltrated protected health information before deploying ransomware. A cross-sectional study published in JAMA Network Open documented the breach as affecting 100 million individuals, with $2.4 billion in response costs and disruption reaching care delivery nationwide. The same study, which analyzed every ransomware attack on HIPAA-covered entities from 2010 to 2024, found that hacking and IT incidents accounted for 88% of all records exposed across that period, a figure that captures how completely data theft has come to define the healthcare breach picture. The case also demonstrated that paying does not resolve the exfiltration problem, since data tied to the breach surfaced even after a substantial ransom was reportedly paid, a reminder that once data has left an organization, no payment reliably brings it back.

 

FAQs

What is the difference between a data breach and data exfiltration?

A breach is unauthorized access to a system, while exfiltration is the actual removal of data from it. Every exfiltration begins with a breach, but an attacker can access a system without successfully taking data out of it. For HIPAA purposes, the distinction affects how an organization assesses and reports an incident.

 

Why doesn't having backups protect against modern ransomware?

Backups let an organization restore encrypted systems without paying for a decryption key, but they do nothing about data that attackers stole before encrypting anything. Double extortion exploits exactly this gap, since the threat to publish or sell exfiltrated patient data remains even after systems are fully recovered from backup.

 

How does data exfiltration trigger HIPAA obligations?

When protected health information is exfiltrated, it is presumed to be a reportable breach unless the organization can demonstrate a low probability that the data was compromised. That triggers patient notification within 60 days, reporting to HHS OCR, and media notification for larger breaches, regardless of whether the organization restored its systems.

 

Why are healthcare vendors such a common source of exfiltrated data?

Vendors like billing companies, coding services, and clearinghouses hold concentrated volumes of patient data from many providers at once, which makes a single breach extraordinarily productive for an attacker. More than 80% of stolen healthcare records in recent years came from third parties rather than hospitals, and every provider that relied on a breached vendor inherits the notification and regulatory consequences.

 

What is the most effective way for a healthcare organization to reduce exfiltration risk?

Stopping attackers at the entry point offers the most advantage, since exfiltration cannot happen without initial access. Pre-delivery email filtering that removes phishing before staff encounters it, combined with data loss prevention that monitors outbound email for protected health information, addresses both the way attackers get in and the way data leaves.