Novo Nordisk announced that it experienced a data breach on June 20 when unauthorized individuals accessed and copied information from internal networks, including clinical trial data. Protected health information (PHI) and demographic information was accessed, but the notice did clarify that direct identifiers such as name, email addresses, phone numbers, etc. were not expected to have been accessed.
Removing direct identifiers can be useful, but does not always mean that data is safe. Firstly, we have no way of knowing from Novo Nordisk’s notice if the dataset was in fact processed according to HIPAA’’s deidentification standard. Secondly, the compliance analysts will still need to ask why the data was collected, how it was de-identified, what other data was included in the set, and could the records have been used to attribute to a specific individual. As the experts at Paubox point out with respect to the Novo Nordisk breach, removing direct identifiers is just part of the solution.
A recent JAMIA article leaves no ambiguity around the legal reality, “Once data qualify as de-identified they are no longer covered by HIPAA and can be used for any purpose, without restriction.” Qualify is the operative term here. Just because a dataset has a label, came with vendor assurances, has a tokenization process, or does not include a name field does not make it HIPAA deidentified.
What qualifies as deidentified data under HIPAA?
The Safe Harbor and Expert Determination are two methods HIPAA prescribes for deidentification.
Safe Harbor requires that certain elements of information be removed from data about an individual and, as appropriate, the individual’s relatives, household members, and employers. The covered entity also must not know that the information can identify an individual by itself or in combination with other information. “The Safe Harbor method requires all 18 personal identifiers to be eliminated,” according to one AMIA Annual Symposium Proceedings Archive article.
Expert Determination allows for more flexibility if a qualified expert applies generally accepted statistical and scientific principles to determine that the risk is very small that the information can be used to identify an individual and documents the methods and results. However, this option may allow you to retain useful variables, but it requires documentation of a defensible analysis.
The Health Affairs Scholar reviewed guidance showing that Safe Harbor also requires a covered entity not to have actual knowledge that remaining information can be used to identify an individual. It could include knowledge that variables can identify someone with a rare diagnosis from a small population; knowledge of highly detailed timelines; free-text notes, images, and data elements; or knowledge that a dataset can be combined with information from outside sources after receipt.
When might the incident still involve a HIPAA violation?
HIPAA applies identifiable information that meets the classification of PHI. There are multiple reasons why the dataset you thought was deidentified was still PHI. These reasons include, but aren’t limited to, leftover identifiers missed, faulty expert determination, or the organization knew that the data left on the records could be used to identify someone.
HIPAA may apply if the attacker hasn’t exported everything they accessed. It’s common to find files extracted from the application still lingering in source records, backups, log files, technical support tickets, email attachments, or even stored in the cloud. PHI stored in those locations will need to be remediated so the incident response team should verify the scope of the attack across all forensic artifacts.
Limited datasets should have all direct identifiers removed but can contain dates, city, ZIP code, and other fields. A limited data set is a restricted version of PHI which excludes the 16 names and numbers. Full deidentification is different as exposure of a limited dataset may result in a HIPAA breach review.
Reidentification risk does not automatically prove a HIPAA breach
A previously cited paper notes, “De-identified information can be re-identified (rendered distinguishable) by using a code, algorithm, or pseudonym that is assigned to individual records.” Healthcare organizations must protect mapping systems and consider the recipient’s data.
Research discussing reidentification attacks must be viewed with caution. One systematic review showed that, of 14 successful attacks against deidentified information, only two used information deidentified under present standards. Further, while analyzing attacks using deidentified health information (only study example) deidentified under present standards, researchers saw two successful attacks out of 15,000. It offers a rate of just 0.013%. Authors point out that most successful attacks used information not properly deidentified. Widespread concerns of the certainty of successful attacks couldn’t be substantiated by the paper’s findings.
Just because someone can achieve successful reidentification doesn’t mean the initial disclosure was in violation of HIPAA. It could have occurred because later applications of the method revealed flaws. The rules may have changed after the data was disclosed. Maybe the recipient obtained more information than they were explicitly entitled to.
Deidentified, pseudonymized, limited, and encrypted are not interchangeable
Healthcare professionals sometimes think of these three terms as interchangeable options for one approved level of data protection when they’re not.
HIPAA defines deidentification as being either Safe Harbor or Expert Determination. Data can be pseudonymized or tokenized by substituting identifiable data with unique codes, but that doesn’t automatically make it HIPAA deidentified data. Limited data sets contain PHI where certain identifying information has been removed, but the information is still considered PHI and is typically required to be accompanied by a data use agreement when used.
Encrypted PHI is identifiable data that has gone through an algorithm that makes the data unreadable by an unintended individual. Encryption can alter whether notifications are necessary depending on the situation and technology, but the information does not stop being PHI.
An encrypted electronic medical record, a limited data set, and a properly deidentified set of data extracted from that record could all lead to different penalties under HIPAA if leaked by the same hacker. The information must first be identified by the covered entity.
Where secure communication fits
In one Paubox Report, 58% of 170 surveyed US healthcare IT decision makers reported suffering an email breach in the past two years. Of respondents who experienced a breach, 47% said their first course of action was to implement stronger encryption following the breach. While these statistics demonstrate a need for tighter controls around data in motion, they also shed light on the false belief that encryption deidentification.
HIPAA compliant email is one necessary tool for reducing risk when your teams need to send PHI, research extracts, or breach notifications. But it won’t grant Safe Harbor compliance, replace nuanced human analysis, protect files post-download, or manage an entire process. Organizations still need access controls, endpoint security, secure storage, vendor management, logging/monitoring, incident response, etc.
FAQs
Can HIPAA deidentified information be outside of the CCPA?
Information that has been deidentified in accordance with HIPAA’s standard deidentification process is generally outside the scope of the California Consumer Privacy Act.
Can the Federal Trade Commission (FTC) Health Breach Notification Rule apply outside of HIPAA?
Certain health apps, connected devices, personal health record vendors, and related businesses that fall outside HIPAA may still be subject to the FTC Health Breach Notification Rule.
Does 42 CFR Part 2 apply when substance use disorder data is deidentified?
Once information is properly deidentified such that it would not reveal that an individual is a patient of a federally assisted substance use disorder program, Part 2’s identifying-records provisions typically won't apply. Part 2 can still apply when substance use disorder data is being collected, transformed, stored, or disclosed.
