U.S. Compliance: Deidentifying PHI to Pass §164.514 Audits
U.S. Compliance: Deidentifying PHI to Pass §164.514 Audits

Under HIPAA, PHI can be de-identified in only two authorized ways: Safe Harbor or Expert Determination, as set out in 45 CFR §164.514. Safe Harbor requires removing 18 specific identifiers, while Expert Determination relies on a qualified expert who applies statistical methods and formally attests that re-identification risk is very small. Properly de-identified data falls outside the Privacy Rule, but it still carries residual re-identification risk that compliance teams must manage, not ignore.
TL;DR:
- Removing all 18 specific identifiers is essential for Safe Harbor compliance, but unstructured text like clinical notes still requires careful manual review.
- Expert Determination involves a qualified analyst applying statistical methods and documenting the risk assessment to ensure the chance of re-identification is negligible.
- Proper de-identification workflows include inventorying data sources, choosing the appropriate method, secure key management, and thorough manual testing of sample records.
- Regular re-assessment using simulated linkage attacks and continuous risk testing are critical because external data availability evolves over time.
- Maintaining detailed documentation, including method rationale, assumptions, and version control, is vital for audit readiness, especially when sharing de-identified data for research or AI training.
Table of Contents
- The HIPAA de-identification standard under §164.514
- Safe Harbor method: the 18 identifiers, ZIP and date rules
- Expert Determination: statistical rigor and required documentation
- Preparing data and an operational workflow for de-identification
- Assessing re-identification risk: testing methods that matter
- Governance and administrative controls that satisfy auditors
- Special cases: clinical notes, images, and genomic data
- Common pitfalls and a short compliance checklist
- How de-identification fits into AI and data programs
- Arosplatforms services for de-identification-ready AI pipelines
- Sources
- FAQ
The HIPAA de-identification standard under §164.514
The Privacy Rule’s de-identification standard rests on a simple test: information is no longer PHI once there is no reasonable basis to believe it can identify an individual. 45 CFR §164.514 lays out two implementation paths that satisfy this test, Expert Determination under §164.514(b)(1) and Safe Harbor under §164.514(b)(2), and covered entities may choose either one depending on their data and intended use.
“No reasonable basis to believe” is a risk threshold, not a guarantee of zero risk. Regulators accept that some small, quantifiable chance of re-identification remains after either method is properly applied. The companion concept, “actual knowledge,” matters just as much operationally: if a covered entity knows that the remaining information could identify a specific person, even after applying Safe Harbor’s 18-identifier removal, the data does not qualify as de-identified. HHS guidance is explicit that this knowledge standard can override a mechanically completed checklist.
The legal payoff for getting this right is significant. Once data meets either standard, it is no longer subject to the Privacy Rule’s use and disclosure restrictions, which is what makes de-identification attractive for research, analytics, and AI model development. That exemption is also why regulators and auditors scrutinize the process closely: an organization that mislabels identifiable data as de-identified has effectively made an unauthorized disclosure of PHI. Getting the method, documentation, and testing right up front avoids that exposure later.
Safe Harbor method: the 18 identifiers, ZIP and date rules
Safe Harbor is the more mechanical of the two paths, but “mechanical” does not mean “simple.” The method requires removing all 18 categories of identifiers listed in 45 CFR §164.514(b)(2), and doing so wherever they appear, not just in the fields where a team expects to find them.
- Names of the patient or relatives, employers, or household members
- Geographic subdivisions smaller than a state, including street address, city, county, and precinct
- All elements of dates (except year) directly related to an individual, including birth date, admission date, discharge date, and date of death
- Telephone and fax numbers
- Email addresses
- Social Security numbers
- Medical record numbers
- Health plan beneficiary numbers
- Account numbers
- Certificate or license numbers
- Vehicle identifiers and serial numbers, including license plate numbers
- Device identifiers and serial numbers
- Web URLs
- IP addresses
- Biometric identifiers, including fingerprints and voiceprints
- Full-face photographs and comparable images
- Any other unique identifying number, characteristic, or code
Two identifiers carry built-in exceptions worth memorizing.
The hardest part of Safe Harbor in practice is not the structured database fields, it is free text. Clinical notes, discharge summaries, and radiology reports routinely contain names, dates, and device serial numbers embedded in prose, and a field-level scrub misses them entirely. HHS guidance confirms that Safe Harbor applies to identifiers “wherever found,” which means unstructured text needs its own redaction pass, typically a combination of automated pattern matching and manual review.
Safe Harbor also permits assigning a re-identification code to de-identified records under §164.514©, as long as the code is not itself derived from PHI and the covered entity keeps the key secured separately. That provision is what allows longitudinal research on de-identified data without compromising the standard, provided key management stays tight.
Expert Determination: statistical rigor and required documentation
Expert Determination suits datasets where Safe Harbor’s blanket removals would destroy too much analytic value, such as detailed geographic data for public health research or precise dates for outcomes studies. Under §164.514(b)(1), a person with appropriate knowledge of statistical and scientific methods for rendering information not individually identifiable applies those methods, documents the analysis, and formally determines that the risk of re-identification is very small.
HIPAA does not define a fixed credential list for who qualifies as an expert. That absence is intentional but also risky for organizations that treat it casually. In practice, most covered entities look for a demonstrable background in biostatistics, epidemiology, or data privacy engineering, plus a track record of prior determinations that would hold up under scrutiny. NCVHS committee recommendations have called for clearer minimum competency guidance precisely because the current standard leaves qualification judgment to the covered entity.
The statistical toolkit an expert draws on typically includes k-anonymity style analysis, which groups records so that no individual is distinguishable from a set number of others sharing the same quasi-identifiers, alongside broader risk modeling that estimates the probability of successful linkage against outside data sources. NIST IR 8053 notes that k-anonymity metrics offer intuitive guarantees for structured quasi-identifiers but can be brittle against creative attacks, while differential privacy provides a mathematical bound on what a query can reveal, at some cost to data utility. Choosing between these approaches, or blending them, depends on how the data will actually be used downstream.
“Very small” risk is a judgment call the expert has to justify, not a number fixed in the regulation. That justification is exactly what auditors and internal compliance reviewers will ask to see later, so the expert’s work product should include the specific methods applied, the assumptions behind them, representative data samples used in testing, the resulting risk metrics, and a clear date and version stamp tying the determination to a specific dataset snapshot. Without that paper trail, a determination that was defensible at the time becomes nearly impossible to defend months later when the dataset, or the surrounding data landscape, has changed.
Preparing data and an operational workflow for de-identification
De-identification works best as a defined workflow, not a one-time scrub before a data release. Compliance and privacy teams that treat it as a repeatable process catch more risk and produce more consistent documentation.
- Inventory every data location. Map structured fields, free-text notes, embedded images, and any genomic or biometric data, and flag fields that could enable linkage across datasets.
- Match the method to the intended use. Choose Safe Harbor when the analytic need tolerates coarse geography and shifted dates, and choose Expert Determination when precise dates, detailed geography, or rare conditions make Safe Harbor’s utility loss too costly.
- Apply preprocessing techniques deliberately. Normalize formats, tokenize identifiers, shift dates consistently per patient, and pseudonymize values with the key stored separately from the working dataset.
- Stand up a Disclosure Review Board. Give a small group of authority to sign off on any release, run staging-environment QA, and perform manual spot checks before data leaves the organization.
- Version and log everything. Keep a record of which method, parameters, and reviewers applied to each dataset snapshot, so any later question about a specific release has a clear answer.
Pro Tip: Run a sample of 50 to 100 records through manual review even after automated de-identification passes, since free text and edge-case fields are where automated tools most often leave identifiers behind.
Secure key handling deserves particular attention in step three. A pseudonymization scheme is only as strong as the separation between the lookup key and the de-identified dataset, and a shared drive or an unencrypted spreadsheet undermines the entire exercise. Philter, the clinical-text de-identification pipeline used at UCSF, pairs date shifting with a securely isolated key store as a practical model for this kind of separation. Teams building patient data pipelines for analytics or AI training often find that data integration best practices developed for other purposes, like consistent field mapping and lineage tracking, translate directly into better de-identification inventories.
Assessing re-identification risk: testing methods that matter
De-identification is not a switch you flip once. NIST frames it as an ongoing risk-management exercise, because the outside data available for linkage attacks keeps growing, which means a dataset judged low-risk today can look different in two years.
Three statistical concepts come up repeatedly in risk assessment, and practitioners benefit from understanding the trade-offs rather than treating them as interchangeable. K-anonymity ensures each record is indistinguishable from a minimum number of others sharing the same quasi-identifiers, which is intuitive but vulnerable when an attacker has outside knowledge about the group. L-diversity extends that protection by requiring diversity in sensitive attribute values within each group, addressing cases where everyone in a k-anonymous group shares the same sensitive diagnosis. T-closeness goes further still, requiring that the sensitive attribute distribution within a group closely match the overall dataset distribution, which reduces the risk that group membership itself leaks information. Differential privacy takes a different approach entirely, adding calibrated noise to query results so that no single record’s presence or absence meaningfully changes the output, at some cost to the precision of any individual analysis.
Beyond the math, practical testing matters. Simulated linkage attacks, where a reviewer attempts to match de-identified records against public voter rolls, social media, or commercial data brokers, reveal weaknesses that statistical formulas alone can miss. Some organizations run internal “tiger team” exercises specifically to try to re-identify their own released data before anyone else does. NIST guidance recommends against relying solely on automated redaction tools for exactly this reason: a scanner can confirm that a field is empty, but only a targeted human or adversarial test can confirm that the remaining information, taken together, still resists identification.
Periodic re-assessment matters just as much as the initial test. A dataset released for one study, then repurposed for a different AI training project two years later, deserves a fresh risk look, not a carryover assumption that the original determination still holds. Combining automated metrics with human review and a documented governance process, rather than either one alone, is the pattern NIST and practitioner literature both point to as more reliable than either approach used in isolation.

Governance and administrative controls that satisfy auditors
The technical de-identification work only holds up if the paperwork behind it does too. Auditors and internal reviewers want to see a clear record of what was done, why, and by whom.
- Method documentation: which approach was used (Safe Harbor checklist or Expert Determination report), including the specific identifiers removed or the statistical methods applied.
- Assumptions and risk results: the reasoning behind key decisions, such as why a ZIP prefix was retained or what quasi-identifiers the expert modeled.
- Version and date stamps: tying every determination to the exact dataset snapshot it covers, since datasets change.
- Data Use Agreements: recipient obligations even for de-identified data, particularly restrictions on attempting re-identification or re-linking to other sources.
- Retention schedule: how long de-identification records, expert reports, and risk assessments are kept, and where.
Data Use Agreements deserve special mention because de-identified status does not eliminate all obligation to the data’s original subjects. Even when data legally falls outside the Privacy Rule, a well-drafted DUA typically prohibits the recipient from attempting re-identification, restricts further redistribution, and requires notification if a breach or re-identification attempt occurs. This is a contractual backstop layered on top of the technical and legal protections, and many covered entities require it as standard practice even when not strictly mandated.
Re-assessment triggers should be written down in advance rather than decided ad hoc: a new analytics use case, a merger that introduces new external datasets capable of linkage, or a multi-year gap since the last review are all reasonable prompts to redo the risk analysis. Organizations building governance programs around AI and analytics often find it useful to formalize these triggers alongside broader AI governance and compliance processes, since the same review cadence and documentation discipline apply to both.
Special cases: clinical notes, images, and genomic data
Structured database fields are the easy part. The harder categories are where most real-world re-identification risk concentrates.
Clinical text carries names, dates, employer references, and device identifiers woven into sentences, which is why redaction pipelines need both automated pattern matching and human review checkpoints. The Philter V1.0 pipeline built at UCSF combined programmatic redaction, date shifting, and third-party verification to produce certified de-identified clinical notes at scale, showing that a staged approach, automated first pass followed by manual spot checks, can meet HIPAA’s standard for high volumes of free text.

Medical images and photographs carry their own risk profile. Full-face photographs are explicitly listed among the 18 Safe Harbor identifiers, and even non-facial images can carry identifying metadata in the file itself, such as device serial numbers or embedded location tags. Removing that metadata and, where facial features appear, applying face-removal or blurring techniques are standard mitigations, though algorithmic redaction of images carries its own error rate and typically needs a manual verification step before release.
Genomic and rare-disease data pose the sharpest challenge of all, because a genetic sequence or a rare diagnosis can be effectively unique to one person or one family, making conventional de-identification techniques far less effective. For these data types, Expert Determination combined with controlled access, rather than public release, is generally the safer path: a limited data set shared under a strict Data Use Agreement with a named, vetted recipient carries far less exposure than an openly published file. The general rule holds across all three categories: the rarer or more distinctive the underlying data, the more the balance should tilt toward controlled access over broad de-identified release.
Common pitfalls and a short compliance checklist
Most de-identification failures trace back to a handful of recurring mistakes, not exotic edge cases.
- Overreliance on automated scanning tools without a manual review pass, which misses identifiers embedded in free text or unusual formats.
- Treating Safe Harbor as a field-by-field checklist rather than checking every location an identifier could appear, including notes and attachments.
- Misapplying the ZIP or age exceptions, such as retaining a low-population three-digit ZIP prefix or failing to bucket ages 90 and older.
- Weak or missing documentation, leaving no record of which method, assumptions, or risk results supported a given release.
- Ignoring linkage risk by evaluating a dataset in isolation instead of considering what outside data could be combined with it.
A short pre-release checklist catches most of these before they become a problem: confirm the method chosen and why, confirm every identifier location including free text and metadata, confirm ZIP and age exceptions were applied correctly, confirm documentation is complete and dated, and confirm a linkage risk review was performed, not just a field scan.
Pro Tip: Before any release, have someone outside the project team try to re-identify five random records using only public information. If they succeed even once, the dataset is not ready.
How de-identification fits into AI and data programs
Every AI project built on patient data runs into the same tension: the more thoroughly you de-identify a dataset, the less analytically useful it becomes, and the less you de-identify it, the more compliance risk you carry into the model. Compliance officers rarely get to make that call alone. It requires input from data science on what utility loss the model can tolerate, from legal on documentation standards, and from IT on where pseudonymization keys actually live.
The organizations that handle this well treat de-identification as a design decision made early in a project’s architecture, not a cleanup step applied right before a dataset ships. That means deciding up front whether a project needs Safe Harbor’s speed and simplicity or Expert Determination’s flexibility, and building the Disclosure Review Board and documentation habits into the project plan from day one rather than retrofitting them under deadline pressure. A readiness assessment or a governance sprint, run before the first dataset moves, tends to surface these decisions while they are still cheap to change.
— arosplatforms team
Arosplatforms services for de-identification-ready AI pipelines
Building a HIPAA-compliant de-identification workflow into an AI project is easier with a partner who has already mapped out the governance, testing, and pipeline architecture. Some AI consultancies design AI systems embedding de-identification, documentation, and disclosure review into the pipeline rather than adding them afterward.
The AI Readiness Assessment helps compliance and data teams decide where Safe Harbor or Expert Determination fits their specific analytics or AI training goals before any data moves. From there, AI Governance & Compliance engagements help formalize Disclosure Review Board processes, documentation standards, and re-assessment triggers, while Custom AI Development builds the actual pipeline, redaction steps, key management, and QA checkpoints included.
If your organization is preparing patient data for research, analytics, or an AI project and needs the governance and pipeline work done right the first time, start with a readiness assessment to see where your current process stands.
This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.
Sources
- Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule
- eCFR :: 45 CFR 164.514 – Other requirements relating to uses and disclosures of protected health information.
- De-Identification of Personal Information (NIST IR 8053)
- NIST guidance (related SP/IR publications) on de-identification and risk management
- Philter V1.0: scalable clinical-text de-identification (UCSF/PubMed Central)
FAQ
Can PHI records be de-identified?
Yes, PHI can be de-identified under HIPAA using one of two authorized methods, Safe Harbor or Expert Determination. Once properly de-identified, the information falls outside the Privacy Rule’s restrictions, though a small residual re-identification risk always remains.
Which are the 18 PHI identifiers?
The 18 Safe Harbor identifiers include names, geographic subdivisions smaller than a state, all dates except year, phone and fax numbers, email addresses, Social Security numbers, medical record and account numbers, device and vehicle identifiers, URLs, IP addresses, biometric data, full-face photographs, and any other unique identifying code. All 18 categories must be removed for Safe Harbor to apply.
What are the 5 basic rules of HIPAA?
HIPAA is generally organized around the Privacy Rule, the Security Rule, the Breach Notification Rule, the Enforcement Rule, and the Omnibus Rule, each governing a different aspect of protected health information handling. De-identification specifically falls under the Privacy Rule’s requirements in 45 CFR §164.514.
What are 5 examples of PHI?
Common examples of protected health information include a patient’s name linked to a diagnosis, a home address tied to a treatment record, a medical record number, a full-face photograph in a patient file, and a date of birth connected to lab results. Each of these falls within the 18 identifier categories that Safe Harbor requires removing before data can be considered de-identified.