Writing / Compliance

Does removing names de-identify patient data for AI?

October 6, 2026 / 7 min read

No. HIPAA recognizes two de-identification methods, and a name is one of eighteen identifiers. Here is what each method needs and where redaction falls short.

No. Taking the patient’s name out of a prompt does not make it de-identified under HIPAA. The Privacy Rule recognizes two methods, and a name is one of eighteen categories of identifier the safe harbor method requires you to remove, along with the condition that you have no actual knowledge the rest could still identify the person. The other method needs a documented statistical determination that the risk is very small. Text that meets neither is still protected health information, which means the rules on disclosure, on business associate agreements, and on breaches all still apply to it. Redaction software that works on a best-effort basis is a technique, not either method.

What does HIPAA mean by de-identified?

The standard is at 45 CFR 164.514(a) and (b). Health information is not individually identifiable if it “does not identify an individual and with respect to which there is no reasonable basis to believe that the information can be used to identify an individual.” A covered entity may conclude that only by one of two routes. The payoff for using either one is real: under 164.502(d)(2), the requirements of the subpart do not apply to information de-identified in accordance with 164.514. The cost of falling short is that they do.

What does the safe harbor method require?

Under 164.514(b)(2), the identifiers of the individual, and of their relatives, employers, or household members, must all be removed, and the covered entity must have no actual knowledge that what remains could be used, alone or with other information, to identify the person. The eighteen categories are:

  • names;
  • geographic subdivisions smaller than a state, including street address, city, county and zip code, with a narrow exception for the first three digits of a zip code in larger areas;
  • all elements of dates except year that relate directly to the person, including birth, admission, discharge and death dates, and all ages over 89;
  • telephone numbers, fax numbers, and email addresses;
  • Social Security numbers, medical record numbers, health plan beneficiary numbers, and account numbers;
  • certificate and license numbers;
  • vehicle identifiers and serial numbers, and device identifiers and serial numbers;
  • web addresses and IP addresses;
  • biometric identifiers, including finger and voice prints;
  • full face photographs and comparable images; and
  • any other unique identifying number, characteristic, or code.

Now read a prompt against that list. Consider an invented, illustrative line: a ninety-one-year-old admitted on March 3 after a fall at home, with a note about the daughter who lives nearby. The name is gone and the line still carries an age over 89, a date of admission, and a detail about a relative. Clinical prompts are mostly free text, and free text is where identifiers hide: in dates, in places, in the reference to a relative or an employer, and in the last catch-all category, any other unique characteristic. Removing a name is the easy category. The others take a review of the whole text.

What is expert determination?

The other route, 164.514(b)(1), requires a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods to apply them and determine that the risk is very small that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify the individual. The expert must also document the methods and results that justify the determination.

Note what that is: an analysis with a written record, done by a qualified person about a particular way of handling information and a particular recipient. It is not a setting on a tool. Whether any process for stripping prompts could be supported that way is a question for a qualified expert and your counsel, and this post makes no claim that one can.

Does redaction software count as de-identification?

Neither method mentions software. The safe harbor asks whether all the identifier categories were removed and whether you lack actual knowledge that the remainder identifies someone. That is a test of the result, applied to the text in front of you. A tool that catches most patterns most of the time can help you get there, and it cannot promise the result, because best-effort is what best-effort means. If one identifier in a prompt slips through, the prompt is not de-identified, and the analysis is the one in whether pasting patient data into an AI tool is a HIPAA breach.

There is a second catch, and it is easy to miss. Under 164.502(d)(1), a covered entity may use protected health information to create de-identified information, but it may disclose protected health information for that purpose only to a business associate. In our reading, sending identifiable text to an outside AI service so that the service can strip it is itself a disclosure to that service, and it needs the same agreement any other disclosure would. Where the agreement stands is covered in is ChatGPT HIPAA compliant.

What about replacing names with codes?

Section 164.514(c) allows a covered entity to assign a code so that de-identified information can be re-identified later, on two conditions: the code is not derived from or related to information about the individual and cannot be translated to identify them, and the covered entity does not use or disclose the code or the mechanism for re-identification for any other purpose. A random placeholder such as “Patient A” with a mapping table that stays inside the organization can fit that description. A placeholder built from a medical record number, or a hash of one, does not, and under 164.502(d)(2) disclosing a code designed to enable re-identification is itself a disclosure of protected health information.

What should a health system have in place?

  • A policy that says removing a name is not de-identification, and that a prompt containing anything on the eighteen-item list is treated as protected health information.
  • A named owner and a written method if the organization wants to use de-identified data with AI at all, decided with counsel, including whether it will rely on safe harbor or on a documented expert determination.
  • A rule that placeholders are random and that any mapping table stays inside the organization.
  • An approved-tools list, because whether a disclosure is permitted depends on the agreement behind the tool.
  • A record of AI use on managed devices, so that when someone asks whether a prompt was clean, there is something to look at besides memory.

Where Verillian fits

Verillian governs AI use on the devices you enroll. A checkpoint on each device sits between your people’s AI tools and agents and the AI providers it supports. For Claude and Claude Code traffic (the Anthropic API format), a tool call your policy bans is removed before your machine can run it; for the other supported providers, it screens and records the usage, and the Claude desktop app and Cursor are recorded only, with no redaction. Each record is signed on the device it came from and hash-chained to the one before it, so a change to its signed fields is detectable, and it stays on your own infrastructure. It cannot show that nothing was omitted. Redaction is best-effort, not a guarantee that every value is caught. The admin server runs where you choose: on-prem or in a private cloud you run. macOS is the supported install today; Windows has an interim scripted installer and Linux builds from source. For a health system, that means the checkpoint can flag, and once an administrator sets redaction on, replace, the identifiers your policy flags from the fixed set of detectors Verillian ships, such as a Social Security number or an email address, both of which are on the safe harbor list. On a fresh install it only flags and does not change the traffic, and a redacted Social Security number leaves the device as [US_SSN_REDACTED]. A fixed set of detectors is not a review of eighteen categories, and Verillian’s redaction is not a de-identification method under 164.514(b) and should not be described as one. What the record can add is evidence of which AI service was used, from which enrolled device, attributed to the user the checkpoint reports, and when. Verillian is built for HIPAA-regulated environments, and no HIPAA certification of a product exists.

Verillian sees only the devices it is installed on. A prompt typed on a personal phone or an unmanaged computer never reaches it.

For what a record has to contain to be worth keeping, see what an AI audit trail is. The related policy question is in what a clinician AI use policy has to cover, and the platform page covers how policy is applied.

Sources

  • 45 CFR 164.514, (a) and (b), de-identification standard and methods, and (c), re-identification codes.
  • 45 CFR 164.502, (d)(1) and (d)(2), uses and disclosures of de-identified protected health information.

All writing

See the record
for yourself

Thirty minutes with your security team. We show policy enforced at execution and the signed chain it produces.