There's a version of healthcare AI that looks finished and isn't. The model works, the demo is impressive — and then someone asks where the patient data goes when it hits the language model, who can see it, and whether there's a signed agreement with every vendor in the path. If those answers weren't designed in from the start, the honest response is a partial rebuild.
HIPAA — the US law governing protected health information — is what makes AI HIPAA compliant, and it doesn't sit on top of a healthcare AI system as a compliance layer you add before launch. It shapes the architecture: where data lives, how it moves, who can touch it, what gets logged, and which vendors you're even allowed to use. Get those decisions right early and compliance is a manageable, ongoing discipline. Get them wrong and you're unwinding foundational choices under pressure.
This guide is engineering guidance, not legal advice. It walks through the technical backbone of a HIPAA-aware AI system — the encryption, audit logging, access control, and vendor-agreement decisions that a security auditor will actually examine — and it's honest about one thing that gets lost in marketing: "HIPAA-compliant" describes a whole program and posture, not a checkbox you tick on a single tool. Certification and attestation require an auditor, not a self-assessment. This sits directly on top of the same foundations as any AI agent handling sensitive data — HIPAA just adds legal teeth.
What Actually Counts as PHI
Before any architecture, you need to know what you're protecting. Protected health information is health information tied to an individual — and the identifier is broader than most teams assume. It's not only the diagnosis or the lab result. It's the eighteen categories of identifiers that, combined with health context, make data protected: names, dates more specific than a year, phone numbers, email addresses, medical record numbers, device identifiers, and more, down to full-face photographs and any other unique identifying number.
This matters enormously for AI, because the risk isn't only in the obvious fields. A free-text clinical note dictated into an AI medical scribe is dense with PHI — patient names, ages, addresses, family history — scattered through prose, not neatly boxed in database columns. An audit log that records "which patient record was accessed" is itself PHI. The retrieval context you feed a model in a medical LLM RAG system is PHI. Treating PHI as "the fields marked sensitive" while free text and metadata leak everywhere is one of the most common design failures in healthcare AI.
The practical consequence: your data-flow map has to trace PHI through every hop — ingestion, storage, the model call, the response, the logs, the analytics. Anywhere PHI travels, HIPAA's requirements travel with it.
The Third-Party LLM Problem: BAAs and Where Data Goes
Here is the challenge specific to AI, the one that doesn't exist in a traditional healthcare app. Most AI systems send data to a third-party model API — and if that data is PHI, you have a vendor in your compliance perimeter who now handles protected health information on your behalf.
Under HIPAA, any vendor that creates, receives, maintains, or transmits PHI for you is a business associate, and you need a signed Business Associate Agreement (BAA) with them. That includes your model provider. The critical, frequently-missed detail: standard and consumer tiers of most LLM APIs do not come with a BAA. The default terms often permit data to be retained or used to improve services, which is incompatible with handling PHI. You typically need a specific enterprise or healthcare arrangement — one that offers a BAA, excludes your data from training, and commits to defined retention and deletion.
So the architecture question becomes concrete: for each model call that touches PHI, is there a BAA covering that provider under the plan we're actually using? If the answer is no, you have three real options — move to a BAA-covered tier, de-identify the data before it leaves your environment, or run a model inside infrastructure you already control. This is exactly the third-party model risk that applies to any AI agent, but in healthcare the consequence of getting it wrong is a reportable breach rather than an awkward conversation.
De-identification and Data Minimisation Before the Model
The cleanest way to reduce risk is to reduce what leaves your controlled environment in the first place. Two techniques do most of the work.
Data minimisation means each model call carries only the PHI the task genuinely requires. If a summarisation prompt needs the clinical narrative but not the patient's full demographic block, don't send the demographic block. This is the least-privilege principle applied to data flow rather than accounts, and it shrinks the blast radius of any single failure.
De-identification goes further: stripping or replacing identifiers so the data reaching the model is no longer PHI. HIPAA recognises two routes — Safe Harbor, which removes the eighteen identifier categories, and Expert Determination, where a qualified statistician certifies that re-identification risk is very small. In practice, many AI systems tokenise or redact identifiers before the model call and re-attach them afterward inside the trusted boundary, so the model works on de-identified text and the identified record never leaves your environment.
Two honest caveats. First, de-identification is genuinely hard on free text — a name buried in a dictated note or a rare condition that is itself identifying can slip through automated redaction, so this needs validation, not blind trust. Second, if de-identification is imperfect or reversible in your design, treat the data as PHI and keep the BAA in place. De-identification is a risk-reduction technique, not a licence to be careless.
Encryption, Access Control and Audit Logging
Three controls form the backbone an auditor expects to see, and each has a specific AI wrinkle.
Encryption at rest and in transit. PHI should be encrypted in transit (TLS on every hop, including the call to the model API) and at rest (encrypted databases, object storage, vector stores, and backups). The AI-specific catch is the vector store: embeddings generated from clinical text are derived from PHI and often reconstruct it, so the vector database holding your RAG index needs the same encryption and access posture as the primary record store. Teams routinely lock down the database and forget the vector index sitting beside it.
Least-privilege access control (IAM). Every human and every service account gets the minimum access required, no more. The AI-specific failure mode is the over-permissioned service account: an AI agent given broad database credentials "to speed up the build" that were never narrowed. If your scribe integration can read every patient's record when it only ever needs the record for the active encounter, that's a finding waiting to happen. Scope every query to the authenticated context, and give the AI's service account exactly the tables and fields it needs — the same excessive-permissions discipline that applies to any agent, here with regulatory weight.
Comprehensive audit logging. HIPAA expects you to know who accessed which PHI, when, and what they did. For an AI system that means logging not just human access but every model call: which record was in context, which user or process triggered it, what action followed. This is what lets you answer an auditor's questions and detect misuse — and it's also why the logs themselves contain PHI and must be protected like any other PHI store. Tamper-evident, access-controlled, retained per policy.
| HIPAA requirement | What it means for an AI system | How to implement |
|---|---|---|
| Encryption in transit & at rest | PHI protected on every hop, including the model call and the vector index | TLS everywhere; encrypted DB, object storage, backups, and vector store |
| Business Associate Agreement | A signed BAA with every vendor that touches PHI, including the LLM and cloud provider | Use BAA-covered enterprise/healthcare tiers; sign the cloud provider BAA |
| Minimum necessary / least privilege | AI service accounts and users see only the PHI the task needs | Scoped service accounts; per-context query scoping; no admin keys |
| Audit controls | Record who accessed which PHI and every model call, when | Immutable, access-controlled logs covering human and model access |
| Access authorisation | Only authorised identities reach PHI or trigger the AI | Role-based access, MFA, per-session isolation, deprovisioning |
| Integrity & transmission security | PHI isn't altered or exposed in transit to third-party APIs | Validated payloads, minimised prompts, redaction before external calls |
HIPAA-Eligible Cloud, and the Program Around It
Most healthcare AI runs on cloud infrastructure, and the major providers offer HIPAA-eligible services — a defined subset of their services that may be used with PHI, provided you sign the cloud provider's BAA and configure them correctly. Two conditions there are load-bearing. You have to actually sign the BAA, and you have to stay within the eligible-service list; standing up PHI on a service outside that list, or without the agreement, doesn't become compliant just because it's the same cloud.
And this is the point where the honest framing matters most. HIPAA compliance is not a single tool you buy or a badge the cloud hands you. It's a program: administrative, physical, and technical safeguards; risk assessments; workforce training; incident-response and breach-notification procedures; and periodic review. The encryption and logging above are the technical safeguards — a necessary part, not the whole. A tool that measures or enforces some controls is helpful, but measuring controls is not the same as being compliant. That's why certification and attestation — whether a HIPAA assessment or a HITRUST certification, which many healthcare buyers now expect — require an independent auditor. No engineering team should tell a client they are "HIPAA-certified" on the strength of their own architecture; that's a conversation for qualified counsel and a third-party assessor.
One related boundary worth flagging: if your AI does more than assist — if it diagnoses or drives clinical decisions autonomously — it may fall under the FDA's Software as a Medical Device (SaMD) framework, an entirely separate regulatory track from HIPAA. Keeping the AI in an assistive, human-reviewed role reduces that exposure and is better clinical practice regardless.
Benefits of HIPAA-Compliant AI Architecture
Designing for HIPAA from the first sprint feels like overhead. In practice it pays back in several concrete ways.
No foundational rebuild later
The most direct benefit is avoiding the retrofit. Data access scoping, the logging model, and which model API sees PHI are foundational choices. Getting them right early means the system you demo is the system you can deploy, rather than a prototype that has to be partly rebuilt once someone asks where patient data goes. That protects both the budget and the launch date, because the expensive surprises happen before code exists rather than after.
Shorter security reviews with healthcare buyers
Hospitals, health systems, and payers run detailed vendor security assessments before letting any system near PHI. A team that can produce a PHI data-flow map, name the BAA-covered tiers in use, and show scoped access and audit logs answers those questionnaires with evidence rather than promises. That shortens procurement cycles and signals maturity to buyers who have seen many vendors promise compliance without the architecture behind it.
A smaller blast radius when something fails
Minimisation, de-identification, least-privilege service accounts, and encryption all limit what any single failure can expose. If a credential leaks or a component is misconfigured, a well-designed system exposes far less PHI than one where every service can read everything. Comprehensive logging also means you can determine what was accessed, which matters for any breach assessment and notification decision.
Access to stronger models without crossing the line
Understanding the BAA landscape lets teams use capable third-party models on covered tiers, or route sensitive steps to self-hosted models, instead of avoiding AI altogether out of caution. De-identification round trips widen the options further. The result is a system that can use the best model for each task while keeping PHI within agreed boundaries.
Evidence ready for auditors
Audit logs, access policies, and documented data flows are exactly what an independent assessor asks for during a HIPAA assessment or HITRUST certification. Building them in as part of the system, rather than assembling them under deadline, makes attestation a review of existing controls instead of a scramble to create them.
HIPAA-Compliant AI Use Cases
These are the common healthcare AI patterns where HIPAA architecture decisions carry the most weight, and how the controls above apply to each.
AI medical scribes
Ambient scribes capture clinician-patient conversations and draft structured notes. The audio and transcript are dense with PHI in free text, so the speech-to-text and summarisation providers both need BAA coverage, recordings need defined retention, and the draft note should be reviewed by the clinician before it enters the record. Done well, clinicians spend less time on documentation without patient data leaking to uncovered services.
Retrieval over clinical documents
RAG systems that answer questions over clinical guidelines, patient records, or care plans store embeddings derived from PHI. The vector index needs the same encryption and access controls as the source records, and queries must be scoped so a user can only retrieve documents they're authorised to see. The outcome is a useful assistant that respects the same access boundaries as the underlying record system.
Patient intake and messaging assistants
Chat or voice assistants that collect symptoms, schedule appointments, or answer billing questions handle identifiers from the first message. Minimising what each model call receives, logging every interaction, and escalating clinical questions to staff keep these assistants helpful without turning them into an unmonitored PHI channel. Transcripts should follow the same retention and access rules as any other patient communication.
Remote patient monitoring
Continuous readings from home devices are PHI from the moment they're collected. BAAs must cover the device platforms and cloud services in the path, and alerts reviewed by clinicians should be logged like any other access. The architecture lets care teams act on trends while keeping the whole chain accounted for, from the device to the phone app to the servers where readings are analysed.
Analytics and model development on de-identified data
Teams building population analytics or training internal models often work on de-identified datasets. Validated de-identification, via Safe Harbor or Expert Determination, reduces obligations, but only if it genuinely holds up against free-text leakage and re-identification risk. Where it can't be validated, the safer course is to keep the dataset inside the PHI boundary with the same controls as production.
Questions to Ask Your Development Team
"For every model call that touches PHI, is there a signed BAA covering that provider under the plan we're on?" The answer should name the provider, the tier, and confirm data is excluded from training. "We use the standard API" is a red flag.
"Show me the PHI data-flow map." They should be able to trace PHI through ingestion, storage, the model call, the response, and the logs — including the vector store. If free text and metadata aren't accounted for, they haven't mapped it properly.
"What does the AI's service account have access to, and how is it scoped?" You want least privilege and per-context scoping, not a broad or admin-level credential.
"What's logged, and are the logs themselves protected as PHI?" Every model call and PHI access should be logged, immutably and access-controlled.
"Are we de-identifying before the model, and how is that validated?" If they claim de-identification, ask how they test it against free-text leakage.
"Who signs off on the compliance posture — is there an auditor or counsel involved?" The right answer acknowledges that the engineering is necessary but not sufficient, and that attestation needs an independent party.
Common HIPAA-Compliant AI Mistakes
Most HIPAA problems in AI systems come from a small set of recurring design decisions. Each is easy to avoid early and expensive to fix late.
Calling the standard API tier with PHI
The fastest way to build a prototype is to use the model provider's default API. Those tiers typically don't include a BAA and may allow retention or use of data. Prototypes built this way often carry real patient data before anyone checks the terms, and the habit persists into production. Confirming BAA coverage on the specific tier in use, before any PHI flows, avoids starting the project with a reportable problem.
Protecting the fields and missing the free text
Teams often secure the obviously sensitive columns, such as diagnoses and record numbers, while PHI flows freely through clinical notes, prompts, model responses, and analytics events. Free text carries names, ages, and family details scattered through prose. Treating the data-flow map as a list of database fields, rather than every hop the data takes, leaves these paths unprotected.
Forgetting the vector store and the logs
Embeddings derived from clinical text and audit logs that record which patient was accessed are both PHI. They frequently sit outside the controls applied to the main database: unencrypted indexes, logs shipped to third-party tools without a BAA, broad read access for debugging. Auditors look for exactly these gaps, because they're where careful teams most often slip.
Shipping over-permissioned service accounts
Broad credentials granted to an AI service "to speed up the build" tend to outlive the build. An agent that can read every patient's record when it only needs the active encounter turns any prompt injection or bug into a much larger exposure. Narrowing service accounts to exact tables, fields, and per-session context is tedious, and skipping it is one of the most common findings.
Treating de-identification or tooling as the finish line
Automated redaction misses identifiers in free text, and a tool that monitors controls doesn't make an organisation compliant. Teams that assume either one settles the question end up overstating their posture to customers. De-identification needs validation, and compliance needs the full program of safeguards, training, and independent attestation.
HIPAA-Compliant AI Best Practices
A well-built system has a recognisable shape. These practices describe it in terms an engineering team can act on.
- Map PHI end to end before writing code. Trace every hop, from ingestion through storage, model calls, responses, logs, analytics, and vector stores, with free text and metadata included. Encrypt PHI at rest and in transit throughout, and update the map whenever a new component or vendor joins the path.
- Put a BAA in place with every PHI-touching vendor. That includes the model provider and the cloud provider, on tiers that exclude your data from training and define retention. Keep a register naming each vendor, tier, and agreement, and check it before adding any new service.
- Send the model the minimum. Minimise each prompt to the PHI the task requires and, where feasible, de-identify before data leaves the trusted boundary. Validate redaction against realistic free text rather than trusting it blindly.
- Make access least-privilege and audited. Use scoped service accounts, per-session isolation, MFA for people, and immutable, access-controlled logs covering both human and model access to PHI. Review permissions regularly and remove what isn't used, including access granted temporarily for debugging or migrations.
- Keep the AI assistive. Keep a licensed human in the decision loop for anything clinical, which supports both patient safety and a clear regulatory position. Document which functions are assistive so that scope creep toward autonomous clinical decisions gets noticed and reviewed.
- Treat compliance as a program. The architecture is one part; risk assessment, workforce training, incident response, and independent attestation are the rest, and everyone involved should know the difference.
- Bring counsel and assessors in early. Involve qualified legal counsel and, where buyers expect it, an independent assessor during design rather than at launch, so their findings shape the system instead of forcing changes to it.
What It Costs and How Long It Takes
Building HIPAA architecture in from the start adds real but bounded effort — the encryption posture, the scoped IAM, the audit logging, the de-identification layer, and the BAA legwork are weeks of work layered alongside core development, not a separate project. The honest caveat is the one from the pillar: healthcare AI has a higher floor than general software, and this is a large part of why. That floor is also the barrier that keeps the work defensible.
The expensive path is the other one. Retrofitting HIPAA architecture onto a working prototype — rescoping data access, rebuilding the logging model, re-routing PHI away from a non-BAA API — typically costs several times what designing it in would have, because these are foundational choices rather than surface features. And beyond the build, certification or attestation (a HIPAA assessment or HITRUST) is its own engagement with an independent auditor, on its own timeline. The same compliance-automation discipline that underpins SOC 2 and ISO programs applies here — HIPAA simply raises the stakes and narrows the acceptable answers.
Related guides
- What healthcare AI development actually involves
- FHIR and EHR integration, explained
- Medical LLMs, RAG and human-in-the-loop
- AI medical scribe: speech to SOAP notes to EHR
- AI agent security: risks and mitigations
- AI compliance automation: SOC 2 and ISO
- Our healthcare AI development services
We Build Healthcare AI That's HIPAA-Aware From the First Decision
We treat HIPAA architecture as part of scoping, not a review bolted on at the end. That means mapping PHI before we write code, choosing BAA-covered model and cloud tiers, scoping access to least privilege, logging every PHI touch, and de-identifying data before it leaves the trusted boundary where we can. We're also honest about the boundary of our role: we build the technical safeguards, and we'll tell you plainly where you need qualified counsel and an independent auditor to certify the program around them.
If you're building a healthcare AI system and want to get the compliance foundations right before they're expensive to change, that's exactly the conversation worth having early.
This architecture discipline runs through everything in our healthcare AI development practice, not just the systems that touch PHI directly.
Talk to us about your platform — no commitment, just a conversation.
Frequently Asked Questions
Does sending patient data to an LLM like GPT or Claude break HIPAA?
Not automatically, but it requires care. Sending PHI to a third-party model API makes that provider a business associate, so you need a signed Business Associate Agreement with them — and the standard or consumer API tiers of most providers do not include one. You typically need an enterprise or healthcare arrangement that offers a BAA, excludes your data from training, and defines retention. The alternatives are to de-identify the data before it leaves your environment, or to run a model inside infrastructure you already control. What breaks HIPAA is sending PHI to a provider with no BAA in place — the design, not the technology, is the deciding factor.
Can a single tool or platform make my healthcare AI "HIPAA-compliant"?
No, and it's worth being precise about this. "HIPAA-compliant" describes a whole program and posture — administrative, physical, and technical safeguards, risk assessments, training, and incident response — not a property of one tool. A vendor can offer HIPAA-eligible services or sign a BAA, which are necessary building blocks, but assembling them correctly and running the surrounding program is your responsibility. A tool that measures or enforces certain controls helps, but measuring controls is not the same as being compliant. Certification or attestation requires an independent auditor, not a self-declaration.
What exactly counts as PHI in an AI system?
Protected health information is health information tied to an individual through any of eighteen identifier categories — names, dates, contact details, medical record numbers, device identifiers, and more. In an AI system, PHI is broader than the obvious database fields: free-text clinical notes, the retrieval context fed to a model, embeddings in a vector store, and audit logs recording which record was accessed all contain or derive from PHI. The common failure is protecting the marked-sensitive fields while free text and metadata leak elsewhere. Your data-flow map has to trace PHI through every hop, including the model call and the logs.
Do we need a BAA with our cloud provider as well as the model provider?
Yes, if PHI touches either. The major cloud providers offer HIPAA-eligible services and will sign a BAA, but two conditions are load-bearing: you have to actually sign it, and you have to keep PHI on the eligible-service list and configure those services correctly. The same logic applies to the model provider and any other vendor in the PHI path. A practical rule: any vendor that creates, receives, maintains, or transmits PHI on your behalf is a business associate and needs a BAA.
Is de-identifying data before the model enough to avoid HIPAA obligations?
It can be, but only if the de-identification is genuine and validated. HIPAA recognises Safe Harbor (removing the eighteen identifier categories) and Expert Determination (a statistician certifying very low re-identification risk). The difficulty is that de-identification is hard on free text — a name buried in a dictated note, or a rare condition that is itself identifying, can slip past automated redaction. If your de-identification is imperfect or reversible, treat the data as PHI and keep the BAA in place. De-identification is a strong risk-reduction technique, not a blanket exemption you can assume without testing.
How does HIPAA relate to SOC 2, HITRUST, and FDA requirements?
They're different frameworks that often coexist. SOC 2 and ISO 27001 are general security-and-controls attestations — a strong foundation that HIPAA sits on top of, covered in our compliance automation guide. HITRUST is a certifiable framework popular in US healthcare that many buyers now expect, and it incorporates HIPAA requirements. The FDA's Software as a Medical Device (SaMD) framework is separate again — it applies if your AI diagnoses or drives clinical decisions autonomously rather than assisting a human. HIPAA governs how you handle PHI; the others govern security posture and, in the FDA's case, clinical function. Keeping the AI assistive and human-reviewed reduces FDA exposure, while HIPAA and a certification like HITRUST address the data and security side.
How much does it cost to make an AI system HIPAA-ready?
There's no fixed price, because the effort depends on how much PHI the system touches, how many vendors sit in the data path, and whether de-identification is part of the design. Designed in from the start, encryption, scoped access, audit logging, and BAA work usually add weeks alongside core development rather than a separate project. Retrofitting the same controls onto a finished prototype typically costs far more. Independent assessment or HITRUST certification is a separate engagement with its own budget.
Where should a team start when building AI that handles PHI?
Start with a data-flow map before writing code. List every place PHI enters, is stored, is sent to a model, appears in responses, and lands in logs or vector stores. Then confirm which vendors in that path will sign a BAA on the tier you plan to use, decide where de-identification or minimisation can reduce exposure, and involve qualified counsel early on the overall compliance program rather than at launch.
Conclusion
The core problem with healthcare AI isn't the model. It's that patient data now flows to places traditional healthcare apps never sent it: third-party model APIs, vector stores built from clinical notes, and logs that record what the AI saw. Each of those hops carries HIPAA obligations, and most of them are decided by early architecture choices.
The key moves are consistent. Map PHI end to end, including free text and metadata. Make sure every vendor in the path, model and cloud provider alike, is covered by a BAA on the tier you actually use. Minimise or de-identify data before it leaves your boundary, and validate that redaction rather than trusting it. Encrypt everywhere, scope service accounts tightly, and log every model call as carefully as any human access.
The caveats matter just as much. Technical safeguards are only one part of a compliance program that also covers risk assessment, training, and incident response, and attestation requires an independent auditor rather than an engineering team's self-assessment. AI that diagnoses or acts autonomously may also bring FDA requirements into play.
Before your next sprint, sketch the PHI data-flow map for your system and mark every vendor without a confirmed BAA. If you want engineering help building those safeguards in from the start, talk to our healthcare AI development team.
