Erik Kangas: Publications & Writing
The Future of Protected Health Information in the Age of AI
Table of Contents
HIPAA became law in 1996, three years before Google Search existed and eleven years before the iPhone. The people who wrote it could not have anticipated wearable heart-rate monitors, at-home genetic testing kits, or a large language model that can infer a medical condition from a photograph it was never shown for that purpose. The law’s core definition of protected health information (PHI), an individual’s health condition, care, or payment for care, tied to an identifier, has barely changed since. The technology surrounding that definition has changed enormously, and it keeps finding ways to generate health information the definition never anticipated.
This creates two separate problems for a security leader, and it’s worth naming them separately because they call for different responses. The first is that health-adjacent data is accumulating outside HIPAA’s reach faster than regulators can extend that reach to cover it. The second is that artificial intelligence is making it easier to convert ordinary, unregulated data, such as a photo, a voice recording, or a purchase history, into something that functions as health information even though nobody classified it that way when it was collected. Neither problem waits for the other to be solved first.
How HIPAA Has Evolved While Technology Moved Faster #
HIPAA’s scope has broadened several times since 1996, and each expansion followed rather than anticipated a change already underway. The original law applied only to covered entities: healthcare providers, health plans, and healthcare clearinghouses. The Privacy Rule and Security Rule, finalized in the early 2000s, defined what counts as PHI and how it must be protected. The Breach Notification Rule, added under the HITECH Act in 2009, required organizations to disclose when PHI had been exposed.
The most consequential expansion came in 2013, when the HITECH Omnibus Rule extended direct liability to business associates: the vendors, billing companies, and marketing firms that process PHI on a covered entity’s behalf without providing care themselves. That single change turned HIPAA compliance from something a hospital’s own staff managed into something an entire vendor ecosystem has to manage together. Information-blocking rules requiring providers to share electronic health records with patients on request have applied since April 2021; civil penalties for violating them are set to begin September 1, 2023.
Set against that timeline, the pace of technological change looks almost unfair to the regulators trying to keep up. Google Search launched in 1998, two years after HIPAA. Facebook in 2004. The iPhone in 2007, the same year 23andMe began selling consumer genetic tests. Consumer-facing AI models capable of generating fluent text arrived in late 2022. Each of these technologies created new ways to collect, infer, or transmit health-relevant information, largely outside the boundary HIPAA had already drawn.
Technology Has Already Expanded What Counts as PHI #
HIPAA’s list of individual identifiers is long: name, address, phone number, Social Security number, medical record number, IP address, photographic image, biometric data, and a catch-all for “any other characteristic that could uniquely identify the individual.” Combined with any of three categories of health information (a person’s condition, their care, or payment for that care), any one of these identifiers is enough to make data PHI.
What has changed is not this definition but the volume and variety of data now capable of satisfying it. Health systems have moved close to universal adoption of electronic health records. Telehealth and video visits, accelerated by the pandemic, produce their own records. Wearable devices and internet-connected medical equipment generate a continuous stream of biometric data outside a clinical setting. None of this was hypothetical when HIPAA was written; it simply didn’t exist yet.
The same technology has also made personalized care more valuable, which is part of what’s driving organizations toward wider data use rather than away from it. A study spanning health systems in Massachusetts, the Netherlands, Norway, and the UK found that continuous IT improvements to track outcome data across the full care cycle, paired with a value-based culture among providers, are central to implementing value-based care (Mjåset et al., NEJM Catalyst, 2020). A separate piece in Medical Economics made a related point about the fee-for-service model it’s meant to replace: without accurate, real-time patient data, providers can’t be held accountable for the outcomes and personalized care that value-based arrangements are supposed to reward (Kaushal, Medical Economics, 2022). Behavioral scientist Amy Bucher, writing in the Journal of Patient Experience, argued that digitally enabled personalization can substitute in part for the individualized attention healthcare staffing shortages make it harder to provide at scale (Bucher, Journal of Patient Experience, 2023).
None of that is an argument against using PHI. It’s an argument for using it deliberately, because the amount of it in circulation, and the number of systems touching it, keeps growing.
Why “Anonymized” Health Data Usually Isn’t #
The deeper problem is that a great deal of data now functions as identifiable health information even when nobody labeled it that way, because re-identification from supposedly anonymized records has turned out to be far easier than most organizations assume.
In 2015, researchers analyzing three months of anonymized credit card transactions from 1.1 million people, spread across roughly 10,000 stores, found that knowing just four spatiotemporal points, meaning where and roughly when someone made a purchase, was enough to re-identify 90 percent of the people in the dataset as unique individuals (de Montjoye et al., Science, 2015). Latanya Sweeney, working independently, matched health records in a supposedly de-identified Washington State hospital discharge dataset to news stories about hospitalizations and re-identified dozens of patients by name (Sweeney, 2013). A 2019 study in Nature Communications, using a generative model to estimate re-identification risk, concluded that 99.98 percent of Americans could be correctly re-identified in almost any dataset using just 15 demographic attributes (Rocher, Hendrickx & de Montjoye, Nature Communications, 2019).
None of that research involved AI directly; it’s a decade-old demonstration that “anonymized” and “de-identified” describe a much weaker guarantee than most people assume. What AI adds is scale and reach. Techniques that once required a research team and a specific dataset to combine can now be approximated with off-the-shelf models applied broadly, and AI systems have also gotten good at inferring health status from data that was never health data to begin with. Deep learning models can classify skin lesions from photographs at a level competitive with dermatologists (Esteva et al., Nature, 2017). More recent work has used voice recordings to detect Alzheimer’s disease (Brain Sciences, 2023) and speech and writing patterns to flag early cognitive decline (UT Southwestern Medical Center, 2023). A photo posted to social media, a voicemail, or a public speaking clip was never collected as health information and was never subject to HIPAA. Run through the right model, it can produce health information anyway.
California’s privacy law has already caught up with this reality in one respect: its definition of biometric information explicitly includes “sleep, health, or exercise data that contain identifying information” (Cal. Civ. Code § 1798.140). Almost any sufficiently detailed behavioral signal, on that definition, can double as a fingerprint.
The Regulatory Boundary Is Moving Too #
HIPAA already treats some categories of health information more strictly than others — psychotherapy notes and substance use disorder records both carry disclosure restrictions well beyond the general rule. That differential treatment is likely to keep expanding rather than staying fixed.
In April 2023, HHS’s Office for Civil Rights proposed a narrower but still significant change: a rule that would bar covered entities from disclosing PHI related to lawful reproductive health care for use in a civil, criminal, or administrative investigation against a patient, a provider, or anyone who helped them obtain that care (HHS OCR, Federal Register, 2023). The proposal followed the Supreme Court’s 2022 Dobbs decision and concerns that out-of-state investigators could subpoena reproductive health records from providers in states where the underlying care remained legal.
Congress has also had several bills in front of it that would extend privacy obligations well past HIPAA’s traditional boundary of covered entities and business associates, none of which has been enacted as of this writing. The Health Data Use and Privacy Commission Act would have created a federal commission to study and recommend health-data privacy standards spanning consumer apps and wearables (117th Congress, 2022). The My Body, My Data Act would specifically restrict retention and disclosure of reproductive health information (118th Congress, 2023). The Data Care Act, repeatedly reintroduced by Senator Brian Schatz, would impose data-security and use-limitation duties on a broad range of online services regardless of whether health data is involved (Data Care Act, 2023). The American Data Privacy and Protection Act, which cleared committee with bipartisan support before expiring at the end of the 117th Congress, would have created the closest thing to a comprehensive federal consumer privacy law the United States has yet considered (American Data Privacy and Protection Act, 2022). Separately, the White House’s Blueprint for an AI Bill of Rights, non-binding guidance from the Office of Science and Technology Policy rather than legislation, calls for protections against algorithmic discrimination and for meaningful human alternatives to automated decisions, principles that would bear directly on AI systems trained on or inferring health data (OSTP, 2022).
None of these has passed. Collectively, they describe the direction regulators and legislators are leaning, which is toward treating a wider range of data as sensitive and toward holding a wider range of organizations accountable for protecting it. HITECH’s 2013 extension of liability to business associates is the precedent: an expansion that seemed aggressive when proposed and unremarkable within a few years of taking effect.
Business Needs vs. Security Risk #
None of this changes the fact that PHI is genuinely useful, and forbidding its use is not a realistic answer for a security leader to give the rest of the organization. Clinical staff need it to educate patients and improve outcomes. Administrative teams need it to reduce no-shows and streamline reimbursement. Marketing needs it to personalize outreach, and accounting needs it to resolve claims. McKinsey’s 2021 personalization research found that 71 percent of consumers expect a personalized experience, and that companies that excel at personalization generate about 40 percent more revenue from those activities than average performers (McKinsey, 2021). A survey of healthcare consumers commissioned by Redpoint Global found a gap between that expectation and reality: while the large majority said relevant communication mattered to them, only about half described themselves as very satisfied with the relevance of what they actually received (Redpoint Global/Dynata, 2021).
The risk side of that ledger is just as real. A 2022 survey of IT professionals conducted by WatchGuard with Gartner Peer Insights found that nearly half of healthcare organizations, about 45 percent, had experienced a data breach in the preceding two years (WatchGuard, 2022). At that rate, a breach is closer to a when than an if, which changes what “security” means in practice. The job isn’t preventing every possible incident; it’s reducing the odds, limiting the blast radius when one happens anyway, and being able to demonstrate that reasonable safeguards were in place. The 2021 HITECH Safe Harbor amendment gives that last point real weight: HHS must now consider whether an organization had recognized security practices, drawn from frameworks like NIST, in place for the preceding 12 months when deciding how severely to penalize a breach or other HIPAA violation (Pub. L. 116-321, 2021). Adopting a cybersecurity framework isn’t just good practice; it’s now a documented mitigating factor if something goes wrong.
Two specific gaps deserve attention because organizations tend to assume they’re solved when they aren’t. The first is patient portals. Restricting all communication to a portal looks like a clean way to keep PHI off insecure channels, but data from the National Cancer Institute’s Health Information National Trends Survey found that only about 40 percent of American adults accessed their patient portal even once in the past 12 months (HINTS Brief 52, National Cancer Institute, 2023). PHI still moves through email, text, and phone calls whether or not an organization wants it to. The second is medical device security. Under Section 3305 of the Consolidated Appropriations Act, 2023, device manufacturers must now submit a plan to the FDA for addressing post-market cybersecurity vulnerabilities and a software bill of materials for the device’s components (FDA, “Cybersecurity,” Consolidated Appropriations Act 2023 §3305); the requirement took effect March 29, 2023, and the FDA said it would begin refusing to accept premarket submissions that omit this information starting October 1, 2023. Devices that predate that requirement, many of them already deployed in clinical settings, don’t benefit from it.
A third gap is less about devices than about code: HHS OCR issued a bulletin in December 2022 stating that tracking technologies like Google Analytics and the Meta Pixel, when placed on pages that handle PHI, generally can’t be used without a signed business associate agreement — and neither Google nor Meta signs one (HHS OCR, 2022). Organizations that added an analytics tag to a patient portal years ago, for reasons having nothing to do with HIPAA, may be out of compliance without realizing it.
Steps to Futureproof Your Compliance Posture #
The organizations best positioned for whatever comes next in this area are the ones treating expansion as the default assumption rather than waiting for OCR to formalize it. A few habits follow from that:
- Treat all data as potential PHI until it’s specifically ruled out, rather than relying on staff to classify it correctly case by case.
- Track technology sprawl deliberately — new tools, plugins, and analytics tags accumulate faster than most security reviews catch them.
- Monitor vendors and business associates on an ongoing basis, not just at contract signing, since HITECH already holds them directly liable.
- Move away from mutual consent as a long-term strategy for handling insecure channels. HIPAA does allow patients to accept the risk of insecure communication if properly warned and documented, but HHS guidance is clear that secure alternatives should be used whenever they’re not meaningfully more expensive or difficult. Encrypting everything by default removes the need to manage that consent process case by case, and it also removes the requirement to limit what’s disclosed over an insecure channel — which, in practice, expands rather than restricts what an organization can safely personalize.
- Adopt a recognized cybersecurity framework and keep it current. It reduces the odds of a breach, and it’s now a documented factor in how HHS assesses penalties when one happens anyway.
Takeaways #
The scope of what counts as identifiable health information is expanding, largely independent of whether HIPAA’s own text keeps pace. Regulators are extending its reach in specific areas, and AI is making it easier to derive health information from data nobody ever intended to be health information. Highly personalized, PHI-driven engagement is a genuine business requirement. The organizations that treat it as valuable and sensitive at the same time, rather than picking one framing and ignoring the other, are the ones that will be ready for whichever direction the law moves next.
References #
- de Montjoye, Y.-A., Radaelli, L., Singh, V. K., & Pentland, A. “Unique in the Shopping Mall: On the Reidentifiability of Credit Card Metadata.” Science, January 30, 2015. science.org/doi/10.1126/science.1256297
- Sweeney, L. “Matching Known Patients to Health Records in Washington State Data.” Data Privacy Lab, Harvard University, 2013. papers.ssrn.com
- Rocher, L., Hendrickx, J. M., & de Montjoye, Y.-A. “Estimating the Success of Re-identifications in Incomplete Datasets Using Generative Models.” Nature Communications 10, 2019. nature.com/articles/s41467-019-10933-3
- California Civil Code § 1798.140 (California Consumer Privacy Act, definition of “biometric information”). codes.findlaw.com
- Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. “Dermatologist-Level Classification of Skin Cancer with Deep Neural Networks.” Nature 542, 2017. doi.org/10.1038/nature21056
- “Artificial Intelligence-Enabled End-To-End Detection and Assessment of Alzheimer’s Disease Using Voice.” Brain Sciences, January 2023. pmc.ncbi.nlm.nih.gov
- UT Southwestern Medical Center. “Speech patterns may hold clues to early Alzheimer’s-related cognitive decline.” April 2023. utsouthwestern.edu
- U.S. Department of Health and Human Services, Office for Civil Rights. “HIPAA Privacy Rule to Support Reproductive Health Care Privacy” (proposed rule). Federal Register, 88 FR 23506, April 17, 2023. federalregister.gov
- Health Data Use and Privacy Commission Act, S. 3620, 117th Congress (2022). congress.gov
- My Body, My Data Act, H.R. 3420, 118th Congress (2023). govtrack.us
- Data Care Act of 2023, S. 744, 118th Congress, reintroduced by Sen. Brian Schatz, March 10, 2023. congress.gov
- American Data Privacy and Protection Act, H.R. 8152, 117th Congress (2022). congress.gov
- White House Office of Science and Technology Policy. “Blueprint for an AI Bill of Rights.” October 2022. whitehouse.gov
- McKinsey & Company. “The Next in Personalization 2021 Report.” 2021. mckinsey.com
- Redpoint Global (Dynata survey). “Healthcare Survey Shows a Need for Consistent, Relevant Engagement.” 2021. redpointglobal.com
- WatchGuard Technologies. “Nearly 50% of Healthcare Organizations Suffer Data Breaches” (Gartner Peer Insights survey summary). 2022. watchguard.com
- Public Law 116-321 (H.R. 7898), amending the HITECH Act to establish a “recognized security practices” safe harbor. Signed January 5, 2021. govinfo.gov
- National Cancer Institute. “Disparities in Patient Portal Communication, Access, and Use.” HINTS Brief 52, June 30, 2023. hints.cancer.gov
- Consolidated Appropriations Act, 2023, Section 3305 (“Ensuring Cybersecurity of Medical Devices,” codified as FD&C Act § 524B), effective March 29, 2023. See U.S. Food and Drug Administration, “Cybersecurity” (Digital Health Center of Excellence). fda.gov
- U.S. Department of Health and Human Services, Office for Civil Rights. “Use of Online Tracking Technologies by HIPAA Covered Entities and Business Associates.” Bulletin, December 1, 2022. hhs.gov
- Mjåset, C., Ikram, U., Nagra, N. S., & Feeley, T. W. “Value-Based Health Care in Four Different Health Care Systems.” NEJM Catalyst, November 10, 2020. catalyst.nejm.org
- Kaushal, A. “How Patient Personalization Can Aid the Transition to Value-Based Care Models.” Medical Economics, June 20, 2022. medicaleconomics.com
- Bucher, A. “The Patient Experience of the Future is Personalized: Using Technology to Scale an N of 1 Approach.” Journal of Patient Experience, April 2023. pmc.ncbi.nlm.nih.gov