A Framework for Investigating Live Medical Data against Privacy Laws

This research has been supported in part by the Secure & Trustworthy Cyberspace (SaTC) program of the U.S. National Science Foundation (NSF)
CNS-2335686 (10.2023 – 09.2025)

Ritwik Banerjee, Principal Investigator
  Chenlu Wang, Doctoral Researcher  →  Research Scientist, Meta
  Weimin Lyu, Doctoral Researcher  →  Applied Scientist, Amazon

Indrakshi Ray, Co-Principal Investigator
  Ethan Myers, Doctoral Researcher
  Yunik Tamrakar, M.S. Researcher  →  Software Engineer, ABC Legal Services

Collaborators

Lorenzo De Carli, Associate Professor of Electrical and Software Engineering, University of Calgary
Shantanu Sharma, Assitant Professor of Computer Science, New Jersey Institute of Technology
Chaoyuan Zuo, Lecturer (tenure-track), School of Journalism & Communication, Nankai University


Digital health systems (mobile apps, clinical databases, health information services, etc.) operate at the intersection of two demands that NLP is uniquely positioned to address: regulatory compliance and information integrity. They must comply with privacy laws that vary by jurisdiction, sector, and enforcement regime; and they must accurately represent the scientific evidence they draw on. When either demand is unmet, patients are exposed to risks they did not consent to — either because an app misrepresents what it does with their data, or because the health information it surfaces misrepresents what the science actually says. This project develops the computational tools and frameworks to detect and measure both classes of failure.

Privacy Compliance in Health Applications

Mobile health apps collect some of the most sensitive personal data that exists — symptoms, medications, reproductive cycles, mental health indicators — and their compliance with the privacy laws governing that data is difficult to verify at scale. We develop NLP models that automatically infer what data a health app collects and processes from its textual description and permission declarations, and test whether that behavior is consistent with what the app actually does [Tamrakar et al. 2025]. This work achieves higher accuracy and lower annotation costs than prior approaches, and provides a practical tool for auditing health app ecosystems at the platform level.

Privacy Laws Across Jurisdictions

The compliance problem is jurisdictional, and not merely technical. A health app that is GDPR-compliant in Europe may violate DPDPA in India; an app compliant with both may still fall short of frameworks emerging in the Middle East. We study how privacy laws across multiple nations (spanning four continents) construct different, and sometimes incompatible, notions of consent, sensitive data, and legitimate interest, and what these differences mean for globally deployed health software and services [Sharma et al. 2026]. This work frames cross-jurisdictional compliance as a natural language inference problem at scale: the meanings of legal terms shift materially across legal corpora, and computational models must account for this to be useful in practice.

Health Information Integrity

A second thread of this research concerns the scientific accuracy of health information in digital media. Ensuring that health claims made in apps, news articles, and social media posts are traceable to legitimate scientific evidence requires both the ability to identify relevant biomedical expertise at scale and the ability to detect when a claim has been distorted or fabricated across language boundaries. We develop a large-scale PubMed-based retrieval framework for biomedical expert finding, enabling automated identification of subject-matter experts capable of evaluating specific health claims (missing reference). Complementing this, we develop cross-lingual models for medical misinformation detection through contrastive claim-evidence reasoning, extending the verification pipeline beyond English-language sources to multilingual health discourse [Zuo and Banerjee 2025]. Both systems contribute to the goal of making the accuracy of health information computationally auditable, not just manually reviewable.

Our novel and data-efficient training paradigm that may underpin several improvements to these models (viz., class distillation with Mahalanobis contrast) is described in detail as part our work on pragmatic language understanding [Wang et al. 2025].

Broader Impact

This project safeguards user privacy and security in the increasingly prevalent use of health apps, which handle sensitive personal data. By addressing regulatory compliance across jurisdictions, improving the interpretability of legal language in technical contexts, and building tools for health information verification, the research contributes both to the protection of individual users and to the broader challenge of making digital health systems accountable to the scientific and legal standards they operate under.


Publications

Sharma, S., Myers, E., De Carli, L., Banerjee, R., and Ray, I. 2026. Local Privacy Laws in a Globalized World. Proceedings of the Sixteenth ACM Conference on Data and Application Security and Privacy, Association for Computing Machinery, 193–204.      (PDF coming soon)

Wang, C., Lyu, W., and Banerjee, R. 2025. Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, 29428–29442.     

Zuo, C. and Banerjee, R. 2025. HDCR: Cross-lingual Medical Misinformation Detection through Contrastive Claim-Evidence Reasoning. 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), IEEE, 1–6.     

Zuo, C., Wang, C., and Banerjee, R. 2025. Large-Scale Biomedical Expert Finding for Health Claim Verification: A PubMed-based Retrieval Framework. 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), IEEE, 4578–4583.     

Tamrakar, Y., Myers, E., Banerjee, R., De Carli, L., and Ray, I. 2025. Harnessing Language Models to Analyze Android App Permission Fidelity. Proceedings of the 22nd Annual International Conference on Privacy, Security, and Trust, IEEE.     



This project page is hosted and maintained by the principal investigator, Dr. Ritwik Banerjee.