CLAP
This is not a single project, but a union of several endeavors in the area of AI and NLP in healthcare.
This research has been supported in part by a SUNY Small Team Multidisciplinary Award and a SUNY Seed Grant.
Ritwik Banerjee, Principal Investigator
Chenlu Wang, Doctoral Researcher → Research Scientist, Meta
Noushin Salek Faramarzi, Doctoral Researcher → NLP Researcher, Boeing AI
Akanksha Dara, M.S. Researcher → Software Engineer, Apple Inc.
Meet Patel, M.S. Researcher → Software Engineer, Yahoo
S. Harika Bandarupally, M.S. Researcher → Software Engineer, Meta
Collaborators
Yejin Choi, Professor, Computer Science, Stanford University
I. V. Ramakrishnan, Professor, Computer Science
Mark C. Henry, Professor and Chair Emeritus, Emergency Medicine
Matthew Perciavalle, Clinical Pharmacist, Emergency Medicine
Farrukh M. Koraishy, Clinical Professor, Nephrology
A problem in healthcare is that in spite of the recent focus on precision medicine, much of the relevant data is not patient-specific, and thus, corroborating relevant information and discarding the rest remains the manual endeavor of clinicians. In the first span of this research, from 2014–2015, we recognized this as a rather complex problem, with several aspects to it such as laboratory tests, prescription drugs, diet, etc. Thus, we developed AI-driven systems that can distill patient-specific information from large amounts of natural language data as well as structured databases. This has led to automatic recommendation of the most relevant laboratory tests for a patient, depending on the precise circumstances [Banerjee et al. 2014], and personalized identification of adverse drug reactions and attribution of a patient’s symptoms to their drug regimen [Banerjee et al. 2015]. In a parallel direction, we also developed NLP systems to automatically generate structured summaries of inpatient radiographic findings identified for clinical follow-up — a practically demanding application requiring comprehension of the highly specialized language of radiology reports [Whiteside et al. 2016].
By 2022, the landscape of natural language processing had transformed dramatically with the advent of transformer-based architectures. We pivoted to leverage these advances, investigating how modern deep learning models could better understand the semantic relationships within clinical documentation. Our work on semantic textual similarity in clinical notes demonstrated that transformer models pretrained on clinical text, when combined with medical ontologies like MeSH, could achieve strong performance without requiring extensive additional clinical training data [Salek Faramarzi et al. 2022]. This represented a shift from feature-based extraction toward contextualized understanding of medical language in a manner that combined deep learning with interpretable structures.
We then applied these transformer-based approaches to a more specific and practically challenging problem: extracting medication events from unstructured clinical narratives. Accurate medication history is fundamental to patient care, yet the complexity of how this information appears in clinical notes had made automated extraction difficult. By systematically evaluating various pretrained language models and incorporating domain-specific training with careful data augmentation, we achieved notable improvements in medication event detection [Salek Faramarzi et al. 2023]. This work underscored an important lesson: even the most sophisticated models require thoughtful preprocessing and domain adaptation to handle the nuances of medical documentation.
Recently, we extended our NLP methodology to an entirely different clinical domain: the analysis of kidney ultrasound reports for chronic kidney disease detection. Here, we returned to a fundamental question about model design — whether word-level or sentence-level approaches better capture diagnostic features. Interestingly, we found that the simpler word-level lexical models outperformed sentence-level analysis for identifying increased kidney echogenicity, a key marker of chronic kidney disease. When applied across over a thousand reports, this approach revealed that bilaterally increased echogenicity was the strongest predictor of chronic kidney disease, with nearly eight-fold increased odds. Through this work, we illustrated how NLP can transform descriptive imaging reports into structured, actionable risk assessments, potentially enabling earlier detection and intervention at scale across healthcare systems [Koraishy et al. 2024], [Wang et al. 2025].
Publications
Banerjee, R., Choi, Y., Piyush, G., Naik, A., and Ramakrishnan, I.V. 2014. Automated Suggestion of Tests for Identifying Likelihood of Adverse Drug Events. Proceedings of the 2014 IEEE International Conference on Healthcare Informatics, IEEE Computer Society, 170–175.
Banerjee, R., Ramakrishnan, I.V., Henry, M., and Perciavalle, M. 2015. Patient Centered Identification, Attribution, and Ranking of Adverse Drug Events. 2015 International Conference on Healthcare Informatics, IEEE Computer Society, 18–27.
Koraishy, F.M., Wang, C., Banerjee, R., et al. 2024. Use of Natural Language Processing and Deep Learning to Analyze Kidney Ultrasound Reports and Their Correlation with CKD Diagnosis. Journal of the American Society of Nephrology 35, 10S.
Salek Faramarzi, N., Dara, A., and Banerjee, R. 2022. Combining Attention-based Models with the MeSH Ontology for Semantic Textual Similarity in Clinical Notes. 2022 IEEE 10th International Conference on Healthcare Informatics (ICHI), IEEE Computer Society, 74–83.
Salek Faramarzi, N., Patel, M., Bandarupally, S.H., and Banerjee, R. 2023. Context-aware Medication Event Extraction from Unstructured Text. Proceedings of the 5th Clinical Natural Language Processing Workshop, Association for Computational Linguistics, 86–95.
Wang, C., Banerjee, R., Kuperstein, H., et al. 2025. Natural language processing for kidney ultrasound analysis: correlating imaging reports with chronic kidney disease diagnosis. Renal Failure 47, 1, 2539938.
Whiteside, I., Ramakrishnan, I.V., Banerjee, R., Balasubramanian, V., Kanaparthi, B.R., and Barish, M. 2016. Development and Evaluation of Natural Language Processing Software to Produce a Summary of Inpatient Radiographic Findings Identified for Follow-Up. Annual Meeting of the Radiological Society of North America.
This project page is hosted and maintained by the principal investigator, Dr. Ritwik Banerjee.