Scollr summary
What this paper is about
This work converts EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose large language models (LLMs) to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data.
Full abstract
Read the full abstract
Electronic health records (EHRs) offer considerable potential for clinical prediction, but their complexity and heterogeneity challenge traditional machine learning. Domain-specific electronic health record foundation models trained on unlabeled EHR data have shown improved predictive accuracy and generalization. However, their development is constrained by limited data access and site-specific vocabularies. We convert EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose large language models (LLMs) to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data. LLM-based embeddings perform on par with a specialized EHR foundation model, CLMBR-T-Base, across 15 clinical tasks from the EHRSHOT benchmark. In an external validation using the UK Biobank, an LLM-based model shows statistically significant improvements for some tasks, which we attribute to higher vocabulary coverage and slightly better generalization. Overall, we reveal a trade-off between the computational efficiency of specialized EHR models and the portability and data independence of LLM-based embeddings.
Direct answer
What can I do from this paper page?
Use this page to scan "Large language models are powerful electronic health record encoders" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Machine Learning in Healthcare research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Hegselmann2026Large,
title = {Large language models are powerful electronic health record encoders},
author = {Stefan Hegselmann and Georg von Arnim and Tillmann Rheude and Noel Kronenberg and David Sontag and Gerhard Hindricks and Roland Eils and Benjamin Wild},
journal = {npj Digital Medicine},
year = {2026},
doi = {10.1038/s41746-026-02915-9},
url = {https://doi.org/10.1038/s41746-026-02915-9}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Machine Learning in Healthcare research papers?
Follow Machine Learning in Healthcare research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Large language models are powerful electronic health record encoders. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app