How the medical database developed at MIT evolved into a global standard for data sharing | MIT News

Before the development of scientific data storage and collaboration in the cloud, medical investigators seeking the success of clinical research had to overcome significant obstacles in the collaboration and collection of important clinical data.
Data was censored and difficult to disseminate, so those who wanted to do research had no choice but to collect it themselves. This not only made the research more expensive, but also made it challenging to compare findings across datasets.
In 1975, researchers studying arrhythmias at MIT and Boston’s Beth Israel Hospital saw another way: the team began collecting and digitizing electrocardiogram recordings with the goal of not only studying them, but also making them available to the wider research community.
The team built their own computers for the process, carefully duplicating the tapes one by one, and creating more than 100,000 annotations for the recordings. The process took years, but by the summer of 1980, the tapes were finally ready. The team originally thought their tool would reach fewer than a dozen academic and industrial groups. But the interest kept coming. During the next ten years, they continued to send about 100 copies.
The data eventually became the first database of the global platform PhysioNet – founded in 1999 at the Harvard-MIT program in Health Sciences and Technology – as a repository of clinical data for complex biological signals.
At the time, that kind of data sharing, which might seem automatic today, was a near-revolutionary idea. “The creation of PhysioNet was an amazing vision,” said Thomas Heldt, Richard J. Cohen (1976) Professor of Medicine and Biomedical Physics, associate director of MIT’s Institute for Medical Engineering and Science, and senior author of the latest paper. Natural Health to assess the impact of the field.
Eventually, those magnetic tapes were mailed to burned CD-ROMs, which then evolved into managed FTP servers on the newly created Internet. Today, as PhysioNet looks back on more than 25 years of operation, the platform hosts hundreds of databases, and has become one of the leading biomedical and clinical data repositories in existence. Last year, more than 15,000 scientific publications were indexed in PhysioNet, and users from more than 180 countries registered on the platform. It is widely used by researchers, manufacturers, and clinical decision makers.
“The impact of the research is really significant,” said Heldt, who is also a professor in MIT’s Department of Electrical Engineering and Computer Science and a principal investigator at the Research Laboratory of Electronics, “and it’s really humbling.”
“It’s really good to see that an idea like this has been proven right and allows so many people.”
Setting the standard
Around 2009, a PhD student named Tom Pollard was conducting research on critically ill patients at one of London’s leading hospital systems. Although the hospital produced large volumes of important clinical data, the infrastructure and processes needed to prepare and support their use for extensive research were still in progress.
“Hospital data is collected primarily to support immediate patient care, with little attention paid to how it can be processed and used for research,” said Pollard, now a research scientist at MIT’s Laboratory for Computational Physiology (LCP), technical director of PhysioNet, and lead author on the study. Natural Health paper.
The problem was not just privacy. Hospital information systems are primarily designed to support patient care and management, not research. Data were fragmented across systems and rarely organized with future reuse in mind, making it difficult and expensive to transform them into relevant research resources.
But Pollard needed details to complete his draft. After searching the Internet, he finally found the Medical Information Mart for Intensive Care (MIMIC), a database of de-identified electronic health records managed by PhysioNet. Seeing its potential, his clinic manager, Kevin Fong, arranged a visit to Boston. Soon after, Fong and Pollard were sitting across the table from Roger Mark, discussing how their teams could work together.
Academic advocacy has long favored publication and specialized analysis over the abstract work involved in preparing data for use by others. That tension still exists today. PhysioNet’s founders embraced a different model, believing that sharing research resources could accelerate discovery and ultimately improve human health, he says. MIIC became the focus of Pollard’s writing, and after completing his PhD, he came to MIT to help build the next generation of databases.
In the years since PhysioNet was founded, the value of sharing research data has gained much wider recognition. The late Roger Mark, MIT distinguished professor of health sciences and technology emeritus and co-founder of PhysioNet, described its mission as building an “accessible international community around data” to “positively impact global health.”
Earlier this year, Mark and the late George Moody, founder of PhysioNet, jointly received the prestigious IEEE Biomedical Engineering Award for their contributions to PhysioNet and biomedical signal processing. IEEE pointed to their “leadership in ECG signal processing and global dissemination of selected biomedical and clinical information, thereby accelerating biomedical research worldwide.”
The platform’s source code, like most of its data, is public. According to the The environment excerpt: “As the field evolved, the PhysioNet community expanded significantly beyond its origins in signal processing and cardiovascular health to include clinical data, critical care and machine learning in healthcare.” People have used that to build their own PhysioNet-esque infrastructure, Heldt said. Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet’s legacy.
Although there are now many services out there that handle the same electronic health data, according to Google DeepMind researcher Vivek Natarajan, both PhysioNet and MIMIC “set the standard,” he says, “and it’s still the standard right now.”
That standard, according to those using the platform, changed the way research was done. Access to data should not be the determining factor in what ideas are possible, according to Ziad Obermeyer, an associate professor at the University of California at Berkeley’s School of Public Health and the College of Computer, Data Science, and Society.
“PhysioNet changed the way I think about the bottleneck in research. Usually it’s not ideas or talent. It’s a conflict. If access to data is slow, expensive, and difficult, the ideas that die first are the ones with the greatest risk, things that probably won’t work, but can change if they do. That’s exactly the wrong model if you want real progress,” he says. “PhysioNet lowers the fixed costs of testing wishful thinking, and that changes what science can do.”
The rise of AI
PhysioNet, once a repository primarily for those working in the biomedical signal processing and healthcare fields, has evolved over its 25 years. Initially, the holdings only included cardiovascular ECG data. PhysioNet is now the leading source of electronic health records, imaging data, and AI software and models.
Especially as artificial intelligence approaches took off, “society changed,” Heldt explained. Those who need signal processing data still use the PhysioNet database, but the user base has grown to include employees at major technology companies, educators, and physicians in all medical fields, as well as researchers in health-related machine learning and AI. Today, that latter group “dominates the user community,” Heldt said.
The platform hosts the highest quality datasets available for healthcare AI research, according to Natarajan, whose research spans AI, science, and medicine and has published several papers using its data.
“It’s been an important foundation that has driven all the progress in AI healthcare over the last decade,” Natarajan said. In addition to using PhysioNet data, he and his colleagues have contributed data to the platform, helping to build the supporting ecosystem that characterizes PhysioNet.
Looking to the coming decades, forum administrators like Heldt and Pollard envision continuing to expand its reach with an annual conference. The team is also preparing to test a new system that will allow users to annotate data and contribute their own information, enriching PhysioNet’s services for the next phase of the platform.
“The kind of research people want to do now requires interdisciplinary work. Mathematicians, computer scientists, doctors, chemists, and nurses must all come together and contribute their knowledge to develop algorithms that are useful to people,” said Pollard. “Society has matured, and advances in AI have expanded both the questions researchers can answer and the ones they believe are possible.”



