Digital health Medical data in research
Medical data is primarily used to provide medical care to patients. However, it can also be used to obtain new findings about health and illnesses as it contains valuable information for medical research.
At a glance
- Medical data is primarily collected and used to provide medical care to patients.
- Medical data harbors major potential for medical research and other purposes for the common good.
- Research involving medical data can improve healthcare and lead to a better understanding of illnesses, for example.
- Strict rules apply to the use of medical data for research, for example to ensure data protection.
What is medical data?
Medical data is all data that could provide information about a person’s health. This includes information such as test results, diagnoses or prescribed treatments.
The data is primarily collected and documented by staff in medical facilities, where it is used to provide medical care to patients. If data is used for the purpose for which it was collected, this is referred to as its primary use.
More precise information about what is classed as medical data and how it is collected on an everyday basis can be found in the article Medical data in everyday life.
However, medical data also has major potential for research. After being collected and used for the purpose of a patient’s medical care, it can also be used for additional purposes for the common good. This is referred to as its secondary use.
Good to know: medical data from everyday medical life contains valuable information, including because it directly depicts the healthcare provided by hospitals and medical practices on a daily basis.
The use of medical data for research can improve healthcare, screening processes and the treatment of patients. Correlations between different medical conditions and their risk factors can be uncovered and researched based on medical data from everyday healthcare. To that end, it should in future be possible to share medical data from electronic patient records for research.
Who uses medical data for what purposes?
Medical data should be made more usable for scientific research, improving healthcare and other purposes for the common good. That means that the use of the medical data should benefit as many people as possible rather than just individuals or companies.
Since October 2025, the Health Data Lab (HDL) has enabled research groups, institutions and companies to use medical data in a secure working environment for certain purposes, including:
- improving the provision of healthcare
- hospital and care structure planning
- scientific research in the field of medicine, medical care and pure research
- national and state reporting and information about citizen health
- the fulfillment of legal responsibilities regarding public health
- the development and monitoring of new treatment methods and medication
- the development of digital health applications and artificial intelligence systems within the field of healthcare
Healthcare facilities, such as university clinics and other hospitals, which independently collect and save medical data can also use this data for additional purposes. For example, to conduct their own scientific research projects, improve patient care or conduct healthcare reporting. The data must be anonymized or pseudonymized prior to any such use.
Answers to questions about the Health Data Lab can be found on the Federal Ministry of Health (BMG) website.
Under what conditions can medical data be used for research?
Medical data can contain sensitive personal information. If it is to be used for research, certain conditions must therefore be fulfilled.
Personal information, such as a person’s name or date of birth, is removed from the medical data before this is used for research purposes. This prevents any insights into the person who provided the data.
Researchers who want to use medical data have to submit an application for their research project. This must indicate that the medical data will only be processed for permitted purposes. If the application is approved, the researchers must only use the medical data for the approved research project and not share it with third parties. Applications will only be approved if projects do not pose any undue risk to the protection of the data.
Medical data can only be used for purposes permitted by law. Furthermore, it can naturally only be used for the purpose for which it was originally collected, i.e. usually for providing care to a patient.
Important: people’s medical data must not be used to their detriment. The use of medical data for market research or to advertise medication is also not permitted.
If a person’s medical data has been collected in different contexts, it can only be combined for specific research projects. This means, for example, that data from an electronic patient record can only be combined with data from a cancer register if this is necessary for a specific research project. This data will otherwise remain stored in different places.
What are anonymization and pseudonymization?
There are basically two ways to protect personal information. With anonymization, certain information, such as a person’s name or address, is removed from the data record. Following anonymization, it is extremely difficult to associate the data with a specific person. This is an effective way of protecting people’s personal information.
With pseudonymization, sensitive personal information about an individual is removed from a data record. However, this information is replaced with a pseudonym.
A pseudonym could, for example, be an alphanumeric sequence, which represents a data record. The people who use pseudonymized data for their research cannot assign the data to individual real people. However, the assigned alphanumeric sequence does still make it possible to link together different data records from a single person. This is not possible with anonymization.
One particularly useful reason for being able to link data records is to add newly acquired medical data to previously collected medical data. This would make it possible, for example, to link information about a specific form of treatment for an illness with the results of a check-up after several years so the effectiveness of different forms of treatment can be compared.
What consent procedures are possible?
An opt-out rule applies to the use of certain medical data. For example, this will be the case for the use of data from electronic patient records (ePAs). Patients can object to the use of their data from the ePA for research purposes. If they do not do this, their consent is assumed. This principle is known as an opt-out rule.
The opposite of the opt-out procedure is an opt-in procedure. With an opt-in procedure, people must actively provide their consent for data to be used.
The opt-out procedure should make it easier to use medical data for research purposes. Opt-out rules generally lead to data being available from a larger number of people than when each and every individual has to actively provide their consent.
The more data available from different people, the better suited it is for certain research projects. The validity of scientific studies can often be improved by larger data volumes. This is important to enable study results to be applicable to the general population.
How else is medical data protected?
Medical data is a special category of personal data. This highly sensitive data is strictly protected and can only be processed under certain conditions. The processing of medical data is regulated by the European General Data Protection Regulation and the German Federal Data Protection Act (BDSG) among other legislation.
Further information on the protection of medical data can be found in the article Medical data in everyday life.
What is the potential of research with medical data?
Medical data contains valuable information with regard to researching health and illness as well as for improving medical care. Using data that has already been collected anyway can also save on the costs that would have been incurred through other forms of data collection for studies.
It is particularly important to use data from Germany and Europe for research. Such data best depicts regional environmental influences, the prevalence of conditions and local healthcare. On the other hand, research results from other countries cannot always be transferred to Germany.
Medical data is particularly important for research in relation to
- medical care (healthcare research)
- personalized medicine
- rare conditions
- clinical decision support systems
- and many other topics
Healthcare research
Healthcare research generally deals with ways to improve the provision of healthcare for everyone. For example, it is about who has access to certain medical services or how limited resources within the healthcare system can be fairly distributed.
Medical data such as that obtained from the electronic patient record depicts the actual everyday situation within the healthcare system. This makes such data particularly valuable for healthcare research.
Personalized medicine
Personalized medicine is about customizing medical treatments based on people’s individual needs and the progression of their condition. For example, medication can work differently depending on a person’s age, gender or other individual factors.
The analysis of large volumes of data makes it possible to compare how medication acts under certain conditions or how different people’s blood values change as a result of treatment. These results improve the ability to predict how individual patients will respond to a treatment and make it possible to modify procedures accordingly.
Rare conditions
When it comes to researching rare conditions, large datasets are required. This is because if medical data from many different people is incorporated into a study, there is a higher probability that several of these people will have been affected by the same rare condition. The outlook of such conditions can then be researched and new findings about potential treatments can be obtained.
Clinical decision support systems
Systems for supporting decisions are computer systems designed to assist doctors with patients’ everyday clinical care.
Such systems provide a warning, for example, if they notice potential interactions between medication on a patient’s medication plan. The medical personnel will receive a notification on their computer and can modify the medication if required. Decision support systems are therefore designed to help avoid errors, among other factors.
Medical data is also used to develop these kinds of systems. Computers can use medical data, for example, to learn of correlations between symptoms and illnesses. Medical personnel can subsequently be supported by the computer system when diagnosing their patients.
What are big data and artificial intelligence?
The increasing digitalization of everyday life means that lots of different data is constantly generated and saved in both the healthcare sector and other areas. These large volumes of data are also known as big data.
Artificial intelligence (AI) can help when analyzing these volumes of data. This is a sub-area of information technology and includes computer systems that are able to learn and resolve problems.
AI-based applications can use large data volumes to detect even weak correlations. This could make it possible to detect even previously unknown medical condition precursors, for example.
Self-learning AI systems are initially trained using data. This teaches them the rules and templates that they should later apply. The more good-quality data available for this training, the more reliable the AI predictions.
What results have been achieved to date through research using medical data?
There are already numerous examples of research using medical data that was originally collected for other purposes.
For example, data that is routinely collected in Germany by health insurance funds is used in pseudonymized form to depict everyday healthcare provision. The data makes it possible to obtain information about the use of medication for certain medical conditions or reasons for people’s incapacity to work. The latter information can be used, among other things, as a basis for developing targeted workplace health promotion measures.
Further research projects involving the secondary use of medical data have been established in the field of cancer or with regard to treatment options for thrombosis, for example.
Bowel cancer screening
A new computer model has been developed based on blood tests and with the help of information from cancer registers. This process used anonymized data from over 450,000 people. All the model needs to detect a heightened risk of bowel cancer are the results of a routine blood test and a person’s age and gender. Any affected patients can then be notified about their heightened risk of bowel cancer and undergo further tests and examinations if necessary.
Cancer
To identify temporal correlations between cancer and other illnesses, anonymized data from doctor’s appointments by people in Taiwan were analyzed. The knowledge about temporal correlations can help doctors identify a heightened risk of cancer in good time and treat accompanying illnesses at an early stage in the event of a cancer diagnosis.
Drug therapy following thrombosis
Information obtained within the scope of everyday healthcare in Germany and Canada was analyzed for a study about therapy following thrombosis. This analysis of anonymized diagnoses and prescribed medication made it possible to compare the side-effects of various anticoagulation drugs, for example.
- Bundesministerium für Bildung und Forschung. Versorgungsforschung. Aufgerufen am 20.11.2025.
- Bundesministerium für Bildung und Forschung, Bundesministerium für Gesundheit, Bundesministerium für Wirtschaft und Energie. Daten helfen heilen. Aufgerufen am 20.11.2025.
- Bundesministerium für Gesundheit. Daten für die Forschung und Versorgung. Aufgerufen am 20.11.2025.
- Bundesministerium für Gesundheit. Fragen und Antworten zum Gesundheitsdatennutzungsgesetz. Aufgerufen am 20.11.2025.
- Bundesministerium für Gesundheit. Gesundheitsdatennutzungsgesetz. Aufgerufen am 20.11.2025.
- Bundesministerium für Gesundheit. Wissenschaftliches Gutachten „Datenspende“. Stand 03/2020.
- Bundesministerium der Justiz. Gesundheitsdatennutzungsgesetz. Stand 03/2024.
- Datenschutzgrundverordnung. Erwägungsgrund 35. Aufgerufen am 20.11.2025.
- Data Saves Lives. DSL DE Logbuch 2022/2023. Aufgerufen am 20.11.2025.
- Data Saves Lives. Early Detection of a Cancer Killer. Aufgerufen am 20.11.2025.
- Data Saves Lives. „Big Data” used for the early identification of other diseases associated with cancer. Aufgerufen am 20.11.2025.
- Datenschutzgrundverordnung. Artikel 4 Begriffsbestimmungen. Aufgerufen am 20.11.2025.
- Datenschutzgrundverordnung. Artikel 9 Verarbeitung besonderer Kategorien personenbezogener Daten. Aufgerufen am 20.11.2025.
- Douros A, Basedow F, Cui Y, Walker J, Enders D, Tagalakis V. Effectiveness and safety of direct oral anticoagulants with antiplatelet agents in patients with venous thromboembolism: A multi‐database cohort study. Res Pract Thromb Haemost. 6:e12643. 2022.
- Fraunhofer-Institut. Diagnose auf Knopfdruck. Aufgerufen am 20.11.2025.
- Sondermann W, Ventzke J, Matusiewicz D, Körber A. Analyse der pharmazeutischen Versorgungssituation von Patienten mit Psoriasis‐Arthritis auf Basis von Routinedaten der Gesetzlichen Krankenversicherung. Journal der Deutschen Dermatologischen Gesellschaft. 16(3), 285–296. 2018.
- Sozialgesetzbuch Fünftes Buch (V). §303e Datenverarbeitung. Aufgerufen am 20.11.2025.
- Statista. Definition Repräsentativität. Aufgerufen am 20.11.2025.
- Zoike E, Bödeker W. Berufliche Tätigkeit und Arbeitsunfähigkeit. Bundesgesundheitsbl. 51, 1155–1163. 2008.
As at: