top of page

Open data is transforming medical AI; but governance is struggling to keep up.

Original paper: Copycats: the many lives of a publicly available medical imaging dataset

Authors: Amelia Jiménez-Sánchez, Natalia-Rozalia Avlona, Dovile Juodelyte, Théo Sourget, Caroline Vang-Larsen, Anna Rogers, Hubert Zając, Veronika Cheplygina



About the researcher

Amelia Jiménez-Sánchez, Postdoctoral Researcher, University of Barcelona, Spain. Find out more here: https://ameliajimenez.github.io/ 


What problem does this paper address, and why does it matter?

This paper investigates what happens after medical datasets are shared publicly, especially when they are copied, re-uploaded, and reused across platforms like Kaggle and HuggingFace. It matters because this may not only violate licenses and contribute to the reproducibility crisis in AI, but also lead to overoptimistic results, which can have serious real-world consequences, directly affecting patients in healthcare AI.


What did this research discover/create?

This paper reviews the 30 most cited computer vision, natural language processing and medical imaging datasets by selecting the top-10 datasets for each field by querying Papers with Code with "Images", "Text", and "Medical" in the Modality field. 


We find that datasets have vague licenses, lack of persistent identifiers and storage, duplicates, and missing metadata across platforms. We highlight differences between medical imaging and computer vision datasets, particularly in the potentially harmful downstream effects from poor adoption of recommended dataset management practices. 


How could this research impact real-world applications?

This study could push researchers and platform-owners to adopt better practices for dataset documentation, licensing, and monitoring, leading to more reliable medical AI models and reducing risks like biased results or unsafe clinical decisions.


Who should care about this work? 

Practitioners, policymakers, general public, researchers.


What is noteworthy about this research?

This work presents a case study of how a skin lesion challenge dataset (ISIC) has been duplicated 640 times on Kaggle. While the size of the original ISIC datasets is 38 GB, Kaggle stores 2.35 TB of data (see Figure 2). Several highly downloaded versions (≈13k downloads) lack original sources or license information. This proliferation of duplicate datasets not only wastes resources but also poses a significant impediment to the reproducibility of research outcomes.


What's the ONE key takeaway you want people to remember? 

Medical imaging is not just "small computer vision".


Funding/sponsorship information for this research

This project has received funding from the Independent Research Council Denmark (DFF) Inge Lehmann 1134-00017B. 


Read More


Have research to share or recommend? https://forms.gle/kaTJEd3gfJV8vtDx8 






Related Posts

See All

2 Comments


When we needed to quickly obtain detailed information from another clinic for a second opinion, serious delays and bureaucratic obstacles appeared. Сaremount medical helped navigate real patient reviews and find an effective solution. Thanks to this we promptly received the necessary data, were able to make an informed treatment decision and now feel much calmer, with clearer access to medical information and greater trust in the process.

Like

The scale of dataset duplication and its impact on medical AI research and clinical decision-making is definitely eye-opening. Whenever I want to switch gears for a while, I enjoy That's Not My Neighbor a clever mystery game that keeps me focused on spotting inconsistencies and hidden clues.

Like
bottom of page