top of page

Open data is transforming medical AI; but governance is struggling to keep up.

Jun 17
2 min read

Original paper: Copycats: the many lives of a publicly available medical imaging dataset

Authors: Amelia Jiménez-Sánchez, Natalia-Rozalia Avlona, Dovile Juodelyte, Théo Sourget, Caroline Vang-Larsen, Anna Rogers, Hubert Zając, Veronika Cheplygina



About the researcher

Amelia Jiménez-Sánchez, Postdoctoral Researcher, University of Barcelona, Spain. Find out more here: https://ameliajimenez.github.io/ 


What problem does this paper address, and why does it matter?

This paper investigates what happens after medical datasets are shared publicly, especially when they are copied, re-uploaded, and reused across platforms like Kaggle and HuggingFace. It matters because this may not only violate licenses and contribute to the reproducibility crisis in AI, but also lead to overoptimistic results, which can have serious real-world consequences, directly affecting patients in healthcare AI.


What did this research discover/create?

This paper reviews the 30 most cited computer vision, natural language processing and medical imaging datasets by selecting the top-10 datasets for each field by querying Papers with Code with "Images", "Text", and "Medical" in the Modality field. 


We find that datasets have vague licenses, lack of persistent identifiers and storage, duplicates, and missing metadata across platforms. We highlight differences between medical imaging and computer vision datasets, particularly in the potentially harmful downstream effects from poor adoption of recommended dataset management practices. 


How could this research impact real-world applications?

This study could push researchers and platform-owners to adopt better practices for dataset documentation, licensing, and monitoring, leading to more reliable medical AI models and reducing risks like biased results or unsafe clinical decisions.


Who should care about this work? 

Practitioners, policymakers, general public, researchers.


What is noteworthy about this research?

This work presents a case study of how a skin lesion challenge dataset (ISIC) has been duplicated 640 times on Kaggle. While the size of the original ISIC datasets is 38 GB, Kaggle stores 2.35 TB of data (see Figure 2). Several highly downloaded versions (≈13k downloads) lack original sources or license information. This proliferation of duplicate datasets not only wastes resources but also poses a significant impediment to the reproducibility of research outcomes.


What's the ONE key takeaway you want people to remember? 

Medical imaging is not just "small computer vision".


Funding/sponsorship information for this research

This project has received funding from the Independent Research Council Denmark (DFF) Inge Lehmann 1134-00017B. 


Read More


Have research to share or recommend? https://forms.gle/kaTJEd3gfJV8vtDx8 






Related Posts

See All

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page