This is the dataset associated with our paper titled "Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets". Artificial Intelligence (AI) is rapidly transforming healthcare, but also raising concerns about algorithmic biases that mostly stem from the training data. It is widely supported that transparent dataset documentation is key to enabling responsible AI development. Several standardized dataset documentation approaches have been established, such as Datasheet, Dataset Nutrition Label, Accountability Documentation, Healthsheet, and Data Card. However, their suitability and usage for health datasets remain unclear. In this paper, we compared all five approaches and evaluated their alignment with the STANDING Together Recommendations for Documentation of Health Datasets. We also investigated their real-world usage and gathered insights from generators and consumers of health datasets.
See this inventory for all related resources, including the paper.
This dataset is structured according to the SPARC Data Structure v3 (c.f. related manuscript and documentation) and curated using the SPARC data curation software SODA v16.3.2.
Simply download this repository to use the data files. Alternatively, you can download it from Zenodo. There is a manifest.xlsx file that describe all the files. See this inventory for a link to the Jupyter notebooks we developed to collect, analyze, and visualize the data.
This work is licensed under a Creative Commons Attribution 4.0 International License.
Use the GitHub issues for submitting feedback or making suggestions. You can also fork the repository and submit a pull request with suggestions. Alternatively, you can send an email to bpatel@calmi2.org
If you use this dataset, please cite the related paper (it will be listed here when available) and also cite this dataset as indicated on the Zenodo page of the dataset.
