Data Availability

The Journal of Healthcare Analytics and Informatics requires authors to make the data, code, and other research artifacts that support the conclusions of their work available to readers, subject to clearly stated and legitimate restrictions. The Journal recognizes that health and clinical data often cannot be shared openly, and it aligns its policy with the FAIR principles (Findable, Accessible, Interoperable, Reusable), which support responsible sharing even where data cannot be made fully public.

Mandatory Data Availability Statement

Every manuscript must include a Data Availability Statement, placed in a dedicated section between the Conclusions and the References. The statement specifies what data underlie the findings reported in the article and how those data may be accessed.

Acceptable forms of data availability

Authors should select the option below that best fits their study. Multiple options may apply.

  1. Data in a public repository. The data are deposited in a public or controlled-access repository (Zenodo, Figshare, Dryad, PhysioNet, the NIH database of Genotypes and Phenotypes (dbGaP), or a discipline-specific repository) with a persistent identifier. Provide the repository name, the persistent identifier or URL, and the access terms.
  2. Data included in the article and supplementary information. The data are fully reported in the article and its supplementary files; no separate dataset is required.
  3. Data available on reasonable request. The data are available from the corresponding author on reasonable request, subject to specified conditions (for example, an institutional data-sharing agreement or ethics committee approval).
  4. Restricted data. The data are subject to restrictions (for example, patient confidentiality, personal health data protection, or ethics approval limits). Specify the nature of the restriction and the conditions under which the data may be obtained. For health data, state the governing framework where relevant (for example, HIPAA, the GDPR, or the applicable national health data regulation).
  5. No new data created. The study used only previously published or publicly available data; cite the original sources.

Patient privacy and de-identification

Where data derive from human participants or patient records, authors must confirm that any shared data have been de-identified in accordance with the applicable data-protection framework, and that sharing is consistent with the consent obtained and the approving ethics committee's terms. Authors must not deposit identifiable patient data in any open repository. Where individual-level data cannot be shared, authors are encouraged to share aggregate data, summary statistics, or synthetic datasets that allow assessment of the results without compromising privacy.

Code and software availability

Where the study uses or produces software or code, authors are required to:

  • Cite all third-party software used, with version numbers, in the methods section;
  • Make any custom code that is essential to reproducing the reported results available in a public repository (GitHub, GitLab, Zenodo) with a persistent identifier;
  • Specify the license under which the code is released. Recommended licenses are MIT, Apache 2.0, GPL, or BSD for code, and CC BY for data.

Materials availability

For studies that involve specific physical or digital materials such as trained clinical models, datasets, or custom test environments, authors should describe the materials in sufficient detail for replication and indicate whether the materials are available on request.

Reproducibility

The Journal encourages submissions to include a Reproducibility section describing the computational environment (operating system, library versions, hardware), the random seeds used, and the steps required to reproduce the reported results from the published code and data. This is particularly relevant for work involving clinical prediction models, machine learning pipelines, and large-scale analyses of health data, where full specification of the environment, cohort definitions, and parameters is essential to independent verification. Where the underlying patient data cannot be shared, authors should provide enough methodological detail, including data dictionaries, inclusion and exclusion criteria, and preprocessing steps, to allow the analysis to be reproduced on equivalent data.

Example statements

Example 1: public repository. "The de-identified dataset analyzed in this study is available in the Zenodo repository at https://doi.org/10.5281/zenodo.EXAMPLE under a CC BY 4.0 license. Custom analysis code is available at https://github.com/EXAMPLE under an MIT license."

Example 2: restricted health data. "The data that support the findings of this study are not publicly available because they contain information that could compromise patient privacy. De-identified data are available from the corresponding author on reasonable request and subject to approval by the relevant ethics committee and a data-sharing agreement."

Example 3: no new data. "No new data were created or analyzed in this study. All data referenced in this work are publicly available from the cited sources."