We conducted a scoping review to characterize code sharing practices for papers developing or evaluating multivariable prediction models in healthcare. We defined a multivariable prediction model as any combination or equation of two or more predictors that is used for individualized predictions to estimate an individual’s probability of the presence of a particular health condition (diagnostic) or whether a particular outcome will occur in the future (prognostic). The review addresses two key questions. First, what proportion of multivariable prediction model studies that cite the TRIPOD or TRIPOD+AI statements report on the availability of analytical code? Second, among studies providing accessible code, what are the structural and documentation characteristics of the code? The review was conducted in accordance with the PRISMA extension for scoping reviews32 (Supplementary Table 1). The protocol for this review was registered (INPLASY202620080)33.
Search strategy and selection criteria
The cohort comprised all PubMed-indexed articles that cited the TRIPOD or TRIPOD+AI statements (www.tripod-statement.org) as of 11 August 2025. Focusing on TRIPOD-citing studies provides a defined and policy-relevant subset of prediction model research in which reporting guidance is already acknowledged. As such, this sampling frame represents a conservative estimate of code sharing practices among studies attentive to reporting standards.
Owing to simultaneous publication across different journals, TRIPOD and TRIPOD+AI have multiple entries in PubMed (12 and 2 entries, respectively). For each entry, citation lists were downloaded from the PubMed ‘Cited by’ section and aggregated programmatically. Duplicate records were removed. We included primary research articles that developed, validated or updated a multivariable prediction model using statistical or machine-learning methods, consistent with definitions in the TRIPOD-Code protocol19. To enable reproducibility without subscription barriers, we restricted inclusion to articles retrievable through the PubMed Central Open Access API. No additional exclusion criteria were applied at this stage.
Data analysis
Data analysis comprised two components: (1) article-level screening and metadata extraction and (2) characterization of associated code repositories. An LLM-assisted pipeline was implemented using a predefined structured output schema (Supplementary Notes 2 and 3). Prompts were iteratively refined through trial and error on a small set of articles and repositories before being fixed for validation and full-cohort analysis. All model outputs were validated against human-annotated subsets. The codebook used by the annotators is provided in Supplementary Note 4.
Article screening and metadata extraction
Full-text articles were retrieved using the PubMed Central Open Access API. Figures, tables and appendices were excluded from the extraction. Eligibility screening and metadata extraction were performed using an LLM (GPT-5.2 2025-12-11). The LLM was prompted using a structured schema to determine eligibility, identify code-availability statements, extract repository links and record the country of the first author’s institution (Supplementary Note 2).
To evaluate the performance of the automated pipeline, two independent reviewers with relevant expertise in prediction model research, clinical AI and computational methods, including the development and evaluation of analytical code (T. S., R. G.) manually annotated 500 randomly selected articles. The sample size was selected pragmatically to balance the need for a sufficiently large and diverse validation set with the feasibility of detailed manual annotation. Disagreements were resolved through group discussion involving both annotators. Performance of the pipeline was then evaluated against the human labels. The LLM achieved a weighted F1 score of 0.97 on the selection criteria and a 92.3% accuracy on repository link retrieval. Full evaluation metrics and error analysis are provided in Supplementary Note 5.
Repositories identified during article screening were retrieved using a custom repository retrieval utility (described below). The utility supported links from GitHub, GitLab, Gitee, Zenodo, Figshare, Open Science Framework (OSF) and digital object identifiers (DOIs) resolving to these providers. Repositories were retrieved as of the date of analysis and processed from their default branch. An article was considered to be sharing code if the retrieval utility successfully downloaded repository contents and detected at least one non-empty source-code file. We also counted articles that stated code was available in supplementary materials, or that linked to a repository hosted on a provider unsupported by our retrieval utility. For these latter cases, we could not independently verify accessibility or content, but we included them to avoid underestimating reported code availability. Code availability was analyzed by publication year, journal and country of first author affiliation. The location and wording of code-sharing statements were also categorized.
Temporal trends in code sharing were evaluated using binary logistic regression, with code-sharing status as the outcome and publication year modeled as a continuous predictor. Results are reported as ORs with 95% CIs. Code-sharing prevalence was compared between articles citing only TRIPOD and those citing only TRIPOD+AI using a two-sided Pearson χ2 test; articles citing both statements were excluded from this comparison. The comparison between TRIPOD-only and TRIPOD+AI-only articles was additionally evaluated using binary logistic regression with publication year included as a continuous covariate, to account for temporal differences between the two groups. Differences in code-sharing prevalence across countries were evaluated using a global Pearson χ2 test of independence among countries represented by more than 10 articles. Differences across journals were similarly evaluated among journals represented by more than 10 articles; because 42.3% of expected cell counts were below five, the journal-level P value was estimated using 100,000 Monte Carlo permutations.
Once code repositories were identified, they were explored and characterized. A repository retrieval utility was developed to compile each repository’s content into a structured text file. These files were subsequently processed by an LLM to extract code properties.
Code repository characterization
Repositories retrieved during article screening were subsequently characterized to assess code documentation, structure and features relevant to reproducibility. A custom utility compiled repository content and metadata into a textual representation according to predefined inclusion rules. The utility first outputs the full repository file tree, ensuring that file presence was always visible to the model regardless of subsequent truncation. README files and code files were included in full. Large binary files (for example, datasets, model checkpoints and compiled artifacts) were excluded. Other file types were truncated to a maximum of 3,000 tokens to avoid exceeding the model context window. Further details on the repository retrieval utility and validation of its extraction approach are provided in Supplementary Note 6.
Characterization was performed using an LLM (GPT-5.2 2025-12-11) prompted with a structured schema defining 14 repository features previously determined by the TRIPOD-Code executive committee (Supplementary Note 3). These features captured elements relevant to computational reproducibility, including the presence of documentation (for example, README), licensing information, specification of dependencies, test frameworks and indications of version control. README purpose and expected outputs were assessed only when a README file was present, dependency versioning only when dependencies were specified and seed control only when stochastic processes were identified; all other features were assessed for every repository. In addition to the structured feature extraction, the LLM generated brief free-text summaries of notable strengths and weaknesses of each repository for qualitative review.
To evaluate performance, two independent reviewers with relevant expertise in prediction model research, clinical AI, and computational methods, including the development and evaluation of analytical code (T. S., R. G.) annotated all retrieved repositories (n = 35), according to the same predefined schema. Disagreements were resolved through group discussion involving both annotators. Model performance was assessed against the human labels, yielding a weighted F1 score of 0.83 over all features. Performance was highest for objective features such as the presence of a README file (F1 = 1.0), LICENSE file (F1 = 1.0) and tests (F1 = 1.0), and lower for more subjective criteria such as adequacy of documentation (F1 = 0.71). Detailed performance metrics and error analysis are provided in Supplementary Note 7. Repository characteristics were summarized descriptively at the cohort level and analyzed in relation to the journal of publication and year.
Temporal trends in Python and R use were evaluated using separate binary logistic regression models, with the presence of each language within a repository as the outcome and publication year modeled as a continuous predictor. Because repositories could contain both languages, Python and R were analyzed as non-mutually-exclusive binary characteristics, and P values were adjusted across the two models using the Benjamini–Hochberg false-discovery-rate procedure.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
