# v2.1 evidence-aligned figure captions

### Figure 1. Evidence layers and inference limits in NMD-VCell

**(a)** DMD-prior, HepG2 perturbation and muscle-context layers are shown with native units and denominators. The primary HepG2 audit comprises 145,473 cells, 2,160 eligible perturbation targets and five non-overlapping 432-target outer folds. **(b)** The evidence path separates the unique-target primary benchmark from historical sensitivity records. Every eligible target receives one primary out-of-fold prediction; the historical design reused 787 targets and never tested 514. **(c)** Descriptive evidence is supported at L1, bounded hypothesis triage is supported with caveats at L2 and mechanism or therapeutic claims at L3+ are unsupported. Linked evidence layers are not interpreted as equivalent validation.

### Figure 2. DMD transcriptional priors are source- and threshold-dependent

**(a)** Source-specific median effects for the 18 genes with the largest cross-source spread among genes measured in at least three sources. Colors are centered at zero and clipped at the 90th percentile of the absolute displayed effects; missing source-gene combinations are grey. **(b)** Pairwise Pearson, Spearman and cosine concordance for the four source vectors. B, baseline pseudobulk; D, delta/DID; S, SEMA3C state; N, NicheNet target state. **(c)** Fraction of genes assigned as direction-conflicted or insufficient as the minimum absolute median source-effect threshold changes. The vertical line marks `|effect|≥0.25`, at which 6,307 genes are conflicted, corresponding to 54.4% of multisource-evaluable genes. The panels describe sensitivity of an integrated direction prior and do not define a disease ground truth.

### Figure 3. Unique-target cross-fitting reveals modest predictive signal

**(a)** Historical repeated holdouts tested 1,646 of 2,160 eligible targets, reused 787 and never tested 514. The replacement five-fold design assigns all 2,160 targets to one non-overlapping outer fold (432 per fold) and uses 32 leakage-controlled external or curated pre-model features; DepMap and HepG2-wide outcome summaries are excluded. **(b)** Mean target-level paired effects with bootstrap 95% confidence intervals for the safe model and three negative controls. Negative RMSE differences and positive cosine differences favor the model. Filled points meet the intended within-model Holm-adjusted directional test. **(c,d)** Effects across frozen tertiles of the primary reliability measure (720 targets per tertile), with bootstrap 95% confidence intervals. Filled points indicate stratum support after within-stratifier Holm correction. The safe model reduces mean RMSE by 5.10% relative to zero and 0.41% relative to train mean, with a raw-cosine gain of 0.00164. Reliability strongly tracks versus-zero RMSE and residual cosine, but not improvement over train mean or raw-direction gain. These data support a statistically detectable but practically modest full-OOF benchmark signal, not strong directional prediction.

### Figure 4. Historical coverage-aware, same-target model comparison

Coverage-aware, same-target model comparison. **(a)** Number of targets available for Linear/PCA, Ridge, GEARS and scGPT in each frozen split. Horizontal lines show the across-split range, circles show individual splits and diamonds show medians; the x axis is logarithmic. **(b)** Available target count versus paired mean RMSE difference on targets shared with ridge. Positive comparator-minus-ridge values favor ridge. **(c)** Number of the five splits in which each comparator had lower RMSE, higher raw cosine or higher residual cosine than ridge. Counts report numerical direction and are not significance counts. **(d)** Split-specific comparator-minus-ridge RMSE differences with paired target-level bootstrap 95% confidence intervals. All contrasts are within split and within the same target set; reduced coverage is not treated as equivalent performance.

### Figure 5. Candidate priority is an evidence map, not a therapeutic ranking

**(a)** Evidence board for 21 candidate perturbations ordered by the full decision-score rank. Columns show priority tier; observed DMD-prior counteralignment score with its range across decision rules; DMD-source direction agreement; top-four and A-tier rule frequencies; GTEx skeletal-muscle median TPM; moderate-dependency and fusion-screen assessment/hit fields; and model-free-minus-full rank shift. A legacy field name containing `rescue` denotes DMD-prior counteralignment and not therapeutic rescue. A dash denotes unassessed fusion evidence. **(b)** Full decision score versus a score that removes model-derived components. Labels mark the four leading genes; ρ is Spearman correlation. **(c)** Counts of Priority-A/B/C rows, fusion-screen coverage and observed fusion hits, together with the L2 claim ceiling. Priority is an operational validation queue, not an efficacy or safety estimate.

### Figure 6. External datasets define an applicability boundary, not validation

**(a)** Fraction of all measured targets, benchmark-eligible targets and the 21 candidates exceeding GTEx skeletal-muscle TPM thresholds of 0.1, 1 and 5. **(b)** Fusion-hit odds ratios for the feasible queue versus other overlapping genes before and after broad-essentiality and moderate-dependency restrictions. Lines are Woolf 95% confidence intervals; q values are Holm adjusted across four overlapping strata. **(c)** Cell support for the eight of 64 neuromuscular sentinel genes with a direct HepG2 perturbation. The dashed line marks 20 cells; none of five DMD/BMD sentinels has direct coverage. **(d)** Patient-muscle and organoid/model Pearson, Spearman and cosine concordance with bootstrap 95% confidence intervals and Holm-adjusted P values. These panels delimit expression visibility, exclusion risk and missing transfer evidence; they are not muscle perturbation validation.

### Supplementary Figure S1. Data provenance and source-scale quality control

**(a)** Log-scaled counts for HepG2 cells, DMD genes, perturbations, benchmark-eligible targets, GTEx-matched targets, GSE293514 overlap genes, neuromuscular sentinels and candidate genes. Counts refer to different statistical units and are not pooled. **(b)** Distribution of cells per non-control HepG2 perturbation on a logarithmic x axis. Dashed and dotted lines mark the 20-cell benchmark threshold and 80-cell candidate-display threshold. **(c)** Matrix showing whether each evidence layer is used for description, benchmarking, triage or context. Analytical roles follow provenance.

### Supplementary Figure S2. Historical repeated-split design audit

Repeated split design and target reuse. **(a)** Pairwise Jaccard similarity among the five 540-gene held-out sets. Split labels show the final two digits of seeds 20260712–20260716. **(b)** Number of eligible targets held out zero to five times across the frozen splits. A total of 787/2,160 targets recur in more than one test set. **(c)** Distribution of cell support for held-out targets in each split; boxes show the interquartile range and median, whiskers follow the standard box-plot convention, and the dashed line marks 20 cells. Repeated splits quantify assignment sensitivity, not independent replication.

### Supplementary Figure S3. Measurement reliability under three replicate definitions

**(a–c)** Violin distributions of target-level replicate cosine, RMSE and sign concordance under cell split-half, batch split and construct-level guide-pair definitions. Horizontal references mark zero cosine and 0.5 sign concordance. **(d)** Median cell- and batch-split cosine after targets are grouped into six cell-support quantiles. Higher cell support is associated with greater aggregate reliability, but target-level heterogeneity remains and guide-pair agreement stays near zero.

### Supplementary Figure S4. Historical repeated-split endpoint detail

Ridge split and regularization detail. **(a)** Mean RMSE for ridge, zero delta and train mean in each frozen split. **(b)** Mean raw cosine for ridge and train mean together with ridge residual cosine after subtraction of the training-set common delta. **(c)** Inner-cross-validation mean RMSE across ridge penalties 0.1–1,000 for each split. **(d)** Mean paired values for RMSE versus zero, RMSE versus train mean, raw cosine versus train mean and residual cosine versus zero; lines span the minimum and maximum split-specific bootstrap 95% confidence limits. Perturbation targets, not cells, are the inferential unit.

### Supplementary Figure S5. Historical same-target comparator forests

Complete same-target comparator analysis. **(a–c)** Split-specific comparator-minus-ridge differences with target-level bootstrap 95% confidence intervals for RMSE, raw cosine and residual cosine. RMSE and cosine have different favorable directions and must be interpreted separately. **(d)** Number of the five split-level contrasts passing Holm correction for each comparator and endpoint. The blue scale denotes evidence for a difference, not evidence that the comparator is better.

### Supplementary Figure S6. Historical coverage and shared-error sensitivity

Coverage and shared-error sensitivity. **(a)** Native model coverage and the common four-model target set in each split. **(b)** Median RMSE difference relative to zero on each model's available set versus the common four-model set; the diagonal denotes no summary shift after restriction. **(c)** Range and mean of off-diagonal same-target model-pair error correlations under four error definitions. Common-set restriction and baseline definition change the apparent similarity of model errors.

### Supplementary Figure S7. DMD source concordance and meta-analysis boundary

**(a–c)** Full 4×4 Pearson, Spearman and cosine concordance matrices for the four DMD source pipelines. Values are calculated on genes shared by each source pair. **(d)** Number of nested contexts and number of genes represented in each harmonized source. None provides independent source-level standard errors. The context rows cannot be treated as independent studies for random-effects meta-analysis.

### Supplementary Figure S8. Feature ablation and leakage sensitivity

**(a)** Five-split mean and range of relative RMSE improvement versus zero delta for 15 feature configurations and negative controls. **(b)** Sixteen largest absolute ablation-versus-full effects across RMSE and residual-cosine comparisons. Positive values favor the ablation; circles denote effects significant in all five split-level Holm-corrected tests and squares denote other effects. **(c)** Counts of features classified as external/curated without held-out target response or as carrying non-zero global dataset-summary risk. **(d)** Feature counts by provenance group. Risk labels describe provenance semantics and do not prove leakage.

### Supplementary Figure S9. Candidate multiverse and evidence completeness

**(a)** Full-score rank versus model-free rank for the 21 candidates; color intensity reflects absolute rank shift and selected genes are labelled. **(b)** Observed DMD-prior counteralignment score and its minimum-to-maximum range across decision rules. **(c)** Median and range of source-specific DMD effects for each candidate. **(d)** Evidence-coverage matrix for positive observed score, at least 80 cells, GTEx muscle TPM above 1, repeated-benchmark inclusion and fusion-screen status. Plus signs denote pass, x denotes fail and dashes denote not assessed. These axes are not collapsed into a therapeutic score.

### Supplementary Figure S10. Skeletal-muscle expression and context transfer

**(a)** Distribution of log10(GTEx skeletal-muscle median TPM + 0.05) across measured targets; the dashed line marks TPM 1. **(b)** Hexbin density of GTEx muscle expression versus the fraction of HepG2 reference cells expressing each gene. Red outlines mark the 21 candidates. **(c)** GTEx muscle TPM for the candidate genes, ordered by expression; green bars exceed TPM 1 and amber bars do not. **(d)** GTEx symbol-match, TPM>1 and TPM>5 coverage for all measured targets, benchmark-eligible targets, candidates and GSE293514-overlap genes. Expression supports context plausibility only.

### Supplementary Figure S11. Fusion-screen and external-context boundary

**(a)** Fusion-hit odds ratios and 95% confidence intervals for the feasible queue before and after essentiality restrictions. **(b)** Continuous observed HepG2 counteralignment score versus the GSE293514 positive-fusion log-fold change for 1,575 overlapping genes; feasible-queue genes and fusion hits are marked. **(c)** Patient and organoid/model Pearson, Spearman and cosine estimates with bootstrap 95% confidence intervals. **(d)** Denominators for the published screen, HepG2 overlap, fusion hits in the overlap, feasible-queue overlap and queue risk flags. Association and expression-context concordance are orthogonal filters, not perturbational muscle validation.

### Supplementary Figure S12. Resource verification and reconstruction boundary

**(a)** Counts of release-manifest files in the 12 largest artifact categories. **(b)** Counts of PASS, WARN and FAIL states among 19 quality gates. **(c)** Numbers of reproducibility capabilities in DOI-pending, browser-QA-pending, blocked, partial, partial-remote and local-clean-rebuild states. **(d)** Percentage of manifest entries present, entries with populated SHA-256 values, independent verification checks passing and quality gates passing. Automated integrity checks do not replace the pending manual browser QA or DOI completion.

### Supplementary Figure S13. Sampling, coverage and reliability audit

**(a)** Target accounting for historical repeated holdouts and final unique-target cross-fitting. **(b)** Bootstrap 95% confidence intervals for the safe-model raw-cosine difference in the superseded unbalanced implementation and corrected balanced implementation. **(c)** scGPT and GEARS availability across the same 2,160-target denominator. **(d)** Standardized mean differences between covered and uncovered targets for prespecified pre-model variables, with bootstrap 95% confidence intervals. Filled points and asterisks mark signals detected after within-model Holm correction. No audited pre-model selection variable is detected for scGPT, whereas GEARS coverage is enriched for DMD-prior and DepMap availability. **(e)** Pearson correlations between four reliability definitions and target-level model effects; asterisks mark Holm-adjusted trend P<0.05. Guide-pair analysis is exploratory because only 133 targets are available. **(f)** Target-level effects against the primary reliability measure; black points are 20 equal-frequency bin means and dashed lines are frozen tertile cutpoints.
