# Major-revision response matrix

Revision target: `v0.6-audit-route-20260716`

Submission category: regular **NAR Data Resources and Analyses**.

Scientific identity: an auditable, direction-aware perturbation-model benchmark and DMD-guided hypothesis-triage resource. DMD is an illustrative stress test and evidence-prior layer; the current HepG2 benchmark is not presented as a DMD or muscle virtual cell.

Status terms: **resolved** means the current evidence and text can close the point; **partial** means the response improves the resource but a specified evidential gap remains; **blocked by new evidence** means existing files cannot validly close the point.

| ID | Reviewer concern | Revision action and evidence | Status | Residual requirement |
|---|---|---|---|---|
| 1 | Disease identity and HepG2 benchmark are mismatched | Adopted the audit-resource route; changed the title and abstract; made HepG2 internal generalization the explicit scope; relegated DMD to an illustrative direction-prior/triage use case; added a five-level claim ladder. | Resolved for framing | Disease-route claims remain prohibited until muscle/DMD perturbation data exist. |
| 2 | Fit to regular NAR Data Resources and Analyses is weak | Retained the regular route, added a task- and audit-axis resource comparison, explicit user tasks, versioned audit tables, failed results, model cards and reproducibility levels. | Partial | A public synchronized snapshot, DOI and a substantive independent disease-context dataset remain required for a strong submission. |
| 3 | A single DMD “consensus” is not statistically defensible | Renamed it the **integrated DMD direction prior**; released source-specific effects, source concordance, leave-one-source/source-only sensitivity and exploratory evidence modules. A formal random-effects meta-analysis was not fabricated because the harmonized source table lacks independent source-level standard errors and contains nested contexts. | Partial | Reconstruct independent study-level estimates and standard errors; then freeze a formal meta-analysis and an external holdout. |
| 4 | Five repeated splits are not independent replications | Audited split membership: 2,160 unique genes, 540 test genes/split, mean pairwise overlap 135.9, mean Jaccard 0.144 and 787 genes tested more than once. All text now calls this retrospective split stability. | Partial | A genuine immutable lockbox and, where biology allows, grouped family/pathway splits must be created prospectively from new or untouched data. |
| 5 | No measurement noise ceiling or biological-replicate reliability | Added target-level cell split-half, batch split and guide-pair/construct audits. Cell/batch median cosine was 0.331/0.333; guide-pair median cosine was 0.006 across 133 auditable targets. Added target reliability table and made target, not cell, the inferential unit. | Partial | Guide-resolved or independent biological-replicate data are needed; current reliability weights are descriptive and were not used for training. |
| 6 | Raw direction must be primary; RMSE/residual direction can be dominated by common effects | Reordered endpoints: raw cosine/sign agreement are primary; RMSE is secondary; common-effect-adjusted residual cosine is diagnostic only. Claim ladder prevents residual-only support from being described as biological direction prediction. | Resolved for reporting | New prospective models must use the frozen endpoint hierarchy. |
| 7 | DepMap may be an outcome proxy or leak; HepG2 row and temporal provenance are problematic | Added a feature-level leakage ledger. Rebuilt DepMap 26Q1 aggregates after excluding HepG2/ACH-000739 and reran ridge v1. Relative RMSE gains remained 8.3–9.6% vs zero, but one residual-cosine comparison worsened after Holm correction. The source file, release, bytes and SHA-256 are recorded. | Partial | The retained 26Q1 release post-dates the benchmark work; an older pre-benchmark snapshot was not available. Therefore “leak-free” and temporal-validation claims are prohibited. |
| 8 | Model comparison lacks model cards, fairness and coverage accounting | Added task-specific cards for zero/train-mean controls, ridge v1/v2, linear-PCA, GEARS, scGPT, STATE and the seen-delta positive control; recorded versions, seeds, native/common coverage, runtime, checkpoint hashes and interpretation boundaries. The deterministic local resource was additionally rebuilt in isolation with 31/31 and 11/11 checks passing. | Partial | Positive-control reproduction on each model's original published dataset and portable independent GPU/full-model reruns remain outstanding. |
| 9 | Candidate A/B/C tiers look like uncalibrated probabilities | Replaced the public-facing interpretation with a two-axis table: **evidence class** and **action priority**. Added source-direction agreement, signature and rule stability, model agreement, dependency, fusion risk, expression context and an immutable disclaimer that these are not effect sizes, P values, treatment probabilities or validation. Historical tier fields remain only for release reproducibility. | Resolved for presentation | Biological calibration still requires prospective muscle-relevant experiments. |
| 10 | “NMD” scope is broader than the evidence | Reframed the 64-gene layer as an NMD coverage-gap map; reported that 56/64 sentinel genes are unsupported and no DMD/BMD sentinel gene is directly perturbed. The scientific title now leads with “Audit,” not a broad disease virtual-cell claim. | Resolved for framing | Broader NMD claims require new directly perturbed disease-relevant datasets. |
| 11 | GTEx, muscle atlas and fusion screen are not external perturbation validation | Reclassified GTEx and scRNA as expression/applicability context and GSE293514 as a fusion-risk filter. Added denominators and target/donor/sample interpretation boundaries; no “validation of rescue” wording remains. | Partial | Independent perturbation transcriptomics or a prespecified functional screen in muscle/DMD context is still absent. |
| 12 | Reproducibility needs browser, data and regeneration levels, DOI and clean rebuild | Added a three-level reproducibility manifest, model/version hashes, release tables, SQLite/JSON/TSV access and explicit unresolved blockers. Completed an isolated deterministic local resource rebuild: 31/31 generator checks, 11/11 verifier checks, 63/63 schemas and 61/63 exact payloads. | Partial | Public synchronization, browser QA, DOI minting, license/contact completion and portable GPU/full-model regeneration are still required. |
| 13 | No task-based user evaluation | Added a frozen external-user protocol covering gene lookup, model comparison, candidate exclusion and unsupported-claim detection, with success, time and error metrics. | Blocked by external users | Run the protocol with independent disease biologists/computational users and publish de-identified task results. |
| 14 | Chronology and analytic forking paths are unclear | Added an analysis chronology. The candidate queue (10 July) predates the fusion audit (12 July); the 16 July decision matrix then incorporated fusion flags. Added a prospective freeze protocol and prohibited retrospective lockbox language. | Partial | Historical analyses were exploratory; prospective confirmation must use registered versions, splits, endpoints and stopping rules. |

## Decision after this revision

The resource is scientifically coherent as an **audit benchmark/resource**, not as a DMD therapeutic-target or muscle virtual-cell paper. The current package supports claim levels L1–L2 only: technical reliability for some aggregate response structure and internal HepG2 perturbation-specific generalization. It does not support cross-cell-type, disease-context or therapeutic claims (L3–L5).

The main remaining desk-reject risks are external rather than editorial: no genuine untouched lockbox, no independent muscle perturbation-response dataset, no formal study-level DMD meta-analysis, no external user evaluation, and no final public DOI-backed release.
