Purpose
Audit one caller-supplied quantitative claim by freezing a compact case with an embedded analysis contract, reproducing that contract under bounded local execution, running authorized sensitivity and targeted evidence work, independently verifying decisive findings, and returning a lean replayable audit package.
Status: Production | Updated: 2026-08-31
Steps
Build compact study case
Validate the target, map only its document and data dependency closure, preserve exact source-backed transformations and document conflicts, and freeze one study case containing primary_target and analysis_contract.
Completion criteria
The child returns one frozen study case with mandatory primary_target, compact analysis_contract, separate document_conflicts, bounded secondary benchmarks, and target-scoped data and code maps, or an explicit blocker.
A unique caller-frozen estimand with a contradictory supplementary document emits case_frozen_with_limits and remains eligible for reproduction.
Outcomes
case_frozenThe claim and every required execution link are fixed without a blocking gap.
case_frozen_with_limitsThe case can execute with visible nonblocking limitations.
case_not_assessableThe target cannot be fixed without an unsupported substantive choice.
operationally_blockedAuthority, privacy, safety, or inaccessible required files block execution.
Preflight and reproduce primary result
Preflight the real case bundle and shared execution capability, run the contract-defined model unchanged, test bounded working-directory or input-binding repairs when needed, and preserve complete scientific outputs in one reproduction package.
Completion criteria
The child receives the real case bundle and frozen analysis_contract, captures sample, missingness, coefficient, standard error, interval, statistic, degrees of freedom and p-value, and returns a self-contained reproduction package.
Execution provenance and scientific comparison remain separate: execution_mode is untouched or technically_repaired_original, while the emitted routine outcome follows comparison_status.
Verified network or DNS denial is not a prerequisite. Memory and storage are operational defaults, and unavailable exact enforcement does not emit not_executable when the analysis can run in an attempt-local working directory.
Outcomes
exact_reproductionThe untouched primary target and required benchmarks reproduce exactly under frozen comparison rules.
within_toleranceThe untouched primary target and required benchmarks satisfy the frozen tolerance rules.
partial_reproductionThe primary target or required benchmarks reproduce only in part and need an explicit continuation decision.
numerical_mismatchThe executable analysis remains numerically inconsistent after the one remediation pass.
not_executableA concrete runtime, dependency, path, permission, isolation, or resource-control blocker prevents execution.
not_assessableThe target cannot be assessed without an unsupported substantive choice.
Decide whether to continue after reproduction
Runs asPeer2Paper OrchestratorReview a partial reproduction or unresolved numerical mismatch and apply the caller's explicit authorization boundary before expensive downstream analysis.
Completion criteria
The supplied decision is continue or stop, names the reason and authorizing person or role, and applies to the current reproduction package.
continue_to_analysis is emitted only with explicit authorization to interpret the partial reproduction or numerical mismatch; otherwise stop_after_reproduction is emitted.
Outcomes
continue_to_analysisThe caller authorizes robustness and targeted research despite the recorded reproduction limitation.
stop_after_reproductionThe run ends with the case contract, document-consistency findings, reproduction package, and diagnostic evidence.
Test bounded sensitivity candidates
Run three result-blind design lanes, rank no more than five candidates, always materialize the conditional second-review artifact, and execute approved analyses through the ready reproduction profile.
Completion criteria
No more than five admissible candidates produce a common-schema sensitivity result, with untested dimensions and execution limits visible.
Outcomes
robustness_readyRequired sensitivity checks completed without residual gaps.
robustness_ready_with_limitsSensitivity checks completed with visible nonblocking gaps.
analysis_not_assessableA blocking scientific or executable gap prevents a defensible sensitivity conclusion.
operationally_blockedAuthority or unsafe execution prevents required work.
sensitivity_blockedSensitivity preparation found a scientific, source/data, or execution-capability blocker.
Answer targeted evidence questions
Run beside sensitivity work after reproduction succeeds or is explicitly continued. Answer no more than three questions raised by the claim, reproduction, or supplied-document conflicts, with no more than three deep-source checkpoints.
Completion criteria
The child returns exact passages, comparability judgments, verification state, and manuscript consequences for no more than three targeted questions and three deep-mapped sources.
Outcomes
evidence_readyRequired targeted evidence is available and supported without residual gaps.
evidence_ready_with_limitsThe targeted search completed with visible nonblocking gaps.
evidence_not_assessableA required evidence question cannot be answered from authorized accessible sources.
operationally_blockedAuthority or access policy prevents required research.
Verify decisive findings
Wait for sensitivity and targeted research, then independently rerun the frozen analysis contract from the real case bundle in a clean workspace with equivalent isolation and resource controls, reopen verdict sources, and produce adjudicated findings.
Completion criteria
The child uses the same frozen analysis contract and source and data hashes, captures complete required numerical comparisons in a clean attempt-local workspace, and returns allowed wording plus concrete manuscript consequences.
Different runner or mount identities do not fail verification when the scientific contract, source and data hashes, dependencies, source immutability, timeout, and recorded environment are equivalent for the claim.
External release stays disabled unless separately approved.
Outcomes
audit_readyAll required findings are independently verified and ready for reporting.
verification_incompleteThe audit can report verified work with explicit unresolved limitations.
operationally_blockedRequired verification cannot proceed safely or with authority.
Publish JSON-first audit package
Build one canonical JSON audit from the verified findings and complete manuscript recommendations, render concise HTML from that JSON, and package replay evidence and scoped release checks.
Completion criteria
One canonical JSON audit, concise HTML report, replay package, and package manifest agree on all numbers, findings, recommendations, and limitations.
Outcomes
audit_completedThe audit and replay package pass every required scoped check.
audit_completed_with_limitsThe audit is releasable with explicit nonblocking limitations.
operationally_blockedA required completeness, replay, authority, or scoped release check blocks packaging.
Sub-routines
Sub-routines
Assemble Final Audit
Purpose
Publish one canonical JSON audit, a concise researcher-facing HTML report, and a replay package with scoped release checks. PDF is outside the required hackathon output.
Steps
Build canonical audit and replay package
Runs asAudit Report BuilderBuild one canonical JSON audit and replay directory from the frozen case, reproduction, sensitivity, targeted evidence, adjudicated findings, and complete manuscript recommendations. Keep raw logs, hashes, prompts, and provenance in the machine package.
Completion criteria
The canonical JSON contains the verdict, reported and reproduced values, numerical differences, sensitivity result, verified conflicts, concrete manuscript recommendations, limitations, and content-addressed references to technical evidence.
The JSON validates against the pinned schema and uses one value for every displayed number and statement.
The replay package contains the shared execution profile, exact commands, 600-second timeout, environments, inputs, outputs, checks, and hashes needed to rerun every scientific result included in the verdict.
Raw logs, IDs, prompts, and detailed provenance remain in the machine package rather than the researcher report.
Every concrete manuscript recommendation from the verified recommendations resource appears in the canonical audit with its affected section or statement, exact change, evidence, urgency, and required or optional status.
The package is internal-only unless audit-request contains separate explicit external-release approval.
Render concise researcher report
Runs asAudit Report BuilderRender a two-to-three-page HTML report from the canonical JSON using the pinned template.
Completion criteria
The HTML presents the verdict, reported versus reproduced result, sensitivity summary, verified contradictions or ambiguities, exact manuscript changes, and remaining limitations.
Every number and allowed statement is read from the canonical JSON, and detailed provenance is linked rather than expanded into large report tables.
The HTML is accessible, readable, and contains no unsupported scientific finding.
Check and package audit
Runs asRelease Quality ReviewerCheck cross-format identity, replay evidence, completeness, and internal release boundaries, then write the final manifest. Run licence and retention checks only when the audit request authorizes public redistribution.
Completion criteria
Every displayed number and verdict statement matches the canonical JSON, and each scientific result has replayable evidence.
Release checks cover completeness, JSON-to-HTML consistency, provenance, replay, accessibility, and internal access boundaries.
External release is disabled unless audit-request contains separate explicit approval. Full licence and retention adjudication runs only with that approval; otherwise the manifest records internal-only scope.
The manifest reports pass, limited, or blocking for each required check and never lists PDF as a required core output.
Emit audit_completed only with no blocking or limited gate; emit audit_completed_with_limits only with no blocking gate and at least one visible limitation.
Outcomes
audit_completedThe audit and replay package pass every required scoped check.
audit_completed_with_limitsThe audit is releasable with explicit nonblocking limitations.
operationally_blockedA required completeness, replay, authority, or scoped release check blocks packaging.
Sub-routines
Build And Freeze Study Case
Purpose
Freeze one caller-supplied quantitative claim into a compact study case containing the mandatory primary target, exact source-backed analysis contract, target dependency closure, bounded secondary benchmarks, and separately preserved document conflicts, without executing a scientific model during intake.
Steps
Validate target claim and map relevant documents
Runs asClaim MapperValidate the caller-supplied target claim and map only the primary paper, supplement, preregistration, and passages needed to identify its reported result and wording.
Completion criteria
One target claim is fixed with exact text and location, population, exposure or intervention, comparator, outcome, timepoint, observation unit, estimand, directional orientation, conclusion-change rules, and known ambiguities.
The primary reported estimate, interval, p-value, sample size, and exact table, figure, or code target are recorded when present.
Only the primary paper, supplement, preregistration, and claim-relevant passages are mapped; unreadable or missing required material is explicit.
Record only values reported in authorized documents and their exact locations. Do not execute supplied analysis code, fit a model, or compute a diagnostic estimate, interval, statistic, p-value, or model sample size.
Map target data and variables
Runs asVariable MapperMap only columns in the primary target dependency closure, including sample filters and variables required to derive those columns. Do not profile unrelated columns or search for participant keys unless target execution requires participant identity.
Completion criteria
Every primary-target element and bounded secondary benchmark has a direct data-field, derivation, sample-filter, or documented missing-link mapping within the target dependency closure.
The map records units, condition codes, derivations, target-relevant missingness, observation unit, and unresolved high-impact ambiguity without profiling unrelated columns.
Single-column uniqueness scans and combination-key searches are prohibited unless participant identity is required by the frozen target; any permitted identity check is bounded to the minimum declared columns.
The checkpoint reports counts and stable lists of mapped columns, inspected columns, and intentionally excluded columns, with inspected columns limited to the target dependency closure.
Map fields, filters, derivations, and target-relevant missingness without fitting a statistical model or calculating the target result.
Map target code and dependencies
Runs asReproduction EngineerIdentify the untouched entrypoint, target-producing code, exact formula, sample filter, transformations, factor levels, coefficient and orientation, dependencies, paths, inputs, outputs, seeds, and no more than three secondary benchmarks.
Completion criteria
The primary command and one unique target-producing code path are identified, or the exact missing or ambiguous entrypoint is recorded.
The code map records source file identities and hashes plus exact source lines and literal expressions for the formula, population filter, transformations, factor levels, reference group, coefficient, orientation, and missing-data declaration.
Expressions present in source are copied exactly into the code map and are not reconstructed from prose or line ranges.
A dependency manifest records languages, lockfiles, imports, packages, system requirements, paths, mounts, permissions, network needs, seeds, and missing files.
The mandatory primary target is separate from no more than three secondary benchmarks.
This step is static code and dependency inspection. Parse-only validation is permitted, but evaluating scientific source, fitting a model, or generating result values is prohibited.
Verify supplied-document consistency
Runs asSource Evidence VerifierCompare the primary paper, supplement, preregistration, and completed code map only where they define the target population, variables, condition codes, analysis, or reported result. Record exact conflicts and manuscript consequences; do not remap unrelated content.
Completion criteria
Every retained conflict links exact passages or code locations, identifies the affected claim or execution element, and is verified or marked unresolved.
The check is limited to the target paper, supplement, preregistration, and completed code map.
Compare supplied document passages and static code locations only. Do not execute the analysis or use a newly calculated model result as document-consistency evidence.
Bind and freeze analysis
Runs asClaim MapperJoin the claim, target-scoped data map, exact source-backed code map, and verified document findings into one frozen study case with a compact analysis_contract section.
Completion criteria
primary_target is mandatory on every frozen case and remains separate from the bounded secondary_benchmarks array.
analysis_contract contains code_hash, data_hash, formula, population_filter, exact transformations, factor levels and reference group, coefficient, orientation, missing-data declaration, required columns, entrypoint, and dependency identity copied from the code and data maps.
document_conflicts remains a separate collection linked to exact supplied passages or code locations and does not overwrite the primary target or analysis contract.
When the caller-frozen estimand maps to one unique supplied code path and the audit request, scientific policy, primary evidence, and code agree, a contradictory supplementary document emits case_frozen_with_limits without stopping reproduction.
Emit case_frozen only when the unique analysis contract has no nonblocking conflict; emit case_not_assessable only when no unique target exists or choosing a code path would change the caller-frozen estimand; emit operationally_blocked only for authority, safety, privacy, or inaccessible required files.
New intake artefacts must not contain source_consistency_probe or any model-derived diagnostic result. Scientific execution evidence begins only in Reproduce Original Result through the official runner.
Outcomes
case_frozenThe caller-frozen estimand maps to one schema-valid executable target without a blocking gap or source conflict.
case_frozen_with_limitsThe executable target is schema-valid and may run, with a visible nonblocking supplied-document conflict or other declared limitation.
case_not_assessableNo unique executable target exists, or choosing a supplied code path would change the caller-frozen estimand.
operationally_blockedAuthority, privacy, safety, or inaccessible required files block execution.
Sub-routines
Reproduce Original Result
Purpose
Execute the exact analysis_contract embedded in the frozen study case through the shared runner, using an attempt-local working copy and bounded technical repairs without changing scientific expressions, then report execution_mode separately from the component-level scientific comparison.
Status: Production | Updated: 2026-08-31
Steps
Freeze reproduction contract and inputs
Runs asReproduction EngineerFreeze the mandatory primary target and embedded analysis_contract, a separate array of no more than three secondary benchmarks, comparison rules, untouched command, data and source hashes, and dependencies.
Completion criteria
The frozen study case contains primary_target and analysis_contract with code and data hashes, formula, population filter, exact transformations, factor levels, coefficient, orientation, missing-data declaration, entrypoint, and dependencies.
The supplied case bundle contains the source and data identities referenced by the analysis contract.
The checkpoint fixes one primary target and a separate array of no more than three secondary benchmarks with reported estimate, interval, statistic, p-value, sample size, and comparison rules.
Preflight shared execution
Runs asReproduction EngineerUse the project execution runner to validate the real runtime, dependencies, input hashes, attempt-local working copy, source immutability, host network policy, operational controls, cooperative CPU settings, and a trivial language command before the study command.
Completion criteria
Execution readiness records the case bundle, source and data hashes, runtime, dependencies, attempt-local working directory, registered input paths, permissions, host network policy, timeout, and available controls with exact evidence.
When execution may proceed, execution-readiness and execution-profile set status to ready and execution_allowed to true. They preserve nonblocking control gaps in limitations rather than inventing another readiness status.
maximum_cpu_cores is best-effort unless execution-policy explicitly requires hard enforcement for safety, authorization, or infrastructure protection. Set available cooperative thread controls to the requested value and record their names and values.
When hard CPU enforcement is unavailable but not explicitly required, include a limitation with type cpu_limit_best_effort, requested_cores from execution-policy, and hard_enforcement_available false; keep status ready and execution_allowed true.
Hard CPU enforcement may block only when execution-policy explicitly marks it as required for safety, authorization, or infrastructure protection and no safe authorized execution path exists.
A missing runtime or dependency, unreadable required file, source or data hash mismatch, missing permission to create a protected attempt-local working copy, unenforceable required network isolation, or another explicit safety or authority requirement may block execution.
Execution profile is always written as an attempt envelope. Operational controls and limitations are preserved for downstream consumers, which trust execution_allowed rather than reinterpreting a best-effort CPU field.
Every scientific command has a 600-second hard timeout. No advisory routine budget is inferred.
Run untouched primary target
Runs asReproduction EngineerWhen a clean supplied entrypoint exists, copy the required inputs into a separate attempt-local working directory and execute the unchanged entrypoint once. Do not generate or use a wrapper in this step; otherwise write a skipped checkpoint for bounded repair.
Completion criteria
The checkpoint is written for every run with status completed, failed, blocked, skipped_for_repair, or not_assessable and records the executable-contract hash.
This step never creates or runs a wrapper. An unresolved working directory, input binding, output capture, or section-selection need writes skipped_for_repair and proceeds to the frozen repair candidates.
Before execution, the runner verifies source and data hashes, parses any R command file without evaluating it, and confirms that the working directory is inside the attempt directory.
A completed attempt proves that the source hashes, population filter, formula, transformations, treatment arms, factor levels, reference group, contrast, missing-data behavior, and coefficient orientation match the executable contract.
A completed or failed attempt records the command, 600-second timeout, working directory, runtime, dependencies, inputs, outputs, logs, environment details, and hashes. Verified inputs are hashed before and after and must remain unchanged.
Model diagnostics include initial eligible rows, model-frame rows, nobs(), excluded-row count, na.action class and excluded-row evidence, coefficient and standard error, test statistic and degrees of freedom, p-value, interval endpoints, interval method, confidence level, and coefficient orientation when scientific source execution reaches them.
A best-effort CPU limitation never changes a ready execution profile to blocked and never prevents the untouched attempt or an eligible bounded repair.
Compare primary result
Runs asReproduction EngineerCompare the untouched primary attempt with the reported result and executable contract. Score estimate, interval, statistic, p-value, and sample independently, and check internal numerical coherence without treating one matching component as evidence for another.
Completion criteria
The checkpoint records reported and reproduced values, absolute and relative differences, tolerance basis, displayed precision, and a separate pass, fail, not_reported, or not_comparable status for estimate, each interval endpoint, statistic, p-value, and sample.
The checkpoint records initial eligible rows, model-frame rows, nobs(), excluded-row count, na.action class and evidence, coefficient standard error, degrees of freedom, interval method, confidence level, and coefficient orientation.
For reported estimate 0.028 and interval [0.006, 0.041], the checkpoint records midpoint 0.0235 and coefficient-to-midpoint discrepancy 0.0045, then determines the interval method from evidence instead of inferring interval validity from the coefficient.
exact_reproduction or within_tolerance is eligible only when every comparable required numerical component passes its frozen rule and the executed model specification matches the contract.
Missing execution evidence is classified from the preceding checkpoint and never treated as a numerical result.
Record execution_mode separately as untouched or technically_repaired_original. Record comparison_status separately as exact_reproduction, within_tolerance, partial_reproduction, numerical_mismatch, or not_executable.
A technically repaired execution is scored under the same component rules as an untouched execution; repair provenance cannot replace or determine the scientific comparison status.
Diagnose failure and freeze repairs
Runs asReproduction EngineerAfter an unchanged execution failure, a skipped clean entrypoint, or a numerical mismatch, freeze no more than three ordered technical repairs from the explicit allowlist. Record denylisted proposals as rejected and never execute them.
Completion criteria
The checkpoint always exists and has status candidates_frozen or not_required.
Allowed repairs are limited to binding the attempt-local working directory, wrapping unchanged source for launch or output capture, selecting an author-declared replication section, binding registered input paths, and capturing runtime dependencies.
Filters, sample rules, formulas, transformations, factor levels, treatment arms, contrasts, missing-data behavior, intervals, and inferential methods cannot change.
Each wrapper repair requires ASCII-safe string encoding with encodeString(value, quote = '"'), forbids dQuote and sQuote, persists the wrapper in the attempt directory, and requires parse-only validation before scientific execution.
When the source declares a replication start, contains an empty setwd call, reads a registered local input, and may fail later during optional output work, rank as candidate 1 a wrapper that runs from the writable attempt directory, binds or copies the registered input there, starts at the author-declared replication section, neutralizes only the empty setwd call, and captures the target model before optional plotting or report-generation failures.
Test repair candidate 1
Runs asReproduction EngineerTest the first frozen repair in an isolated execution copy when assigned and needed; otherwise write a skipped or not_required checkpoint.
Completion criteria
The first candidate always writes a checkpoint with attempted, parse_failed, completed, failed, skipped, or not_required status and the exact reason.
A generated R wrapper is stored in the attempt directory, uses ASCII-safe encoded string literals, contains no dQuote, sQuote, or typographic quotes, and passes parse-only validation before supplied scientific source is evaluated.
An attempted repair records the wrapper content and hash, operational change, unchanged scientific expressions, source and data hashes before and after, the 600-second timeout, execution log, and outcome.
A successful working-directory, launch, section-selection, output-capture, input-binding, or dependency-capture repair is classified technically_repaired_original.
A parse failure or runtime failure consumes only this bounded candidate, preserves its evidence, and allows the next frozen candidate to run. A repair that changes the analysis contract cannot support patched_reproduction.
A completed repaired run captures initial eligible rows, model rows, nobs(), excluded rows and na.action evidence, coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.
For the declared replication-section repair, the wrapper runs through the official runner from the attempt directory, binds or copies the registered CSV into that directory, begins at the author-declared replication marker, neutralizes only setwd("") in that section, and captures the target model and diagnostics before optional plotting.
The checkpoint records execution_mode technically_repaired_original plus component-level comparison evidence; it does not use repaired execution as a scientific verdict.
Missing hard CPU affinity cannot skip this repair when execution_allowed is true.
Test repair candidate 2
Runs asReproduction EngineerTest the second frozen repair only when the first did not yield a valid reproduced primary result; otherwise write a skipped or not_required checkpoint.
Completion criteria
The second candidate always writes a checkpoint with attempted, parse_failed, completed, failed, skipped, or not_required status and the exact reason.
A generated R wrapper is stored in the attempt directory, uses ASCII-safe encoded string literals, contains no dQuote, sQuote, or typographic quotes, and passes parse-only validation before supplied scientific source is evaluated.
An attempted repair records the wrapper content and hash, operational change, unchanged scientific expressions, source and data hashes before and after, the 600-second timeout, execution log, and outcome.
A successful working-directory, launch, section-selection, output-capture, input-binding, or dependency-capture repair is classified technically_repaired_original.
A parse failure or runtime failure consumes only this bounded candidate, preserves its evidence, and allows the next frozen candidate to run. A repair that changes the analysis contract cannot support patched_reproduction.
A completed repaired run captures initial eligible rows, model rows, nobs(), excluded rows and na.action evidence, coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.
Missing hard CPU affinity cannot skip this repair when execution_allowed is true.
Test repair candidate 3
Runs asReproduction EngineerTest the third frozen repair only when earlier candidates did not yield a valid reproduced primary result; otherwise write a skipped or not_required checkpoint.
Completion criteria
The third candidate always writes a checkpoint with attempted, parse_failed, completed, failed, skipped, or not_required status and the exact reason.
A generated R wrapper is stored in the attempt directory, uses ASCII-safe encoded string literals, contains no dQuote, sQuote, or typographic quotes, and passes parse-only validation before supplied scientific source is evaluated.
An attempted repair records the wrapper content and hash, operational change, unchanged scientific expressions, source and data hashes before and after, the 600-second timeout, execution log, and outcome.
A successful working-directory, launch, section-selection, output-capture, input-binding, or dependency-capture repair is classified technically_repaired_original.
A parse failure or runtime failure consumes only this bounded candidate, preserves its evidence, and allows the next frozen candidate to run. A repair that changes the analysis contract cannot support patched_reproduction.
A completed repaired run captures initial eligible rows, model rows, nobs(), excluded rows and na.action evidence, coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.
Missing hard CPU affinity cannot skip this repair when execution_allowed is true.
Run secondary target 1
Runs asReproduction EngineerAfter the primary result succeeds unchanged or through a valid repair, run the first required secondary benchmark if assigned. Otherwise write not_assigned, blocked, or skipped.
Completion criteria
The checkpoint always exists and records one benchmark comparison or the exact reason it was not run.
The scientific command runs when the primary target has valid comparable execution evidence under either execution_mode, including a partial component comparison; it does not run after not_executable or not_assessable.
Run secondary target 2
Runs asReproduction EngineerAfter the primary result succeeds unchanged or through a valid repair, run the second required secondary benchmark if assigned. Otherwise write not_assigned, blocked, or skipped.
Completion criteria
The checkpoint always exists and records one benchmark comparison or the exact reason it was not run.
The scientific command runs when the primary target has valid comparable execution evidence under either execution_mode, including a partial component comparison; it does not run after not_executable or not_assessable.
Run secondary target 3
Runs asReproduction EngineerAfter the primary result succeeds unchanged or through a valid repair, run the third required secondary benchmark if assigned. Otherwise write not_assigned, blocked, or skipped.
Completion criteria
The checkpoint always exists and records one benchmark comparison or the exact reason it was not run.
The scientific command runs when the primary target has valid comparable execution evidence under either execution_mode, including a partial component comparison; it does not run after not_executable or not_assessable.
Package reproduction result
Runs asReproduction EngineerAssemble the durable readiness, primary, repair, and secondary checkpoints into a compact comparison package and replay directory without rerunning analysis.
Completion criteria
The reproduction package records or content-addresses the frozen analysis contract, case-bundle identities, execution profile, commands, 600-second timeout, environment, inputs, outputs, logs, hashes, repairs, numerical diagnostics, component comparisons, and replay instructions.
Every completed checkpoint remains durable, and a later failure cannot erase earlier successful evidence.
A successful permitted wrapper records execution_mode technically_repaired_original. The routine outcome follows comparison_status and never follows the repair label.
comparison_status exact_reproduction or within_tolerance requires every comparable required estimate, interval endpoint, statistic, p-value, and sample component to pass its frozen rule, regardless of execution_mode.
A missing hard CPU limit is recorded as cpu_limit_best_effort and cannot produce not_executable when execution_allowed is true. Required network isolation, source protection, and explicit safety or authority controls remain gates.
External release remains disabled unless separately approved.
comparison_status is partial_reproduction when valid execution yields a mixed component result, including a matching estimate, statistic, and p-value with a failing reported interval. A paper-omitted model sample is recorded as not_reported and never silently treated as a match.
Outcomes
exact_reproductionEvery comparable required component reproduces exactly under frozen rules, using either execution_mode.
within_toleranceEvery comparable required component passes its frozen tolerance rule, using either execution_mode.
partial_reproductionValid execution produced a mixed component comparison, such as matching estimate, statistic, and p-value with a failing interval.
numerical_mismatchThe executable analysis remains materially inconsistent after the bounded remediation pass and does not qualify as a partial component match.
not_executableA missing runtime or dependency, unreadable required input, source or data hash mismatch, permission failure, required isolation failure, or inability to create a protected attempt-local working copy prevents every scientific attempt. Best-effort CPU enforcement alone cannot emit this outcome.
not_assessableThe target cannot be assessed without an unsupported substantive choice.
Sub-routines
Research Methods And Evidence
Purpose
Answer no more than three targeted questions raised by the frozen claim, reproduction result, or supplied-document conflicts, deep-map no more than three authorized sources through checkpointed Silk-aware workers, and package exact passages and manuscript consequences. Sensitivity results may trigger a bounded follow-up only during final verification.
Status: Production | Updated: 2026-08-29
Prerequisites
- Call
browser_silk_loadbefore opening each mapped literature site used by the active step. - Verify that each source is open access or otherwise authorized by the active audit before opening full text.
- Reuse the mapped funnels below; none of these read-only literature funnels requires authentication.
- Saved Browser Silk funnels available to the three bounded source slots:
pmc.ncbi.nlm.nih.gov/open_and_extract_articlerequirespmcidand returns visible full text, headings, tables, figures, supplements, and references.arxiv.org/open_and_extract_full_textrequiresarxiv_idand returns article text and structure.journals.sagepub.com/open_and_extract_articlerequiresdoiand returns visible article text and structure.www.cambridge.org/open_and_extract_articlerequiresarticle_slugand returns visible article text and structure.rameliaz.github.io/open_pdfrequiresdocument_path; the local PDF parser supplies page and text locations when browser text is unavailable.www.bmj.com/open_and_extract_articlerequiresarticle_pathand returns visible article text and structure.www.nber.org/open_working_paper_pdfrequiresdocument_path; the local PDF parser supplies page and text locations when browser text is unavailable.osf.io/project inventoryverifies project identity and authorized download records; registered local copies supply full text when available.
Steps
Freeze targeted research questions
Runs asResearcherFreeze no more than three evidence questions that arise from the frozen claim, reproduction comparison, or verified supplied-document conflicts. Do not start a broad literature review or require a completed sensitivity map.
Completion criteria
The brief contains no more than three questions, and each question names the finding it can confirm, contradict, or qualify.
Each question has source classes, search concepts, inclusion and comparison rules, authorization limits, and a stopping rule.
The initial brief does not depend on sensitivity results; any later sensitivity-triggered source check is deferred to final verification.
Search, screen, and assign source slots
Runs asLiterature SearcherUse toone-scholar and authorized sources to answer the frozen questions, metadata-screen broadly, deduplicate versions, and assign no more than three deep-source slots. Use Toone's persistent browser and complete saved Silk flows for recurring interactive sources.
Completion criteria
Every query, source, timestamp, version, access status, screening decision, and deduplication key is recorded.
No more than three deep-source slots are assigned; extra candidates remain metadata-only.
Each slot names its authorized access route, exact evidence target, and mapped Silk funnel when applicable.
Map selected source 1
Runs asSource ExtractorProcess only source slot 1. Load the mapped domain before navigation, replay the complete saved Silk funnel, and write one independent checkpoint; if no source is assigned, write not_assigned without browsing.
Completion criteria
Source slot 1 always produces one checkpoint with status mapped, limited, inaccessible, failed, or not_assigned.
This step is one independent Source Extractor execution in the three-way fan-out. It processes only its assigned slot and never takes another slot's source.
Mapped evidence includes source identity, version, exact passage and location, context, comparability, supported or contradicted statement, manuscript consequence, access status, and replay provenance.
Read maximum_seconds_per_full_text from execution-policy. Complete any started Silk funnel, then converge immediately to the checkpoint without opening another source or beginning optional extraction work when the limit is reached.
For selector_not_found, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. If it still fails, preserve the last failure evidence in the checkpoint; never switch to manual browsing.
Map selected source 2
Runs asSource ExtractorProcess only source slot 2. Load the mapped domain before navigation, replay the complete saved Silk funnel, and write one independent checkpoint; if no source is assigned, write not_assigned without browsing.
Completion criteria
Source slot 2 always produces one checkpoint with status mapped, limited, inaccessible, failed, or not_assigned.
This step is one independent Source Extractor execution in the three-way fan-out. It processes only its assigned slot and never takes another slot's source.
Mapped evidence includes source identity, version, exact passage and location, context, comparability, supported or contradicted statement, manuscript consequence, access status, and replay provenance.
Read maximum_seconds_per_full_text from execution-policy. Complete any started Silk funnel, then converge immediately to the checkpoint without opening another source or beginning optional extraction work when the limit is reached.
For selector_not_found, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. If it still fails, preserve the last failure evidence in the checkpoint; never switch to manual browsing.
Map selected source 3
Runs asSource ExtractorProcess only source slot 3. Load the mapped domain before navigation, replay the complete saved Silk funnel, and write one independent checkpoint; if no source is assigned, write not_assigned without browsing.
Completion criteria
Source slot 3 always produces one checkpoint with status mapped, limited, inaccessible, failed, or not_assigned.
This step is one independent Source Extractor execution in the three-way fan-out. It processes only its assigned slot and never takes another slot's source.
Mapped evidence includes source identity, version, exact passage and location, context, comparability, supported or contradicted statement, manuscript consequence, access status, and replay provenance.
Read maximum_seconds_per_full_text from execution-policy. Complete any started Silk funnel, then converge immediately to the checkpoint without opening another source or beginning optional extraction work when the limit is reached.
For selector_not_found, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. If it still fails, preserve the last failure evidence in the checkpoint; never switch to manual browsing.
Assemble targeted evidence
Runs asResearcherJoin the three source checkpoints into a targeted evidence package, preserving metadata-only candidates, access limits, exact passages, comparability judgments, and concrete manuscript consequences.
Completion criteria
Every retained evidence item has an exact passage and location, the statement it supports or contradicts, a comparability assessment, verification state, and proposed manuscript consequence.
The package joins all three independent source checkpoints, including limited, inaccessible, failed, and not_assigned slots, without serializing or rerunning a sibling slot.
The package answers no more than three questions, contains no more than three deeply processed sources, distinguishes inaccessible and unassigned slots, and exposes remaining gaps.
Evidence used in a verdict is marked for independent reopening during verification; a second complete extraction review is not run here.
Write an immutable run-scoped evidence snapshot alongside the selected per-run evidence package; no prior-run replacement read is required.
Outcomes
evidence_readyRequired targeted evidence is available and supported without residual gaps.
evidence_ready_with_limitsThe targeted search completed with visible nonblocking gaps.
evidence_not_assessableA required evidence question cannot be answered from authorized accessible sources.
operationally_blockedAuthority or access policy prevents required research.
Sub-routines
Stress Test Analysis
Purpose
Test three bounded families of scientifically defensible alternatives that preserve the frozen estimand, approve no more than five result-blind candidate cards, execute them through the shared local runner with bounded commands, and report one interpretable sensitivity result.
Status: Production | Updated: 2026-08-31
Steps
Freeze sensitivity contract
Runs asRobustness LeadPreflight the supplied case bundle and shared local execution capability, confirm that reproduction may continue, and freeze the estimand, conclusion rules, admissibility rules, candidate cap, and stopping rules before alternative results exist.
Completion criteria
The supplied case bundle resolves the frozen source and data hashes, and the runner can create an attempt-local working copy with the required runtime, dependencies, registered inputs, source protection, and 600-second timeout.
Verified network or DNS denial is not required. Memory and storage are operational defaults and unavailable exact enforcement does not block sensitivity work.
The estimand fixes population, treatment or exposure, comparator, outcome, time, effect, observation unit, direction, practical threshold, and conclusion categories before candidate execution.
Emit sensitivity_ready when the baseline is reproduced or explicitly authorized and the local execution preflight passes; emit sensitivity_blocked only for a scientific, source/data, runtime, dependency, permission, or source-protection blocker.
The candidate pool is capped at five and each candidate changes one primary analytical choice.
Outcomes
sensitivity_readyThe frozen contract and shared execution capability are ready for bounded sensitivity design and execution.
sensitivity_blockedA scientific, source/data, or execution-capability blocker prevents sensitivity work.
Design data and measurement candidates
Runs asData Choices AnalystPropose no more than two high-value candidates covering sample inclusion, missingness, outcome or exposure definitions, and transformations while preserving the frozen estimand.
Completion criteria
The lane returns at most two candidate cards, each with one primary change, scientific rationale, supplied support, unchanged estimand elements, expected sample consequence, classification, diagnostics, and compatibility constraints.
Candidate cards contain no alternative-analysis outcomes.
Design model specification candidates
Runs asModel And Covariate AnalystPropose no more than two high-value candidates covering covariates, functional form, weighting, or interactions while preserving the frozen estimand.
Completion criteria
The lane returns at most two candidate cards, each with one primary change, scientific rationale, supplied support, unchanged estimand elements, expected sample consequence, classification, diagnostics, and compatibility constraints.
Post-treatment, collider, unavailable, and meaning-changing specifications are excluded before results exist.
Design inference candidates
Runs asInference AnalystPropose no more than two high-value candidates covering standard errors, clustering, multiplicity, or confidence intervals under the frozen design.
Completion criteria
The lane returns at most two candidate cards, each with one primary change, scientific rationale, supplied support, unchanged estimand elements, classification, diagnostics, and compatibility constraints.
Inference choices identify their design basis and contain no alternative-analysis outcomes.
Rank and freeze sensitivity candidates
Runs asRobustness LeadJoin the three lanes, remove duplicates and incompatible combinations, and freeze a ranked pool of three to five candidate cards before results are available.
Completion criteria
The frozen pool contains three to five candidates when that many defensible options exist, with stable IDs, one primary change each, rationale, support, unchanged estimand elements, expected sample consequence, type, cost, and no result fields.
Excluded and unselected options retain concise reasons, and the selected pool covers more than one decision lane when scientifically possible.
Run primary blind admissibility review
Runs asBlind Method Reviewer AlphaJudge each candidate from results-blind cards under the frozen admissibility rules and identify whether a second review is required for a dispute or high-impact judgment.
Completion criteria
Every candidate receives accept, supplementary, reject, or unresolved with criterion-level reasons and no result exposure.
Emit second_review_required only for a disagreement risk or high-impact admissibility judgment; otherwise emit primary_review_sufficient.
Outcomes
primary_review_sufficientEvery candidate can be resolved from the primary blind review.
second_review_requiredAt least one disputed or high-impact judgment needs a second blind review.
Record conditional second blind review
Runs asBlind Method Reviewer BetaAlways write the second-review artifact. Review only candidates flagged as disputed or high-impact by the primary reviewer; when none are flagged, emit not_required without repeating the primary review.
Completion criteria
The artifact always exists.
Flagged candidates receive an independent criterion-level blind judgment with no result exposure.
Every unflagged candidate is marked not_required with the primary-review reference; no substantive second review is performed for it.
Resolve candidate admissibility
Runs asRobustness LeadApply the frozen disagreement rule once and produce the final set of executable sensitivity candidates without reconsidering admissibility later.
Completion criteria
No more than five accepted or supplementary candidates are executable, and every rejected or unresolved candidate retains its review evidence.
Each approved candidate preserves the estimand and has one primary analytical change.
Execute approved sensitivity analyses
Runs asReproduction EngineerUse the shared runner to execute each approved candidate as an isolated worker, compare outputs under one result schema, and preserve failed or untested candidates.
Completion criteria
Execution begins only after sensitivity_ready routing and after ranked candidates, blind-review-alpha, and the always-materialized blind-review-beta checkpoint have joined.
Each approved candidate runs against the frozen source and data hashes through the shared runner from a separate attempt-local working directory and has a command, 600-second timeout, inputs, outputs, logs, diagnostics, estimate, interval, p-value, sample size, status, and hashes.
Verified inputs remain unchanged. Host network policy and operational memory or storage defaults are recorded but are not scientific-validity gates.
Candidate execution is capped at five; failed, timed-out, and untested candidates remain visible without open-ended repair.
Build sensitivity result
Runs asRobustness LeadSummarize stability, fragility, coverage, and the smallest defensible conclusion change from approved candidate runs.
Completion criteria
The result states how many approved analyses preserve direction and statistical support, the estimate and interval range, any valid direction reversal, the smallest defensible conclusion-changing choice, and whether changes arise from estimate size, uncertainty, or sample composition.
Tested and untested dimensions, failed candidates, coverage, and limitations are explicit; significance counts alone are not used as the conclusion.
Write an immutable run-scoped sensitivity snapshot alongside the selected per-run result; no prior-run replacement read is required.
Outcomes
robustness_readyRequired sensitivity checks completed without residual gaps.
robustness_ready_with_limitsSensitivity checks completed with visible nonblocking gaps.
analysis_not_assessableA blocking scientific or executable gap prevents a defensible sensitivity conclusion.
operationally_blockedAuthority or unsafe execution prevents required work.
Sub-routines
Verify And Adjudicate Findings
Purpose
Independently rerun the primary and verdict-relevant sensitivity results from the same frozen analysis contract and source and data hashes in a separate local working directory, then verify decisive source evidence and adjudicate manuscript changes.
Status: Production | Updated: 2026-08-31
Steps
Freeze important verification targets
Runs asVerification LeadSelect only verdict-relevant reproduction, sensitivity, document, and source findings, freeze their evidence identities, and gate statistical verification against the executable contract and reproduction execution identities.
Completion criteria
Every selected finding has one precise statement, claim and estimand links, importance reason, allowed verification method, source or run identity, hashes, and required completion rule.
The packet freezes the primary_target and analysis_contract from frozen-study-case, including source and data hashes, exact expressions, coefficient, orientation, and comparison rules.
The supplied case bundle resolves the frozen source and data hashes, and the runner can create a separate attempt-local working directory with the required runtime, dependencies, source protection, and 600-second timeout.
Verified network or DNS denial is not required. Memory and storage are operational defaults, and unavailable exact enforcement is recorded without blocking a runnable verification command.
Previously decided candidate admissibility is preserved, and external release remains disabled unless the audit request contains separate approval.
Independently rerun statistical evidence
Runs asStatistical Evidence VerifierPerform an independent execution of the frozen analysis_contract and only verdict-relevant sensitivities from a separate local working directory using the same source and data hashes.
Completion criteria
Preflight proves access to the case bundle, frozen source and data hashes, required runtime and dependencies, registered inputs, a separate attempt-local working directory, source protection, and a 600-second timeout.
The verification run executes the exact frozen formula, filter, transformations, factor levels, treatment or exposure structure, coefficient, orientation, missing-data behavior, interval method, and inferential method.
The primary run records command, working directory, timeout, environment, inputs, outputs, logs and hashes; eligible rows, model rows, nobs(), excluded rows and na.action evidence; coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.
Estimate, interval endpoints, statistic, p-value, and sample are compared independently with the reported and reproduction values.
A different runner hash, workspace path, mount identity, or available operational control does not fail verification when the frozen scientific contract, source and data hashes, dependencies, source immutability, and recorded environment are equivalent for the claim.
Verified network or DNS denial and exact memory or storage enforcement are not prerequisites. A scientific-contract or source/data hash mismatch remains a precise blocker and cannot support audit_ready.
Independently reopen source evidence
Runs asSource Evidence VerifierReopen only supplied-document conflicts and external passages used in the verdict or recommendations, using authorized local sources, toone-scholar, or the saved Browser Silk funnel for each mapped domain. Perform full licence adjudication only when audit-request says the package will be publicly redistributed.
Completion criteria
Every verdict-relevant source statement has an independent identity, version, exact-passage, context, correction or retraction, and comparability check or a visible access limitation.
Internal audits record access authority and provenance. Full licence verification is required only when public redistribution is intended or authorized.
For a selector_not_found failure, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. Preserve the last failure evidence for runtime attempt handling; never switch to manual browsing.
No snippet-only statement passes verification and sources unrelated to the verdict are not reopened.
One bounded locator correction may be made; unresolved source problems remain limitations.
If a sensitivity result raises a new verdict-critical methods question, perform at most one bounded follow-up using the same source limits and record it in source-verification.
Adjudicate findings and manuscript changes
Runs asScientific AdjudicatorJoin the independent statistical and source checks, assign final finding status and allowed wording, and write specific manuscript changes without introducing new analyses or source claims.
Completion criteria
Every retained finding is verified, verified with limits, rejected, incomplete, or disputed, with severity, confidence, exact evidence, allowed wording, and no duplicated admissibility review.
Each manuscript recommendation names the affected section or statement, the exact requested change, its evidence, urgency, and whether it is required or optional.
audit_ready requires an independent rerun from a separate local working directory with the same frozen scientific contract and source and data hashes, complete required numerical comparisons, protected source inputs, recorded environment details, and no unresolved scientific blocker.
Different runner, workspace, mount, network-control, memory-control, or storage-control identities do not block audit_ready by themselves.
Weak nonblocking findings are dropped; operationally_blocked is reserved for missing authority or an unsafe or impossible required verification.
Outcomes
audit_readyAll required findings are independently verified and ready for reporting.
verification_incompleteThe audit can report verified work with explicit unresolved limitations.
operationally_blockedRequired verification cannot proceed safely or with authority.
Agents
Audit Report Builder
Renders verified canonical objects into consistent HTML, PDF, JSON, and replay-package outputs.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read only frozen verified canonical objects and approved wording. Write project://peer2paper/audits/{runId}/delivery/audit.html, audit.pdf, audit.json, replay/, and manifest.json.
- Render reproduction, robustness, and evidence-alignment statuses separately from one pinned template version and preserve exact numbers, order, links, limitations, and hashes across formats.
- NEVER add a scientific finding, promote rejected material, or rewrite a verified conclusion beyond its allowed wording.
Release Quality Reviewer
Checks completeness, consistency, provenance, replay, privacy, accessibility, licensing, and release boundaries.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read the rendered audit package and all declared manifests. Write project://peer2paper/audits/{runId}/delivery/release-gates.json.
- Check numerical consistency, evidence links, replay commands, expected hashes, language rules, accessibility, privacy, redaction, licences, and access restrictions.
- NEVER repair scientific content during quality review or authorize external release beyond the dispatched release policy.
Claim Mapper
Extracts candidate claims and connects the selected claim to exact reported results and source locations.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read paper documents and the artifact registry. Write project://peer2paper/audits/{runId}/study-case/claim-records.json and selected-claim.json.
- Separate measured results from interpretation and record population, exposure or treatment, comparator, outcome, timepoint, effect measure, direction, scope, and reported-result links.
- NEVER finalize an ambiguous target claim or fabricate a missing source location; emit a typed unresolved decision instead.
Variable Mapper
Connects paper concepts, dataset fields, and code references while preserving units, timepoints, roles, and derivations.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read paper documents, dataset maps, code project, and selected claim. Write project://peer2paper/audits/{runId}/study-case/variable-maps.json and evidence-graph.json.
- Preserve raw names and label normalized and derived meanings with transformation rules, confidence, provenance, and review state.
- NEVER guess high-impact mappings, missing-value meanings, observation units, nesting, or scientifically meaningful recodes.
Peer2Paper Orchestrator
Orchestrates the Peer2Paper agentic workflow from intake through completion, coordinating work, tracking dependencies and outcomes, and escalating blockers or approval-sensitive actions without exceeding delegated authority.
- Workflow orchestration
- Task coordination
- Dependency and status tracking
- Blocker escalation
- Outcome reporting
Reproduction Engineer
Builds clean R or Python execution recipes, reruns supplied analyses, compares results, and tests minimal technical repairs.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read project://peer2paper/audits/{runId}/study-case/frozen-study-case.json and its referenced execution copies. Write declared files under project://peer2paper/audits/{runId}/reproduction/.
- Use deterministic repository scripts and isolated execution facilities supplied to the step for environment detection, clean runs, numeric comparison, hashing, and replay checks.
- Run supplied analysis twice from separate clean states when executable. NEVER call a scientific change an exact reproduction or enable network access without explicit permission.
- Write project://peer2paper/audits/{runId}/reproduction/reproduction-package.json with one allowed reproduction status and linked run evidence.
Data Choices Analyst
Proposes defensible sample, missingness, variable-construction, and population choices without seeing candidate outcomes.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read frozen study-case mappings, the estimand, and conclusion rules. Write separate lane files under project://peer2paper/audits/{runId}/robustness/decision-lanes/data/.
- Propose sample, exclusion, missingness, variable-construction, and population options with scientific rationale and compatibility constraints before execution.
- NEVER choose a method because of an observed estimate, interval, direction, or p-value.
Inference Analyst
Proposes valid clustering, weighting, uncertainty, and multiplicity choices for the frozen design and estimand.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read the frozen study case, study design, estimand, and execution constraints. Write inference lane files under project://peer2paper/audits/{runId}/robustness/decision-lanes/inference/.
- Propose clustering, weights, standard errors, uncertainty procedures, and multiplicity rules with diagnostic requirements and compatibility constraints.
- NEVER approve an inference method after seeing whether it changes support.
Model And Covariate Analyst
Proposes design-compatible statistical models and adjustment sets while preserving the frozen estimand.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read the frozen study case, study design, estimand, and variable map. Write model and covariate lane files under project://peer2paper/audits/{runId}/robustness/decision-lanes/modeling/.
- Specify model families, functional forms, interactions, adjustment sets, and compatibility constraints with explicit rationale.
- NEVER treat every control combination as valid or change the scientific question silently.
Robustness Lead
Freezes the estimand and conclusion rules, integrates analysis-decision lanes, registers candidates, and produces the robustness map.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read the frozen study case and reproduction package. Write project://peer2paper/audits/{runId}/robustness/estimand.json, conclusion-rules.json, analysis-decision-space.json, candidate-registry.json, and robustness-map.json.
- Register candidate rationale and decision combinations before result execution, merge only compatible choices, and preserve invalid, failed, and untested regions.
- NEVER reveal candidate results to method reviewers or claim universal robustness from bounded coverage.
Blind Method Reviewer Alpha
Independently judges candidate-analysis validity from results-blind packets.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read only results-blind packets and declared design materials. Write project://peer2paper/audits/{runId}/verification/method-review-alpha.json.
- Judge same-estimand fit, design fit, exclusions, controls, missingness, inference, rationale, and reproducibility using declared criteria.
- NEVER inspect hidden candidate outcomes or coordinate a verdict with the other method reviewer.
Blind Method Reviewer Beta
Provides a second independent judgment of candidate-analysis validity from results-blind packets.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read only results-blind packets and declared design materials. Write project://peer2paper/audits/{runId}/verification/method-review-beta.json.
- Judge same-estimand fit, design fit, exclusions, controls, missingness, inference, rationale, and reproducibility using declared criteria.
- NEVER inspect hidden candidate outcomes or coordinate a verdict with the other method reviewer.
Scientific Adjudicator
Resolves verification disagreements and locks final validity, severity, confidence, and allowed wording.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read independent method reviews, statistical checks, source checks, conflict records, and frozen conclusion rules. Write project://peer2paper/audits/{runId}/verification/final-adjudication.json.
- Apply declared resolution rules, preserve dissent and limitations, and return typed correction, evidence, ready, incomplete, or disputed outcomes.
- NEVER introduce a new analysis, source claim, or unsupported accusation during adjudication.
Source Evidence Verifier
Independently opens sources and checks whether exact passages support proposed literature findings.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Use toone-scholar and Toone's persistent browser with reusable Silk flows to verify authorized source versions, full text, corrections, retractions, and exact passages. Write project://peer2paper/audits/{runId}/verification/source-checks.json.
- Check bibliographic identity, context, population, measures, methods, results, licence, and comparability independently of the source extractor.
- NEVER accept a citation from a snippet, bypass access controls, or broaden the claim beyond the verified passage.
Statistical Evidence Verifier
Independently clean-reruns important analyses and validates statistical findings against frozen rules.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read approved candidate definitions, execution recipes, frozen inputs, and native run outputs. Write project://peer2paper/audits/{runId}/verification/statistical-checks.json and independent-runs/.
- Use supplied deterministic scripts and isolated execution facilities to rerun important analyses and verify IDs, samples, variables, models, seeds, diagnostics, estimates, intervals, and classifications.
- NEVER reuse a proposing analyst's unsupported conclusion or alter the frozen candidate during verification.
Verification Lead
Freezes verification inputs, reconciles independent checks, preserves disagreements, and issues final adjudications.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Read frozen reproduction, robustness, and literature packages. Write project://peer2paper/audits/{runId}/verification/verification-packet.json, resolved-finding-graph.json, and adjudications.json.
- Keep proposed methods separate from observed results until method review is complete, preserve review disagreements and dependency chains, and assign validity, severity, confidence, and allowed wording separately.
- NEVER erase rejected findings from history, convert uncertainty into confidence, or adjudicate without the required independent checks.
Literature Searcher
Runs reproducible scholarly searches and records complete source, query, access, and screening logs.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Use toone-scholar for Europe PMC, PubMed, citation, preprint, and open-access lookup. Use browser and browser-silk for recurring web sources through Toone's persistent browser session and saved Silk flows.
- Read the frozen research brief. Write project://peer2paper/audits/{runId}/research/search-log.json and source-registry.json.
- Annotate and save reusable Silk flows for recurring browser sources while working. NEVER bypass paywalls, store credentials, or cite snippets as evidence.
Researcher
Conducts organization-wide research, gathers and verifies information from available internal integrations and web sources, and delivers evidence-based findings while respecting access and approval boundaries.
- Organizational research
- Web research
- Source verification
- Research synthesis
Source Extractor
Maps accessible full texts into structured study objects and exact evidence passages with provenance.
- Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
- Use toone-scholar and Toone's persistent browser with reusable Silk flows to open authorized full texts and supplements. Write files under project://peer2paper/audits/{runId}/research/full-text/ and extracted-study-objects.json.
- Preserve raw, normalized, and derived layers; attach stable section, page, table, figure, and passage locations to every extracted field.
- NEVER guess unreadable content, claim equivalence between unlike measures, or treat abstract-only evidence as full-text evidence.
Requirements
- Included
- Audit Report Builder
- Release Quality Reviewer
- Claim Mapper
- Variable Mapper
- Peer2Paper Orchestrator
- Reproduction Engineer
- Data Choices Analyst
- Inference Analyst
- Model And Covariate Analyst
- Robustness Lead
- Blind Method Reviewer Alpha
- Blind Method Reviewer Beta
- Scientific Adjudicator
- Source Evidence Verifier
- Statistical Evidence Verifier
- Verification Lead
- Literature Searcher
- Researcher
- Source Extractor
- Routine-owned versioned schemas for canonical findings, claims, actions, limitations, and frozen audit objects.
- Project-owned pinned template used for the researcher-facing HTML report.
- Routine-owned pinned JSON Schema for the machine-readable audit.
- Project-owned versioned scientific execution runner used by reproduction, robustness, and independent statistical verification. A missing runner file blocks execution.
- You provide
- Local directory containing the paper or draft, data, code, environment files, supplements, data dictionaries, and authorized supporting sources.
- Required audit scope containing one exact target claim and location, permissions, privacy classification, retention rules, public-redistribution intent, and external-search authorization.
- Required scientific rules for estimands, tolerances, admissibility, comparability, severity, disagreement, and allowed defaults. Supply an empty JSON object only when the routine's recorded defaults are intended.
- Execution rules for runtime, attempt-local working directories, host network policy, operational memory and storage defaults, a 600-second command timeout, source immutability, source access, and internal-only release.
- Required only after partial_reproduction or numerical_mismatch. Supply a JSON decision of continue or stop, with a reason and the authorizing person or role.
- Produces
- Target claim and relevant document map (json)
- Target data and variable map (json)
- Target code and dependency map (json)
- Verified supplied-document consistency findings (json)
- Compact frozen study case contract (json)
- Frozen reproduction contract and inputs (json)
- Execution readiness (json)
- Execution profile attempt (json)
- Untouched primary run checkpoint (json)
- Primary numerical comparison (json)
- Frozen repair candidates (json)
- Repair attempt 1 checkpoint (json)
- Repair attempt 2 checkpoint (json)
- Repair attempt 3 checkpoint (json)
- Secondary target 1 checkpoint (json)
- Secondary target 2 checkpoint (json)
- Secondary target 3 checkpoint (json)
- Per-target reproduction evidence (directory)
- Reproduction package (json)
- Frozen estimand and conclusion rules (json)
- Data and measurement candidates (json)
- Model specification candidates (json)
- Inference candidates (json)
- Ranked sensitivity candidate pool (json)
- Primary blind admissibility review (json)
- Conditional second blind review (json)
- Approved sensitivity candidates (json)
- Isolated sensitivity runs (json)
- Sensitivity result and coverage map (json)
- Sensitivity version history (directory)
- Targeted research questions (json)
- Screened source shortlist and slot plan (json)
- Source evidence checkpoint 1 (json)
- Source evidence checkpoint 2 (json)
- Source evidence checkpoint 3 (json)
- Assembled targeted full-text evidence (json)
- Targeted literature evidence package (json)
- Targeted evidence version history (directory)
- Frozen important verification targets (json)
- Independent statistical rerun (json)
- Independent source and document verification (json)
- Verified finding package (json)
- Concrete manuscript recommendations (json)
- Canonical audit JSON (json)
- Replay package (directory)
- Researcher audit report HTML (html)
- Scoped release checks (json)
- Audit package manifest (json)
MCP servers
browserbrowser-silktoone-scholar
