toone

Routine

micro1/peer2paper

Audit one caller-supplied quantitative claim by freezing a compact case with an embedded analysis contract, reproducing that contract under bounded local execution, running authorized sensitivity and targeted evidence work, independently verifying decisive findings, and returning a lean replayable audit package.

By
Matheus Paranhos
Approved
License
toone-community-v1
Steps
47
Agents
19
Sub-routines
6

Purpose

Audit one caller-supplied quantitative claim by freezing a compact case with an embedded analysis contract, reproducing that contract under bounded local execution, running authorized sensitivity and targeted evidence work, independently verifying decisive findings, and returning a lean replayable audit package.

Status: Production | Updated: 2026-08-31

Steps

  1. Build compact study case

    Validate the target, map only its document and data dependency closure, preserve exact source-backed transformations and document conflicts, and freeze one study case containing primary_target and analysis_contract.

    Completion criteria

    • The child returns one frozen study case with mandatory primary_target, compact analysis_contract, separate document_conflicts, bounded secondary benchmarks, and target-scoped data and code maps, or an explicit blocker.

    • A unique caller-frozen estimand with a contradictory supplementary document emits case_frozen_with_limits and remains eligible for reproduction.

    Outcomes

    • case_frozen

      The claim and every required execution link are fixed without a blocking gap.

    • case_frozen_with_limits

      The case can execute with visible nonblocking limitations.

    • case_not_assessable

      The target cannot be fixed without an unsupported substantive choice.

    • operationally_blocked

      Authority, privacy, safety, or inaccessible required files block execution.

    • Routine peer2paper-operations/build-and-freeze-study-case
  2. Preflight and reproduce primary result

    Preflight the real case bundle and shared execution capability, run the contract-defined model unchanged, test bounded working-directory or input-binding repairs when needed, and preserve complete scientific outputs in one reproduction package.

    Completion criteria

    • The child receives the real case bundle and frozen analysis_contract, captures sample, missingness, coefficient, standard error, interval, statistic, degrees of freedom and p-value, and returns a self-contained reproduction package.

    • Execution provenance and scientific comparison remain separate: execution_mode is untouched or technically_repaired_original, while the emitted routine outcome follows comparison_status.

    • Verified network or DNS denial is not a prerequisite. Memory and storage are operational defaults, and unavailable exact enforcement does not emit not_executable when the analysis can run in an attempt-local working directory.

    Outcomes

    • exact_reproduction

      The untouched primary target and required benchmarks reproduce exactly under frozen comparison rules.

    • within_tolerance

      The untouched primary target and required benchmarks satisfy the frozen tolerance rules.

    • partial_reproduction

      The primary target or required benchmarks reproduce only in part and need an explicit continuation decision.

    • numerical_mismatch

      The executable analysis remains numerically inconsistent after the one remediation pass.

    • not_executable

      A concrete runtime, dependency, path, permission, isolation, or resource-control blocker prevents execution.

    • not_assessable

      The target cannot be assessed without an unsupported substantive choice.

    • Routine peer2paper-operations/reproduce-original-result
  3. Decide whether to continue after reproduction

    Runs asPeer2Paper Orchestrator

    Review a partial reproduction or unresolved numerical mismatch and apply the caller's explicit authorization boundary before expensive downstream analysis.

    Completion criteria

    • The supplied decision is continue or stop, names the reason and authorizing person or role, and applies to the current reproduction package.

    • continue_to_analysis is emitted only with explicit authorization to interpret the partial reproduction or numerical mismatch; otherwise stop_after_reproduction is emitted.

    Outcomes

    • continue_to_analysis

      The caller authorizes robustness and targeted research despite the recorded reproduction limitation.

    • stop_after_reproduction

      The run ends with the case contract, document-consistency findings, reproduction package, and diagnostic evidence.

  4. Test bounded sensitivity candidates

    Run three result-blind design lanes, rank no more than five candidates, always materialize the conditional second-review artifact, and execute approved analyses through the ready reproduction profile.

    Completion criteria

    • No more than five admissible candidates produce a common-schema sensitivity result, with untested dimensions and execution limits visible.

    Outcomes

    • robustness_ready

      Required sensitivity checks completed without residual gaps.

    • robustness_ready_with_limits

      Sensitivity checks completed with visible nonblocking gaps.

    • analysis_not_assessable

      A blocking scientific or executable gap prevents a defensible sensitivity conclusion.

    • operationally_blocked

      Authority or unsafe execution prevents required work.

    • sensitivity_blocked

      Sensitivity preparation found a scientific, source/data, or execution-capability blocker.

    • Routine peer2paper-operations/stress-test-analysis
  5. Answer targeted evidence questions

    Run beside sensitivity work after reproduction succeeds or is explicitly continued. Answer no more than three questions raised by the claim, reproduction, or supplied-document conflicts, with no more than three deep-source checkpoints.

    Completion criteria

    • The child returns exact passages, comparability judgments, verification state, and manuscript consequences for no more than three targeted questions and three deep-mapped sources.

    Outcomes

    • evidence_ready

      Required targeted evidence is available and supported without residual gaps.

    • evidence_ready_with_limits

      The targeted search completed with visible nonblocking gaps.

    • evidence_not_assessable

      A required evidence question cannot be answered from authorized accessible sources.

    • operationally_blocked

      Authority or access policy prevents required research.

    • Routine peer2paper-operations/research-methods-and-evidence
  6. Verify decisive findings

    Wait for sensitivity and targeted research, then independently rerun the frozen analysis contract from the real case bundle in a clean workspace with equivalent isolation and resource controls, reopen verdict sources, and produce adjudicated findings.

    Completion criteria

    • The child uses the same frozen analysis contract and source and data hashes, captures complete required numerical comparisons in a clean attempt-local workspace, and returns allowed wording plus concrete manuscript consequences.

    • Different runner or mount identities do not fail verification when the scientific contract, source and data hashes, dependencies, source immutability, timeout, and recorded environment are equivalent for the claim.

    • External release stays disabled unless separately approved.

    Outcomes

    • audit_ready

      All required findings are independently verified and ready for reporting.

    • verification_incomplete

      The audit can report verified work with explicit unresolved limitations.

    • operationally_blocked

      Required verification cannot proceed safely or with authority.

    • Routine peer2paper-operations/verify-and-adjudicate-findings
  7. Publish JSON-first audit package

    Build one canonical JSON audit from the verified findings and complete manuscript recommendations, render concise HTML from that JSON, and package replay evidence and scoped release checks.

    Completion criteria

    • One canonical JSON audit, concise HTML report, replay package, and package manifest agree on all numbers, findings, recommendations, and limitations.

    Outcomes

    • audit_completed

      The audit and replay package pass every required scoped check.

    • audit_completed_with_limits

      The audit is releasable with explicit nonblocking limitations.

    • operationally_blocked

      A required completeness, replay, authority, or scoped release check blocks packaging.

    • Routine peer2paper-operations/assemble-final-audit

Sub-routines

Sub-routines

Assemble Final Audit

Purpose

Publish one canonical JSON audit, a concise researcher-facing HTML report, and a replay package with scoped release checks. PDF is outside the required hackathon output.

Steps

  1. Build canonical audit and replay package

    Runs asAudit Report Builder

    Build one canonical JSON audit and replay directory from the frozen case, reproduction, sensitivity, targeted evidence, adjudicated findings, and complete manuscript recommendations. Keep raw logs, hashes, prompts, and provenance in the machine package.

    Completion criteria

    • The canonical JSON contains the verdict, reported and reproduced values, numerical differences, sensitivity result, verified conflicts, concrete manuscript recommendations, limitations, and content-addressed references to technical evidence.

    • The JSON validates against the pinned schema and uses one value for every displayed number and statement.

    • The replay package contains the shared execution profile, exact commands, 600-second timeout, environments, inputs, outputs, checks, and hashes needed to rerun every scientific result included in the verdict.

    • Raw logs, IDs, prompts, and detailed provenance remain in the machine package rather than the researcher report.

    • Every concrete manuscript recommendation from the verified recommendations resource appears in the canonical audit with its affected section or statement, exact change, evidence, urgency, and required or optional status.

    • The package is internal-only unless audit-request contains separate explicit external-release approval.

  2. Render concise researcher report

    Runs asAudit Report Builder

    Render a two-to-three-page HTML report from the canonical JSON using the pinned template.

    Completion criteria

    • The HTML presents the verdict, reported versus reproduced result, sensitivity summary, verified contradictions or ambiguities, exact manuscript changes, and remaining limitations.

    • Every number and allowed statement is read from the canonical JSON, and detailed provenance is linked rather than expanded into large report tables.

    • The HTML is accessible, readable, and contains no unsupported scientific finding.

  3. Check and package audit

    Runs asRelease Quality Reviewer

    Check cross-format identity, replay evidence, completeness, and internal release boundaries, then write the final manifest. Run licence and retention checks only when the audit request authorizes public redistribution.

    Completion criteria

    • Every displayed number and verdict statement matches the canonical JSON, and each scientific result has replayable evidence.

    • Release checks cover completeness, JSON-to-HTML consistency, provenance, replay, accessibility, and internal access boundaries.

    • External release is disabled unless audit-request contains separate explicit approval. Full licence and retention adjudication runs only with that approval; otherwise the manifest records internal-only scope.

    • The manifest reports pass, limited, or blocking for each required check and never lists PDF as a required core output.

    • Emit audit_completed only with no blocking or limited gate; emit audit_completed_with_limits only with no blocking gate and at least one visible limitation.

    Outcomes

    • audit_completed

      The audit and replay package pass every required scoped check.

    • audit_completed_with_limits

      The audit is releasable with explicit nonblocking limitations.

    • operationally_blocked

      A required completeness, replay, authority, or scoped release check blocks packaging.

Sub-routines

Build And Freeze Study Case

Purpose

Freeze one caller-supplied quantitative claim into a compact study case containing the mandatory primary target, exact source-backed analysis contract, target dependency closure, bounded secondary benchmarks, and separately preserved document conflicts, without executing a scientific model during intake.

Steps

  1. Validate target claim and map relevant documents

    Runs asClaim Mapper

    Validate the caller-supplied target claim and map only the primary paper, supplement, preregistration, and passages needed to identify its reported result and wording.

    Completion criteria

    • One target claim is fixed with exact text and location, population, exposure or intervention, comparator, outcome, timepoint, observation unit, estimand, directional orientation, conclusion-change rules, and known ambiguities.

    • The primary reported estimate, interval, p-value, sample size, and exact table, figure, or code target are recorded when present.

    • Only the primary paper, supplement, preregistration, and claim-relevant passages are mapped; unreadable or missing required material is explicit.

    • Record only values reported in authorized documents and their exact locations. Do not execute supplied analysis code, fit a model, or compute a diagnostic estimate, interval, statistic, p-value, or model sample size.

  2. Map target data and variables

    Runs asVariable Mapper

    Map only columns in the primary target dependency closure, including sample filters and variables required to derive those columns. Do not profile unrelated columns or search for participant keys unless target execution requires participant identity.

    Completion criteria

    • Every primary-target element and bounded secondary benchmark has a direct data-field, derivation, sample-filter, or documented missing-link mapping within the target dependency closure.

    • The map records units, condition codes, derivations, target-relevant missingness, observation unit, and unresolved high-impact ambiguity without profiling unrelated columns.

    • Single-column uniqueness scans and combination-key searches are prohibited unless participant identity is required by the frozen target; any permitted identity check is bounded to the minimum declared columns.

    • The checkpoint reports counts and stable lists of mapped columns, inspected columns, and intentionally excluded columns, with inspected columns limited to the target dependency closure.

    • Map fields, filters, derivations, and target-relevant missingness without fitting a statistical model or calculating the target result.

  3. Map target code and dependencies

    Runs asReproduction Engineer

    Identify the untouched entrypoint, target-producing code, exact formula, sample filter, transformations, factor levels, coefficient and orientation, dependencies, paths, inputs, outputs, seeds, and no more than three secondary benchmarks.

    Completion criteria

    • The primary command and one unique target-producing code path are identified, or the exact missing or ambiguous entrypoint is recorded.

    • The code map records source file identities and hashes plus exact source lines and literal expressions for the formula, population filter, transformations, factor levels, reference group, coefficient, orientation, and missing-data declaration.

    • Expressions present in source are copied exactly into the code map and are not reconstructed from prose or line ranges.

    • A dependency manifest records languages, lockfiles, imports, packages, system requirements, paths, mounts, permissions, network needs, seeds, and missing files.

    • The mandatory primary target is separate from no more than three secondary benchmarks.

    • This step is static code and dependency inspection. Parse-only validation is permitted, but evaluating scientific source, fitting a model, or generating result values is prohibited.

  4. Verify supplied-document consistency

    Runs asSource Evidence Verifier

    Compare the primary paper, supplement, preregistration, and completed code map only where they define the target population, variables, condition codes, analysis, or reported result. Record exact conflicts and manuscript consequences; do not remap unrelated content.

    Completion criteria

    • Every retained conflict links exact passages or code locations, identifies the affected claim or execution element, and is verified or marked unresolved.

    • The check is limited to the target paper, supplement, preregistration, and completed code map.

    • Compare supplied document passages and static code locations only. Do not execute the analysis or use a newly calculated model result as document-consistency evidence.

  5. Bind and freeze analysis

    Runs asClaim Mapper

    Join the claim, target-scoped data map, exact source-backed code map, and verified document findings into one frozen study case with a compact analysis_contract section.

    Completion criteria

    • primary_target is mandatory on every frozen case and remains separate from the bounded secondary_benchmarks array.

    • analysis_contract contains code_hash, data_hash, formula, population_filter, exact transformations, factor levels and reference group, coefficient, orientation, missing-data declaration, required columns, entrypoint, and dependency identity copied from the code and data maps.

    • document_conflicts remains a separate collection linked to exact supplied passages or code locations and does not overwrite the primary target or analysis contract.

    • When the caller-frozen estimand maps to one unique supplied code path and the audit request, scientific policy, primary evidence, and code agree, a contradictory supplementary document emits case_frozen_with_limits without stopping reproduction.

    • Emit case_frozen only when the unique analysis contract has no nonblocking conflict; emit case_not_assessable only when no unique target exists or choosing a code path would change the caller-frozen estimand; emit operationally_blocked only for authority, safety, privacy, or inaccessible required files.

    • New intake artefacts must not contain source_consistency_probe or any model-derived diagnostic result. Scientific execution evidence begins only in Reproduce Original Result through the official runner.

    Outcomes

    • case_frozen

      The caller-frozen estimand maps to one schema-valid executable target without a blocking gap or source conflict.

    • case_frozen_with_limits

      The executable target is schema-valid and may run, with a visible nonblocking supplied-document conflict or other declared limitation.

    • case_not_assessable

      No unique executable target exists, or choosing a supplied code path would change the caller-frozen estimand.

    • operationally_blocked

      Authority, privacy, safety, or inaccessible required files block execution.

Sub-routines

Reproduce Original Result

Purpose

Execute the exact analysis_contract embedded in the frozen study case through the shared runner, using an attempt-local working copy and bounded technical repairs without changing scientific expressions, then report execution_mode separately from the component-level scientific comparison.

Status: Production | Updated: 2026-08-31

Steps

  1. Freeze reproduction contract and inputs

    Runs asReproduction Engineer

    Freeze the mandatory primary target and embedded analysis_contract, a separate array of no more than three secondary benchmarks, comparison rules, untouched command, data and source hashes, and dependencies.

    Completion criteria

    • The frozen study case contains primary_target and analysis_contract with code and data hashes, formula, population filter, exact transformations, factor levels, coefficient, orientation, missing-data declaration, entrypoint, and dependencies.

    • The supplied case bundle contains the source and data identities referenced by the analysis contract.

    • The checkpoint fixes one primary target and a separate array of no more than three secondary benchmarks with reported estimate, interval, statistic, p-value, sample size, and comparison rules.

  2. Preflight shared execution

    Runs asReproduction Engineer

    Use the project execution runner to validate the real runtime, dependencies, input hashes, attempt-local working copy, source immutability, host network policy, operational controls, cooperative CPU settings, and a trivial language command before the study command.

    Completion criteria

    • Execution readiness records the case bundle, source and data hashes, runtime, dependencies, attempt-local working directory, registered input paths, permissions, host network policy, timeout, and available controls with exact evidence.

    • When execution may proceed, execution-readiness and execution-profile set status to ready and execution_allowed to true. They preserve nonblocking control gaps in limitations rather than inventing another readiness status.

    • maximum_cpu_cores is best-effort unless execution-policy explicitly requires hard enforcement for safety, authorization, or infrastructure protection. Set available cooperative thread controls to the requested value and record their names and values.

    • When hard CPU enforcement is unavailable but not explicitly required, include a limitation with type cpu_limit_best_effort, requested_cores from execution-policy, and hard_enforcement_available false; keep status ready and execution_allowed true.

    • Hard CPU enforcement may block only when execution-policy explicitly marks it as required for safety, authorization, or infrastructure protection and no safe authorized execution path exists.

    • A missing runtime or dependency, unreadable required file, source or data hash mismatch, missing permission to create a protected attempt-local working copy, unenforceable required network isolation, or another explicit safety or authority requirement may block execution.

    • Execution profile is always written as an attempt envelope. Operational controls and limitations are preserved for downstream consumers, which trust execution_allowed rather than reinterpreting a best-effort CPU field.

    • Every scientific command has a 600-second hard timeout. No advisory routine budget is inferred.

  3. Run untouched primary target

    Runs asReproduction Engineer

    When a clean supplied entrypoint exists, copy the required inputs into a separate attempt-local working directory and execute the unchanged entrypoint once. Do not generate or use a wrapper in this step; otherwise write a skipped checkpoint for bounded repair.

    Completion criteria

    • The checkpoint is written for every run with status completed, failed, blocked, skipped_for_repair, or not_assessable and records the executable-contract hash.

    • This step never creates or runs a wrapper. An unresolved working directory, input binding, output capture, or section-selection need writes skipped_for_repair and proceeds to the frozen repair candidates.

    • Before execution, the runner verifies source and data hashes, parses any R command file without evaluating it, and confirms that the working directory is inside the attempt directory.

    • A completed attempt proves that the source hashes, population filter, formula, transformations, treatment arms, factor levels, reference group, contrast, missing-data behavior, and coefficient orientation match the executable contract.

    • A completed or failed attempt records the command, 600-second timeout, working directory, runtime, dependencies, inputs, outputs, logs, environment details, and hashes. Verified inputs are hashed before and after and must remain unchanged.

    • Model diagnostics include initial eligible rows, model-frame rows, nobs(), excluded-row count, na.action class and excluded-row evidence, coefficient and standard error, test statistic and degrees of freedom, p-value, interval endpoints, interval method, confidence level, and coefficient orientation when scientific source execution reaches them.

    • A best-effort CPU limitation never changes a ready execution profile to blocked and never prevents the untouched attempt or an eligible bounded repair.

  4. Compare primary result

    Runs asReproduction Engineer

    Compare the untouched primary attempt with the reported result and executable contract. Score estimate, interval, statistic, p-value, and sample independently, and check internal numerical coherence without treating one matching component as evidence for another.

    Completion criteria

    • The checkpoint records reported and reproduced values, absolute and relative differences, tolerance basis, displayed precision, and a separate pass, fail, not_reported, or not_comparable status for estimate, each interval endpoint, statistic, p-value, and sample.

    • The checkpoint records initial eligible rows, model-frame rows, nobs(), excluded-row count, na.action class and evidence, coefficient standard error, degrees of freedom, interval method, confidence level, and coefficient orientation.

    • For reported estimate 0.028 and interval [0.006, 0.041], the checkpoint records midpoint 0.0235 and coefficient-to-midpoint discrepancy 0.0045, then determines the interval method from evidence instead of inferring interval validity from the coefficient.

    • exact_reproduction or within_tolerance is eligible only when every comparable required numerical component passes its frozen rule and the executed model specification matches the contract.

    • Missing execution evidence is classified from the preceding checkpoint and never treated as a numerical result.

    • Record execution_mode separately as untouched or technically_repaired_original. Record comparison_status separately as exact_reproduction, within_tolerance, partial_reproduction, numerical_mismatch, or not_executable.

    • A technically repaired execution is scored under the same component rules as an untouched execution; repair provenance cannot replace or determine the scientific comparison status.

  5. Diagnose failure and freeze repairs

    Runs asReproduction Engineer

    After an unchanged execution failure, a skipped clean entrypoint, or a numerical mismatch, freeze no more than three ordered technical repairs from the explicit allowlist. Record denylisted proposals as rejected and never execute them.

    Completion criteria

    • The checkpoint always exists and has status candidates_frozen or not_required.

    • Allowed repairs are limited to binding the attempt-local working directory, wrapping unchanged source for launch or output capture, selecting an author-declared replication section, binding registered input paths, and capturing runtime dependencies.

    • Filters, sample rules, formulas, transformations, factor levels, treatment arms, contrasts, missing-data behavior, intervals, and inferential methods cannot change.

    • Each wrapper repair requires ASCII-safe string encoding with encodeString(value, quote = '"'), forbids dQuote and sQuote, persists the wrapper in the attempt directory, and requires parse-only validation before scientific execution.

    • When the source declares a replication start, contains an empty setwd call, reads a registered local input, and may fail later during optional output work, rank as candidate 1 a wrapper that runs from the writable attempt directory, binds or copies the registered input there, starts at the author-declared replication section, neutralizes only the empty setwd call, and captures the target model before optional plotting or report-generation failures.

  6. Test repair candidate 1

    Runs asReproduction Engineer

    Test the first frozen repair in an isolated execution copy when assigned and needed; otherwise write a skipped or not_required checkpoint.

    Completion criteria

    • The first candidate always writes a checkpoint with attempted, parse_failed, completed, failed, skipped, or not_required status and the exact reason.

    • A generated R wrapper is stored in the attempt directory, uses ASCII-safe encoded string literals, contains no dQuote, sQuote, or typographic quotes, and passes parse-only validation before supplied scientific source is evaluated.

    • An attempted repair records the wrapper content and hash, operational change, unchanged scientific expressions, source and data hashes before and after, the 600-second timeout, execution log, and outcome.

    • A successful working-directory, launch, section-selection, output-capture, input-binding, or dependency-capture repair is classified technically_repaired_original.

    • A parse failure or runtime failure consumes only this bounded candidate, preserves its evidence, and allows the next frozen candidate to run. A repair that changes the analysis contract cannot support patched_reproduction.

    • A completed repaired run captures initial eligible rows, model rows, nobs(), excluded rows and na.action evidence, coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.

    • For the declared replication-section repair, the wrapper runs through the official runner from the attempt directory, binds or copies the registered CSV into that directory, begins at the author-declared replication marker, neutralizes only setwd("") in that section, and captures the target model and diagnostics before optional plotting.

    • The checkpoint records execution_mode technically_repaired_original plus component-level comparison evidence; it does not use repaired execution as a scientific verdict.

    • Missing hard CPU affinity cannot skip this repair when execution_allowed is true.

  7. Test repair candidate 2

    Runs asReproduction Engineer

    Test the second frozen repair only when the first did not yield a valid reproduced primary result; otherwise write a skipped or not_required checkpoint.

    Completion criteria

    • The second candidate always writes a checkpoint with attempted, parse_failed, completed, failed, skipped, or not_required status and the exact reason.

    • A generated R wrapper is stored in the attempt directory, uses ASCII-safe encoded string literals, contains no dQuote, sQuote, or typographic quotes, and passes parse-only validation before supplied scientific source is evaluated.

    • An attempted repair records the wrapper content and hash, operational change, unchanged scientific expressions, source and data hashes before and after, the 600-second timeout, execution log, and outcome.

    • A successful working-directory, launch, section-selection, output-capture, input-binding, or dependency-capture repair is classified technically_repaired_original.

    • A parse failure or runtime failure consumes only this bounded candidate, preserves its evidence, and allows the next frozen candidate to run. A repair that changes the analysis contract cannot support patched_reproduction.

    • A completed repaired run captures initial eligible rows, model rows, nobs(), excluded rows and na.action evidence, coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.

    • Missing hard CPU affinity cannot skip this repair when execution_allowed is true.

  8. Test repair candidate 3

    Runs asReproduction Engineer

    Test the third frozen repair only when earlier candidates did not yield a valid reproduced primary result; otherwise write a skipped or not_required checkpoint.

    Completion criteria

    • The third candidate always writes a checkpoint with attempted, parse_failed, completed, failed, skipped, or not_required status and the exact reason.

    • A generated R wrapper is stored in the attempt directory, uses ASCII-safe encoded string literals, contains no dQuote, sQuote, or typographic quotes, and passes parse-only validation before supplied scientific source is evaluated.

    • An attempted repair records the wrapper content and hash, operational change, unchanged scientific expressions, source and data hashes before and after, the 600-second timeout, execution log, and outcome.

    • A successful working-directory, launch, section-selection, output-capture, input-binding, or dependency-capture repair is classified technically_repaired_original.

    • A parse failure or runtime failure consumes only this bounded candidate, preserves its evidence, and allows the next frozen candidate to run. A repair that changes the analysis contract cannot support patched_reproduction.

    • A completed repaired run captures initial eligible rows, model rows, nobs(), excluded rows and na.action evidence, coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.

    • Missing hard CPU affinity cannot skip this repair when execution_allowed is true.

  9. Run secondary target 1

    Runs asReproduction Engineer

    After the primary result succeeds unchanged or through a valid repair, run the first required secondary benchmark if assigned. Otherwise write not_assigned, blocked, or skipped.

    Completion criteria

    • The checkpoint always exists and records one benchmark comparison or the exact reason it was not run.

    • The scientific command runs when the primary target has valid comparable execution evidence under either execution_mode, including a partial component comparison; it does not run after not_executable or not_assessable.

  10. Run secondary target 2

    Runs asReproduction Engineer

    After the primary result succeeds unchanged or through a valid repair, run the second required secondary benchmark if assigned. Otherwise write not_assigned, blocked, or skipped.

    Completion criteria

    • The checkpoint always exists and records one benchmark comparison or the exact reason it was not run.

    • The scientific command runs when the primary target has valid comparable execution evidence under either execution_mode, including a partial component comparison; it does not run after not_executable or not_assessable.

  11. Run secondary target 3

    Runs asReproduction Engineer

    After the primary result succeeds unchanged or through a valid repair, run the third required secondary benchmark if assigned. Otherwise write not_assigned, blocked, or skipped.

    Completion criteria

    • The checkpoint always exists and records one benchmark comparison or the exact reason it was not run.

    • The scientific command runs when the primary target has valid comparable execution evidence under either execution_mode, including a partial component comparison; it does not run after not_executable or not_assessable.

  12. Package reproduction result

    Runs asReproduction Engineer

    Assemble the durable readiness, primary, repair, and secondary checkpoints into a compact comparison package and replay directory without rerunning analysis.

    Completion criteria

    • The reproduction package records or content-addresses the frozen analysis contract, case-bundle identities, execution profile, commands, 600-second timeout, environment, inputs, outputs, logs, hashes, repairs, numerical diagnostics, component comparisons, and replay instructions.

    • Every completed checkpoint remains durable, and a later failure cannot erase earlier successful evidence.

    • A successful permitted wrapper records execution_mode technically_repaired_original. The routine outcome follows comparison_status and never follows the repair label.

    • comparison_status exact_reproduction or within_tolerance requires every comparable required estimate, interval endpoint, statistic, p-value, and sample component to pass its frozen rule, regardless of execution_mode.

    • A missing hard CPU limit is recorded as cpu_limit_best_effort and cannot produce not_executable when execution_allowed is true. Required network isolation, source protection, and explicit safety or authority controls remain gates.

    • External release remains disabled unless separately approved.

    • comparison_status is partial_reproduction when valid execution yields a mixed component result, including a matching estimate, statistic, and p-value with a failing reported interval. A paper-omitted model sample is recorded as not_reported and never silently treated as a match.

    Outcomes

    • exact_reproduction

      Every comparable required component reproduces exactly under frozen rules, using either execution_mode.

    • within_tolerance

      Every comparable required component passes its frozen tolerance rule, using either execution_mode.

    • partial_reproduction

      Valid execution produced a mixed component comparison, such as matching estimate, statistic, and p-value with a failing interval.

    • numerical_mismatch

      The executable analysis remains materially inconsistent after the bounded remediation pass and does not qualify as a partial component match.

    • not_executable

      A missing runtime or dependency, unreadable required input, source or data hash mismatch, permission failure, required isolation failure, or inability to create a protected attempt-local working copy prevents every scientific attempt. Best-effort CPU enforcement alone cannot emit this outcome.

    • not_assessable

      The target cannot be assessed without an unsupported substantive choice.

Sub-routines

Research Methods And Evidence

Purpose

Answer no more than three targeted questions raised by the frozen claim, reproduction result, or supplied-document conflicts, deep-map no more than three authorized sources through checkpointed Silk-aware workers, and package exact passages and manuscript consequences. Sensitivity results may trigger a bounded follow-up only during final verification.

Status: Production | Updated: 2026-08-29

Prerequisites

  • Call browser_silk_load before opening each mapped literature site used by the active step.
  • Verify that each source is open access or otherwise authorized by the active audit before opening full text.
  • Reuse the mapped funnels below; none of these read-only literature funnels requires authentication.
  • Saved Browser Silk funnels available to the three bounded source slots:
    • pmc.ncbi.nlm.nih.gov / open_and_extract_article requires pmcid and returns visible full text, headings, tables, figures, supplements, and references.
    • arxiv.org / open_and_extract_full_text requires arxiv_id and returns article text and structure.
    • journals.sagepub.com / open_and_extract_article requires doi and returns visible article text and structure.
    • www.cambridge.org / open_and_extract_article requires article_slug and returns visible article text and structure.
    • rameliaz.github.io / open_pdf requires document_path; the local PDF parser supplies page and text locations when browser text is unavailable.
    • www.bmj.com / open_and_extract_article requires article_path and returns visible article text and structure.
    • www.nber.org / open_working_paper_pdf requires document_path; the local PDF parser supplies page and text locations when browser text is unavailable.
    • osf.io / project inventory verifies project identity and authorized download records; registered local copies supply full text when available.

Steps

  1. Freeze targeted research questions

    Runs asResearcher

    Freeze no more than three evidence questions that arise from the frozen claim, reproduction comparison, or verified supplied-document conflicts. Do not start a broad literature review or require a completed sensitivity map.

    Completion criteria

    • The brief contains no more than three questions, and each question names the finding it can confirm, contradict, or qualify.

    • Each question has source classes, search concepts, inclusion and comparison rules, authorization limits, and a stopping rule.

    • The initial brief does not depend on sensitivity results; any later sensitivity-triggered source check is deferred to final verification.

  2. Search, screen, and assign source slots

    Runs asLiterature Searcher

    Use toone-scholar and authorized sources to answer the frozen questions, metadata-screen broadly, deduplicate versions, and assign no more than three deep-source slots. Use Toone's persistent browser and complete saved Silk flows for recurring interactive sources.

    Completion criteria

    • Every query, source, timestamp, version, access status, screening decision, and deduplication key is recorded.

    • No more than three deep-source slots are assigned; extra candidates remain metadata-only.

    • Each slot names its authorized access route, exact evidence target, and mapped Silk funnel when applicable.

  3. Map selected source 1

    Runs asSource Extractor

    Process only source slot 1. Load the mapped domain before navigation, replay the complete saved Silk funnel, and write one independent checkpoint; if no source is assigned, write not_assigned without browsing.

    Completion criteria

    • Source slot 1 always produces one checkpoint with status mapped, limited, inaccessible, failed, or not_assigned.

    • This step is one independent Source Extractor execution in the three-way fan-out. It processes only its assigned slot and never takes another slot's source.

    • Mapped evidence includes source identity, version, exact passage and location, context, comparability, supported or contradicted statement, manuscript consequence, access status, and replay provenance.

    • Read maximum_seconds_per_full_text from execution-policy. Complete any started Silk funnel, then converge immediately to the checkpoint without opening another source or beginning optional extraction work when the limit is reached.

    • For selector_not_found, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. If it still fails, preserve the last failure evidence in the checkpoint; never switch to manual browsing.

  4. Map selected source 2

    Runs asSource Extractor

    Process only source slot 2. Load the mapped domain before navigation, replay the complete saved Silk funnel, and write one independent checkpoint; if no source is assigned, write not_assigned without browsing.

    Completion criteria

    • Source slot 2 always produces one checkpoint with status mapped, limited, inaccessible, failed, or not_assigned.

    • This step is one independent Source Extractor execution in the three-way fan-out. It processes only its assigned slot and never takes another slot's source.

    • Mapped evidence includes source identity, version, exact passage and location, context, comparability, supported or contradicted statement, manuscript consequence, access status, and replay provenance.

    • Read maximum_seconds_per_full_text from execution-policy. Complete any started Silk funnel, then converge immediately to the checkpoint without opening another source or beginning optional extraction work when the limit is reached.

    • For selector_not_found, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. If it still fails, preserve the last failure evidence in the checkpoint; never switch to manual browsing.

  5. Map selected source 3

    Runs asSource Extractor

    Process only source slot 3. Load the mapped domain before navigation, replay the complete saved Silk funnel, and write one independent checkpoint; if no source is assigned, write not_assigned without browsing.

    Completion criteria

    • Source slot 3 always produces one checkpoint with status mapped, limited, inaccessible, failed, or not_assigned.

    • This step is one independent Source Extractor execution in the three-way fan-out. It processes only its assigned slot and never takes another slot's source.

    • Mapped evidence includes source identity, version, exact passage and location, context, comparability, supported or contradicted statement, manuscript consequence, access status, and replay provenance.

    • Read maximum_seconds_per_full_text from execution-policy. Complete any started Silk funnel, then converge immediately to the checkpoint without opening another source or beginning optional extraction work when the limit is reached.

    • For selector_not_found, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. If it still fails, preserve the last failure evidence in the checkpoint; never switch to manual browsing.

  6. Assemble targeted evidence

    Runs asResearcher

    Join the three source checkpoints into a targeted evidence package, preserving metadata-only candidates, access limits, exact passages, comparability judgments, and concrete manuscript consequences.

    Completion criteria

    • Every retained evidence item has an exact passage and location, the statement it supports or contradicts, a comparability assessment, verification state, and proposed manuscript consequence.

    • The package joins all three independent source checkpoints, including limited, inaccessible, failed, and not_assigned slots, without serializing or rerunning a sibling slot.

    • The package answers no more than three questions, contains no more than three deeply processed sources, distinguishes inaccessible and unassigned slots, and exposes remaining gaps.

    • Evidence used in a verdict is marked for independent reopening during verification; a second complete extraction review is not run here.

    • Write an immutable run-scoped evidence snapshot alongside the selected per-run evidence package; no prior-run replacement read is required.

    Outcomes

    • evidence_ready

      Required targeted evidence is available and supported without residual gaps.

    • evidence_ready_with_limits

      The targeted search completed with visible nonblocking gaps.

    • evidence_not_assessable

      A required evidence question cannot be answered from authorized accessible sources.

    • operationally_blocked

      Authority or access policy prevents required research.

Sub-routines

Stress Test Analysis

Purpose

Test three bounded families of scientifically defensible alternatives that preserve the frozen estimand, approve no more than five result-blind candidate cards, execute them through the shared local runner with bounded commands, and report one interpretable sensitivity result.

Status: Production | Updated: 2026-08-31

Steps

  1. Freeze sensitivity contract

    Runs asRobustness Lead

    Preflight the supplied case bundle and shared local execution capability, confirm that reproduction may continue, and freeze the estimand, conclusion rules, admissibility rules, candidate cap, and stopping rules before alternative results exist.

    Completion criteria

    • The supplied case bundle resolves the frozen source and data hashes, and the runner can create an attempt-local working copy with the required runtime, dependencies, registered inputs, source protection, and 600-second timeout.

    • Verified network or DNS denial is not required. Memory and storage are operational defaults and unavailable exact enforcement does not block sensitivity work.

    • The estimand fixes population, treatment or exposure, comparator, outcome, time, effect, observation unit, direction, practical threshold, and conclusion categories before candidate execution.

    • Emit sensitivity_ready when the baseline is reproduced or explicitly authorized and the local execution preflight passes; emit sensitivity_blocked only for a scientific, source/data, runtime, dependency, permission, or source-protection blocker.

    • The candidate pool is capped at five and each candidate changes one primary analytical choice.

    Outcomes

    • sensitivity_ready

      The frozen contract and shared execution capability are ready for bounded sensitivity design and execution.

    • sensitivity_blocked

      A scientific, source/data, or execution-capability blocker prevents sensitivity work.

  2. Design data and measurement candidates

    Runs asData Choices Analyst

    Propose no more than two high-value candidates covering sample inclusion, missingness, outcome or exposure definitions, and transformations while preserving the frozen estimand.

    Completion criteria

    • The lane returns at most two candidate cards, each with one primary change, scientific rationale, supplied support, unchanged estimand elements, expected sample consequence, classification, diagnostics, and compatibility constraints.

    • Candidate cards contain no alternative-analysis outcomes.

  3. Design model specification candidates

    Runs asModel And Covariate Analyst

    Propose no more than two high-value candidates covering covariates, functional form, weighting, or interactions while preserving the frozen estimand.

    Completion criteria

    • The lane returns at most two candidate cards, each with one primary change, scientific rationale, supplied support, unchanged estimand elements, expected sample consequence, classification, diagnostics, and compatibility constraints.

    • Post-treatment, collider, unavailable, and meaning-changing specifications are excluded before results exist.

  4. Design inference candidates

    Runs asInference Analyst

    Propose no more than two high-value candidates covering standard errors, clustering, multiplicity, or confidence intervals under the frozen design.

    Completion criteria

    • The lane returns at most two candidate cards, each with one primary change, scientific rationale, supplied support, unchanged estimand elements, classification, diagnostics, and compatibility constraints.

    • Inference choices identify their design basis and contain no alternative-analysis outcomes.

  5. Rank and freeze sensitivity candidates

    Runs asRobustness Lead

    Join the three lanes, remove duplicates and incompatible combinations, and freeze a ranked pool of three to five candidate cards before results are available.

    Completion criteria

    • The frozen pool contains three to five candidates when that many defensible options exist, with stable IDs, one primary change each, rationale, support, unchanged estimand elements, expected sample consequence, type, cost, and no result fields.

    • Excluded and unselected options retain concise reasons, and the selected pool covers more than one decision lane when scientifically possible.

  6. Run primary blind admissibility review

    Runs asBlind Method Reviewer Alpha

    Judge each candidate from results-blind cards under the frozen admissibility rules and identify whether a second review is required for a dispute or high-impact judgment.

    Completion criteria

    • Every candidate receives accept, supplementary, reject, or unresolved with criterion-level reasons and no result exposure.

    • Emit second_review_required only for a disagreement risk or high-impact admissibility judgment; otherwise emit primary_review_sufficient.

    Outcomes

    • primary_review_sufficient

      Every candidate can be resolved from the primary blind review.

    • second_review_required

      At least one disputed or high-impact judgment needs a second blind review.

  7. Record conditional second blind review

    Runs asBlind Method Reviewer Beta

    Always write the second-review artifact. Review only candidates flagged as disputed or high-impact by the primary reviewer; when none are flagged, emit not_required without repeating the primary review.

    Completion criteria

    • The artifact always exists.

    • Flagged candidates receive an independent criterion-level blind judgment with no result exposure.

    • Every unflagged candidate is marked not_required with the primary-review reference; no substantive second review is performed for it.

  8. Resolve candidate admissibility

    Runs asRobustness Lead

    Apply the frozen disagreement rule once and produce the final set of executable sensitivity candidates without reconsidering admissibility later.

    Completion criteria

    • No more than five accepted or supplementary candidates are executable, and every rejected or unresolved candidate retains its review evidence.

    • Each approved candidate preserves the estimand and has one primary analytical change.

  9. Execute approved sensitivity analyses

    Runs asReproduction Engineer

    Use the shared runner to execute each approved candidate as an isolated worker, compare outputs under one result schema, and preserve failed or untested candidates.

    Completion criteria

    • Execution begins only after sensitivity_ready routing and after ranked candidates, blind-review-alpha, and the always-materialized blind-review-beta checkpoint have joined.

    • Each approved candidate runs against the frozen source and data hashes through the shared runner from a separate attempt-local working directory and has a command, 600-second timeout, inputs, outputs, logs, diagnostics, estimate, interval, p-value, sample size, status, and hashes.

    • Verified inputs remain unchanged. Host network policy and operational memory or storage defaults are recorded but are not scientific-validity gates.

    • Candidate execution is capped at five; failed, timed-out, and untested candidates remain visible without open-ended repair.

  10. Build sensitivity result

    Runs asRobustness Lead

    Summarize stability, fragility, coverage, and the smallest defensible conclusion change from approved candidate runs.

    Completion criteria

    • The result states how many approved analyses preserve direction and statistical support, the estimate and interval range, any valid direction reversal, the smallest defensible conclusion-changing choice, and whether changes arise from estimate size, uncertainty, or sample composition.

    • Tested and untested dimensions, failed candidates, coverage, and limitations are explicit; significance counts alone are not used as the conclusion.

    • Write an immutable run-scoped sensitivity snapshot alongside the selected per-run result; no prior-run replacement read is required.

    Outcomes

    • robustness_ready

      Required sensitivity checks completed without residual gaps.

    • robustness_ready_with_limits

      Sensitivity checks completed with visible nonblocking gaps.

    • analysis_not_assessable

      A blocking scientific or executable gap prevents a defensible sensitivity conclusion.

    • operationally_blocked

      Authority or unsafe execution prevents required work.

Sub-routines

Verify And Adjudicate Findings

Purpose

Independently rerun the primary and verdict-relevant sensitivity results from the same frozen analysis contract and source and data hashes in a separate local working directory, then verify decisive source evidence and adjudicate manuscript changes.

Status: Production | Updated: 2026-08-31

Steps

  1. Freeze important verification targets

    Runs asVerification Lead

    Select only verdict-relevant reproduction, sensitivity, document, and source findings, freeze their evidence identities, and gate statistical verification against the executable contract and reproduction execution identities.

    Completion criteria

    • Every selected finding has one precise statement, claim and estimand links, importance reason, allowed verification method, source or run identity, hashes, and required completion rule.

    • The packet freezes the primary_target and analysis_contract from frozen-study-case, including source and data hashes, exact expressions, coefficient, orientation, and comparison rules.

    • The supplied case bundle resolves the frozen source and data hashes, and the runner can create a separate attempt-local working directory with the required runtime, dependencies, source protection, and 600-second timeout.

    • Verified network or DNS denial is not required. Memory and storage are operational defaults, and unavailable exact enforcement is recorded without blocking a runnable verification command.

    • Previously decided candidate admissibility is preserved, and external release remains disabled unless the audit request contains separate approval.

  2. Independently rerun statistical evidence

    Runs asStatistical Evidence Verifier

    Perform an independent execution of the frozen analysis_contract and only verdict-relevant sensitivities from a separate local working directory using the same source and data hashes.

    Completion criteria

    • Preflight proves access to the case bundle, frozen source and data hashes, required runtime and dependencies, registered inputs, a separate attempt-local working directory, source protection, and a 600-second timeout.

    • The verification run executes the exact frozen formula, filter, transformations, factor levels, treatment or exposure structure, coefficient, orientation, missing-data behavior, interval method, and inferential method.

    • The primary run records command, working directory, timeout, environment, inputs, outputs, logs and hashes; eligible rows, model rows, nobs(), excluded rows and na.action evidence; coefficient, standard error, statistic, degrees of freedom, p-value, interval endpoints and method, confidence level, and orientation.

    • Estimate, interval endpoints, statistic, p-value, and sample are compared independently with the reported and reproduction values.

    • A different runner hash, workspace path, mount identity, or available operational control does not fail verification when the frozen scientific contract, source and data hashes, dependencies, source immutability, and recorded environment are equivalent for the claim.

    • Verified network or DNS denial and exact memory or storage enforcement are not prerequisites. A scientific-contract or source/data hash mismatch remains a precise blocker and cannot support audit_ready.

  3. Independently reopen source evidence

    Runs asSource Evidence Verifier

    Reopen only supplied-document conflicts and external passages used in the verdict or recommendations, using authorized local sources, toone-scholar, or the saved Browser Silk funnel for each mapped domain. Perform full licence adjudication only when audit-request says the package will be publicly redistributed.

    Completion criteria

    • Every verdict-relevant source statement has an independent identity, version, exact-passage, context, correction or retraction, and comparability check or a visible access limitation.

    • Internal audits record access authority and provenance. Full licence verification is required only when public redistribution is intended or authorized.

    • For a selector_not_found failure, perform at most one in-attempt Silk heal, retry the affected flow step once, and push a healed map. Preserve the last failure evidence for runtime attempt handling; never switch to manual browsing.

    • No snippet-only statement passes verification and sources unrelated to the verdict are not reopened.

    • One bounded locator correction may be made; unresolved source problems remain limitations.

    • If a sensitivity result raises a new verdict-critical methods question, perform at most one bounded follow-up using the same source limits and record it in source-verification.

  4. Adjudicate findings and manuscript changes

    Runs asScientific Adjudicator

    Join the independent statistical and source checks, assign final finding status and allowed wording, and write specific manuscript changes without introducing new analyses or source claims.

    Completion criteria

    • Every retained finding is verified, verified with limits, rejected, incomplete, or disputed, with severity, confidence, exact evidence, allowed wording, and no duplicated admissibility review.

    • Each manuscript recommendation names the affected section or statement, the exact requested change, its evidence, urgency, and whether it is required or optional.

    • audit_ready requires an independent rerun from a separate local working directory with the same frozen scientific contract and source and data hashes, complete required numerical comparisons, protected source inputs, recorded environment details, and no unresolved scientific blocker.

    • Different runner, workspace, mount, network-control, memory-control, or storage-control identities do not block audit_ready by themselves.

    • Weak nonblocking findings are dropped; operationally_blocked is reserved for missing authority or an unsafe or impossible required verification.

    Outcomes

    • audit_ready

      All required findings are independently verified and ready for reporting.

    • verification_incomplete

      The audit can report verified work with explicit unresolved limitations.

    • operationally_blocked

      Required verification cannot proceed safely or with authority.

Agents

Audit Report Builder

Renders verified canonical objects into consistent HTML, PDF, JSON, and replay-package outputs.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read only frozen verified canonical objects and approved wording. Write project://peer2paper/audits/{runId}/delivery/audit.html, audit.pdf, audit.json, replay/, and manifest.json.
  • Render reproduction, robustness, and evidence-alignment statuses separately from one pinned template version and preserve exact numbers, order, links, limitations, and hashes across formats.
  • NEVER add a scientific finding, promote rejected material, or rewrite a verified conclusion beyond its allowed wording.

Release Quality Reviewer

Checks completeness, consistency, provenance, replay, privacy, accessibility, licensing, and release boundaries.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read the rendered audit package and all declared manifests. Write project://peer2paper/audits/{runId}/delivery/release-gates.json.
  • Check numerical consistency, evidence links, replay commands, expected hashes, language rules, accessibility, privacy, redaction, licences, and access restrictions.
  • NEVER repair scientific content during quality review or authorize external release beyond the dispatched release policy.

Claim Mapper

Extracts candidate claims and connects the selected claim to exact reported results and source locations.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read paper documents and the artifact registry. Write project://peer2paper/audits/{runId}/study-case/claim-records.json and selected-claim.json.
  • Separate measured results from interpretation and record population, exposure or treatment, comparator, outcome, timepoint, effect measure, direction, scope, and reported-result links.
  • NEVER finalize an ambiguous target claim or fabricate a missing source location; emit a typed unresolved decision instead.

Variable Mapper

Connects paper concepts, dataset fields, and code references while preserving units, timepoints, roles, and derivations.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read paper documents, dataset maps, code project, and selected claim. Write project://peer2paper/audits/{runId}/study-case/variable-maps.json and evidence-graph.json.
  • Preserve raw names and label normalized and derived meanings with transformation rules, confidence, provenance, and review state.
  • NEVER guess high-impact mappings, missing-value meanings, observation units, nesting, or scientifically meaningful recodes.

Peer2Paper Orchestrator

Orchestrates the Peer2Paper agentic workflow from intake through completion, coordinating work, tracking dependencies and outcomes, and escalating blockers or approval-sensitive actions without exceeding delegated authority.

  • Workflow orchestration
  • Task coordination
  • Dependency and status tracking
  • Blocker escalation
  • Outcome reporting

Reproduction Engineer

Builds clean R or Python execution recipes, reruns supplied analyses, compares results, and tests minimal technical repairs.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read project://peer2paper/audits/{runId}/study-case/frozen-study-case.json and its referenced execution copies. Write declared files under project://peer2paper/audits/{runId}/reproduction/.
  • Use deterministic repository scripts and isolated execution facilities supplied to the step for environment detection, clean runs, numeric comparison, hashing, and replay checks.
  • Run supplied analysis twice from separate clean states when executable. NEVER call a scientific change an exact reproduction or enable network access without explicit permission.
  • Write project://peer2paper/audits/{runId}/reproduction/reproduction-package.json with one allowed reproduction status and linked run evidence.

Data Choices Analyst

Proposes defensible sample, missingness, variable-construction, and population choices without seeing candidate outcomes.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read frozen study-case mappings, the estimand, and conclusion rules. Write separate lane files under project://peer2paper/audits/{runId}/robustness/decision-lanes/data/.
  • Propose sample, exclusion, missingness, variable-construction, and population options with scientific rationale and compatibility constraints before execution.
  • NEVER choose a method because of an observed estimate, interval, direction, or p-value.

Inference Analyst

Proposes valid clustering, weighting, uncertainty, and multiplicity choices for the frozen design and estimand.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read the frozen study case, study design, estimand, and execution constraints. Write inference lane files under project://peer2paper/audits/{runId}/robustness/decision-lanes/inference/.
  • Propose clustering, weights, standard errors, uncertainty procedures, and multiplicity rules with diagnostic requirements and compatibility constraints.
  • NEVER approve an inference method after seeing whether it changes support.

Model And Covariate Analyst

Proposes design-compatible statistical models and adjustment sets while preserving the frozen estimand.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read the frozen study case, study design, estimand, and variable map. Write model and covariate lane files under project://peer2paper/audits/{runId}/robustness/decision-lanes/modeling/.
  • Specify model families, functional forms, interactions, adjustment sets, and compatibility constraints with explicit rationale.
  • NEVER treat every control combination as valid or change the scientific question silently.

Robustness Lead

Freezes the estimand and conclusion rules, integrates analysis-decision lanes, registers candidates, and produces the robustness map.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read the frozen study case and reproduction package. Write project://peer2paper/audits/{runId}/robustness/estimand.json, conclusion-rules.json, analysis-decision-space.json, candidate-registry.json, and robustness-map.json.
  • Register candidate rationale and decision combinations before result execution, merge only compatible choices, and preserve invalid, failed, and untested regions.
  • NEVER reveal candidate results to method reviewers or claim universal robustness from bounded coverage.

Blind Method Reviewer Alpha

Independently judges candidate-analysis validity from results-blind packets.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read only results-blind packets and declared design materials. Write project://peer2paper/audits/{runId}/verification/method-review-alpha.json.
  • Judge same-estimand fit, design fit, exclusions, controls, missingness, inference, rationale, and reproducibility using declared criteria.
  • NEVER inspect hidden candidate outcomes or coordinate a verdict with the other method reviewer.

Blind Method Reviewer Beta

Provides a second independent judgment of candidate-analysis validity from results-blind packets.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read only results-blind packets and declared design materials. Write project://peer2paper/audits/{runId}/verification/method-review-beta.json.
  • Judge same-estimand fit, design fit, exclusions, controls, missingness, inference, rationale, and reproducibility using declared criteria.
  • NEVER inspect hidden candidate outcomes or coordinate a verdict with the other method reviewer.

Scientific Adjudicator

Resolves verification disagreements and locks final validity, severity, confidence, and allowed wording.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read independent method reviews, statistical checks, source checks, conflict records, and frozen conclusion rules. Write project://peer2paper/audits/{runId}/verification/final-adjudication.json.
  • Apply declared resolution rules, preserve dissent and limitations, and return typed correction, evidence, ready, incomplete, or disputed outcomes.
  • NEVER introduce a new analysis, source claim, or unsupported accusation during adjudication.

Source Evidence Verifier

Independently opens sources and checks whether exact passages support proposed literature findings.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Use toone-scholar and Toone's persistent browser with reusable Silk flows to verify authorized source versions, full text, corrections, retractions, and exact passages. Write project://peer2paper/audits/{runId}/verification/source-checks.json.
  • Check bibliographic identity, context, population, measures, methods, results, licence, and comparability independently of the source extractor.
  • NEVER accept a citation from a snippet, bypass access controls, or broaden the claim beyond the verified passage.
  • MCP servers browser, browser-silk, toone-scholar

Statistical Evidence Verifier

Independently clean-reruns important analyses and validates statistical findings against frozen rules.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read approved candidate definitions, execution recipes, frozen inputs, and native run outputs. Write project://peer2paper/audits/{runId}/verification/statistical-checks.json and independent-runs/.
  • Use supplied deterministic scripts and isolated execution facilities to rerun important analyses and verify IDs, samples, variables, models, seeds, diagnostics, estimates, intervals, and classifications.
  • NEVER reuse a proposing analyst's unsupported conclusion or alter the frozen candidate during verification.

Verification Lead

Freezes verification inputs, reconciles independent checks, preserves disagreements, and issues final adjudications.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Read frozen reproduction, robustness, and literature packages. Write project://peer2paper/audits/{runId}/verification/verification-packet.json, resolved-finding-graph.json, and adjudications.json.
  • Keep proposed methods separate from observed results until method review is complete, preserve review disagreements and dependency chains, and assign validity, severity, confidence, and allowed wording separately.
  • NEVER erase rejected findings from history, convert uncertainty into confidence, or adjudicate without the required independent checks.

Literature Searcher

Runs reproducible scholarly searches and records complete source, query, access, and screening logs.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Use toone-scholar for Europe PMC, PubMed, citation, preprint, and open-access lookup. Use browser and browser-silk for recurring web sources through Toone's persistent browser session and saved Silk flows.
  • Read the frozen research brief. Write project://peer2paper/audits/{runId}/research/search-log.json and source-registry.json.
  • Annotate and save reusable Silk flows for recurring browser sources while working. NEVER bypass paywalls, store credentials, or cite snippets as evidence.
  • MCP servers browser, browser-silk, toone-scholar

Researcher

Conducts organization-wide research, gathers and verifies information from available internal integrations and web sources, and delivers evidence-based findings while respecting access and approval boundaries.

  • Organizational research
  • Web research
  • Source verification
  • Research synthesis
  • MCP servers browser, browser-silk

Source Extractor

Maps accessible full texts into structured study objects and exact evidence passages with provenance.

  • Read only routine-projected inputs and artefacts under project://peer2paper/audits/{runId}/. Write only the declared output artefacts for the active step. NEVER overwrite raw submissions, invent missing evidence, or treat prose as the canonical result.
  • Use toone-scholar and Toone's persistent browser with reusable Silk flows to open authorized full texts and supplements. Write files under project://peer2paper/audits/{runId}/research/full-text/ and extracted-study-objects.json.
  • Preserve raw, normalized, and derived layers; attach stable section, page, table, figure, and passage locations to every extracted field.
  • NEVER guess unreadable content, claim equivalence between unlike measures, or treat abstract-only evidence as full-text evidence.
  • MCP servers browser, browser-silk, toone-scholar

Requirements

Included
  • Audit Report Builder
  • Release Quality Reviewer
  • Claim Mapper
  • Variable Mapper
  • Peer2Paper Orchestrator
  • Reproduction Engineer
  • Data Choices Analyst
  • Inference Analyst
  • Model And Covariate Analyst
  • Robustness Lead
  • Blind Method Reviewer Alpha
  • Blind Method Reviewer Beta
  • Scientific Adjudicator
  • Source Evidence Verifier
  • Statistical Evidence Verifier
  • Verification Lead
  • Literature Searcher
  • Researcher
  • Source Extractor
  • Routine-owned versioned schemas for canonical findings, claims, actions, limitations, and frozen audit objects.
  • Project-owned pinned template used for the researcher-facing HTML report.
  • Routine-owned pinned JSON Schema for the machine-readable audit.
  • Project-owned versioned scientific execution runner used by reproduction, robustness, and independent statistical verification. A missing runner file blocks execution.
You provide
  • Local directory containing the paper or draft, data, code, environment files, supplements, data dictionaries, and authorized supporting sources.
  • Required audit scope containing one exact target claim and location, permissions, privacy classification, retention rules, public-redistribution intent, and external-search authorization.
  • Required scientific rules for estimands, tolerances, admissibility, comparability, severity, disagreement, and allowed defaults. Supply an empty JSON object only when the routine's recorded defaults are intended.
  • Execution rules for runtime, attempt-local working directories, host network policy, operational memory and storage defaults, a 600-second command timeout, source immutability, source access, and internal-only release.
  • Required only after partial_reproduction or numerical_mismatch. Supply a JSON decision of continue or stop, with a reason and the authorizing person or role.
Produces
  • Target claim and relevant document map (json)
  • Target data and variable map (json)
  • Target code and dependency map (json)
  • Verified supplied-document consistency findings (json)
  • Compact frozen study case contract (json)
  • Frozen reproduction contract and inputs (json)
  • Execution readiness (json)
  • Execution profile attempt (json)
  • Untouched primary run checkpoint (json)
  • Primary numerical comparison (json)
  • Frozen repair candidates (json)
  • Repair attempt 1 checkpoint (json)
  • Repair attempt 2 checkpoint (json)
  • Repair attempt 3 checkpoint (json)
  • Secondary target 1 checkpoint (json)
  • Secondary target 2 checkpoint (json)
  • Secondary target 3 checkpoint (json)
  • Per-target reproduction evidence (directory)
  • Reproduction package (json)
  • Frozen estimand and conclusion rules (json)
  • Data and measurement candidates (json)
  • Model specification candidates (json)
  • Inference candidates (json)
  • Ranked sensitivity candidate pool (json)
  • Primary blind admissibility review (json)
  • Conditional second blind review (json)
  • Approved sensitivity candidates (json)
  • Isolated sensitivity runs (json)
  • Sensitivity result and coverage map (json)
  • Sensitivity version history (directory)
  • Targeted research questions (json)
  • Screened source shortlist and slot plan (json)
  • Source evidence checkpoint 1 (json)
  • Source evidence checkpoint 2 (json)
  • Source evidence checkpoint 3 (json)
  • Assembled targeted full-text evidence (json)
  • Targeted literature evidence package (json)
  • Targeted evidence version history (directory)
  • Frozen important verification targets (json)
  • Independent statistical rerun (json)
  • Independent source and document verification (json)
  • Verified finding package (json)
  • Concrete manuscript recommendations (json)
  • Canonical audit JSON (json)
  • Replay package (directory)
  • Researcher audit report HTML (html)
  • Scoped release checks (json)
  • Audit package manifest (json)

MCP servers

  • browser
  • browser-silk
  • toone-scholar