Evidence over claims. Assurance over automation.

Active empirical study

Do structured workflow scaffolds reduce unsupported claims?

The study tests whether procedural controls for provenance, uncertainty, disconfirmation, and final claim auditing improve source support in AI-assisted research without winning by making answers empty, vague, or excessively cautious.

Research question

Is unsupported output partly a workflow reliability problem?

The proposal treats unsupported claims as errors that can survive several stages: source review, synthesis, inference, drafting, and editing. The intervention is procedural rather than rhetorical.

The study does not assume that extra structure works. It tests whether the controls improve actual claim support or merely produce more convincing process artifacts.

Controls under test

  • Source-to-claim binding and provenance
  • Explicit uncertainty and inference labels
  • Counterevidence or disconfirmation checks
  • Final claim audit with removal or downgrade of weak claims

Experimental design

Four conditions separate formatting from evidence discipline

The official measurement unit is the substantive claim extracted from the visible final answer. Model-generated scaffold tables are treated as model self-report, not as the authoritative claim registry.

01

Baseline

Ordinary source-packet research without an added evidence-control scaffold.

02

Format only

Visible claim structure without provenance, audit, or disconfirmation discipline.

03

Provenance scaffold

Claims identify source, inference, uncertainty, or lack of support.

04

Full scaffold

Provenance, disconfirmation, uncertainty labels, final audit, and removal or downgrade of weak claims.

Measures

The study must not win by saying less

Support measures are paired with usefulness, coverage, false-caution, inspectability, and cost measures.

Primary measures

  • Unsupported-claim rate
  • Source-attribution accuracy
  • Overconfidence
  • Counterevidence recall
  • Claims corrected or removed during audit

Guardrail measures

  • Usefulness and coverage
  • False caution
  • Inspectability
  • Reviewer time-to-problem
  • Tokens, model calls, and runtime

Claim verdicts

  • Supported
  • Partially supported
  • Reasonable inference
  • Unsupported
  • Contradicted or not checkable

Measurement architecture

Generation, candidate evidence, and support judgment remain separate.

The Research Scaffold Harness runs experimental conditions and preserves traceable run artifacts. Evidence Bundler nominates candidate passages and seals evidence bundles. Claim Audit Lab applies controlled support verdicts for later human calibration.

The harness does not verify claims. Candidate passage coverage is not a support verdict. CAL is a measurement channel, not automatic ground truth.

Integrity requirements before a result

  • The official answer surface is defined
  • Zero-answer and malformed outputs are dispositioned
  • Claim extraction is calibrated against human coding
  • Support verdicts are human-calibrated
  • Coverage and false-caution tradeoffs are measured

Current interpretation

No valid scaffold-effect conclusion is claimed yet.

The central causal question remains open. The study will not report a confirmatory difference until the answer-surface protocol, claim extraction, candidate evidence bundles, support judgments, and human calibration are sufficiently settled.

A positive, negative, mixed, or false-caution result would all be informative. The objective is to identify whether the bundled workflow creates real reliability gains and, if so, which components are necessary.

Publication restraint: completing the apparatus or producing a matrix of outputs does not establish that the intervention works. Method progress and causal evidence are different claims.

Need a research question treated with the same controls?

The Evidence Intelligence Research Sprint applies a bounded protocol, source register, competing interpretations, explicit uncertainty, and a decision-focused briefing to a client question.