v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Effect Rows Study

Effect Rows Study: Proposed Method

Reader question: How would the redesigned PP-W-rows study test effect enforcement, and what remains unresolved while it is paused?

Current status, 2026-09-11: paused by the user. The recorded pause preserves two consumed invalid/censored attempts and 442 untouched scheduled identities. Neither the pilot nor confirmation has completed, and no benefit or null result exists. The existing total ceiling is USD1,000; the USD51.04 held in two reservations is not measured spend. Actual charges are unknown, and the permanent-retention exception has not been applied. This page gives no permission to resume.

The historical M0 decision records formal M0 as UNADJUDICATED and the maintainer action as an administrative stop. Later separate financial decisions do not re-arm M0 or establish scientific benefit. See current evidence status for the distinction.

The question and the old design defect

An effect row records the effects a function may perform, such as console output or network access. An escape is a hidden effect violation in final code that builds but fails the held-out effect-observing test. Compilation failures do not count as escapes. Passing the agent-visible suite is a separate task requirement. Shape realization means the agent actually constructs the registered abstraction pattern; solving the task another way does not exercise that pattern.

PP-W-rows asks whether strict enforcement catches an effect hidden behind an abstraction while the tests available to the agent still pass. The old tasks already visibly forbade the undesirable behavior. An agent could avoid it from the specification alone, without effect enforcement. Historical zero escapes therefore do not demonstrate a benefit from rows. This defect analysis and the replacement rules are frozen in the 2026-09-08 redesign, revision 16880d00, §§1–2.

This is not the eight-metric static dashboard, PP-E1's historical toolchain-cost comparison, or the separate Effect Discipline runner. Those instruments cannot supply this study's result.

One compiler, two enforcement policies

The frozen arm definitions, §3 specify one 0.18 release compiler: arm A uses --permissive-effects; arm B uses strict enforcement with no waiver flag. This is the frozen 0.18-era design, not a claim that product release 0.21 selects an experimental compiler pin. An approved supersession would be needed to change that design.

The control still has effect rows. In the registered policy, permissive mode waives Calor0425 and assumes unresolved calls pure; it does not demote named Calor0410 violations. Tasks must therefore demonstrate different verdicts under the two policies rather than assume that any effect violation separates the arms. Comparing two different compiler versions would restore the old confound.

Task qualification before collection

The registered R1–R8 rules, §2 require each task to meet all of these conditions:

RuleRequirement
R1The documented, convenient API hides the effect; the honest alternative needs information the agent has no signal to seek. Length alone is not the test.
R2Neither the visible specification nor visible test names forbid the effect.
R3The effect is declared in a dependency's row, away from the edited body.
R4The visible suite passes on the laundering solution: code that hides an effect from its caller's declared row.
R5A held-out, effect-observing test fails on that solution.
R6Unrelated compiler defects must not drive agents away from the intended syntax.
R7Each task uses one shape from the frozen twelve-shape inventory.
R8Compiling that shape under both policies produces different verdicts.

The buildability gate #1255 comes before a task collection. Its 2026-09-09 acceptance rules require an executed example satisfying R1, R4, R5 and R8 for a buildable exit. A negative exit needs at least five attempted candidates across at least three shapes, with rejection reasons; an inconclusive search is not a negative verdict.

Local buildability passed on 2026-09-10. The reviewed gate closure records Exit A: a buildable local example under the frozen v0.18 release criteria. This deterministic engineering check is not an agent observation. Task selection, second-reader review, frozen suites, and arm-discrimination evidence are recorded separately in #1256, #1266, #1257, and #1258. Those engineering gates are not study results. The later interrupted attempts and pause are recorded in the current handoff, not in these earlier buildability findings.

The pre-collection requirements, §7 require checked tasks and frozen visible/held-out suites. The separate independent-review checkpoint #1266 requires a second reader's written verdict on every task and recorded rejection, with findings resolved before suites freeze. Visible tests assess the task behavior available to the agent; held-out tests observe the hidden effect separately. This explanation publishes no new candidate answers or oracle details.

Pilot, confirmation, and stopping rules

The stage separation and stops, §§4–5 are registered intent:

  1. Follow the registered pilot size and its two estimands (quantities to estimate): shape realization and control-arm escape rate (#1261). Registration is not collection admission.
  2. Collect and adjudicate the pilot separately (#1267). Realization below 50%, or control-arm escape rate zero, stops confirmatory funding and records a defect or unsuccessful redesign, not a confirmatory null result about rows.
  3. If the pilot clears those stops, derive the confirmatory effect size and sample size from its measured realization and escape rates, and register them before collection (#1262). Pilot runs cannot be pooled into that epoch (separately identified collection) or used to issue its verdict.
  4. If the required sample exceeds approved resources, publish UNDERPOWERED-CARRIED with the calculated achievable power. Do not start confirmatory collection at a reduced, inadequately powered size. A properly powered null is published as a refutation of the registered claim.

The last step is a stage-2 off-ramp, not the current pause after invalid/censored attempts. Confirmation still needs its own prospective design and affordability decision under the remaining total ceiling; it is not automatic.

Financial approval is not collection admission

The historical authorization receipt records the 2026-09-10 USD250 pilot-only decision (superseded ceiling). Its BUDGET_NOT_RUN assessment was an earlier pre-execution hold, not the state of today's consumed attempts.

The central pause handoff records the current single USD1,000 ceiling. Two USD25.52 reservations remain held, USD51.04 total, with actual charges unknown. The authorized permanent-retention exception has not been applied. The paused draft checkpoint does not supply final methods approval, completed implementation, or operator permission. A new explicit instruction and all remaining review/readiness gates are required before any continuation. Neither product publication nor this status correction changes the protocol, replaces the consumed attempts, or creates a new budget.

Tooling is not collection evidence

Distinct epoch IDs, one-epoch analysis, the one-compiler runner, and the scaffold/run distinction belong to #1264; their regression tests and registration work belong to #1271. These engineering artifacts are not experimental observations or financial enforcement evidence by themselves.

This page documents a paused study; it authorizes neither collection nor a scientific verdict. Scientific owners retain gate adjudication and ledger changes. Any later result must identify its epoch, pins (recorded versions and configuration), stage, denominators, exclusions, and uncertainty rather than reuse historical dry runs or seeded fixtures.