v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Effect Discipline

Effect Discipline Runner

Reader question: What tasks does this runner execute, and what does its score establish?

This outcome-based runner is separate from the historical eight-metric dashboard and the PP-W-rows agent study. It supplies neither that study's results nor a measure of production incident prevention. Estimation-mode calculator signals are not executed task results.

The recorded outcome records a user-directed pause after two invalid/censored attempts, not a zero-effect result or a completed pilot. This product release does not authorize resuming research.

Tasks and execution

Task categoryIllustrative scenario, not a sourced incident
Flaky-test preventionA report reads the current time instead of receiving it as an input.
Security boundariesA parser makes an unexpected network request.
Side-effect transparencyA string utility also writes a log.
Cache safetyA price function reads a changing exchange rate.

The runner generates or retrieves cached code for each language, compiles it, and executes the task's test cases. Optional C# analyzers and pattern warnings provide additional diagnostics. Tests observe only their encoded behaviors, not every possible hidden effect.

Scoring rule

The scorer uses 50% test-pass fraction + 50% all-tests-pass indicator. Compilation failure scores zero; no tests also scores zero. The indicator is one only when a nonempty test set all passes. Both terms derive from the same tests: the field called “bug prevention” is not an independent determinism proof. Effect annotations and [Pure] earn no bonus; maintainability and pattern warnings do not enter the score.

Equal scores when both languages pass are a consequence of that rule, not an observed tie. A passing sample does not prove purity or general cache safety. See effect coverage and verification limits.

Run an explicit-input example

With .NET 10 and Calor installed, save this complete illustrative program as Price.calr in an empty directory:

Calor
§M{m001:PriceExample}
  §F{f001:Main:pub} () -> void
    §E{cw}
    §B{unitPrice:i32} 10
    §B{quantity:i32} 3
    §B{rate:i32} 2
    §P (* (* unitPrice quantity) rate)
Bash
calor run Price.calr

Expected output:

Plain Text
60

The rate is explicit rather than fetched. This is an arithmetic example, not a financial model, memoization proof, or measured benchmark result.

Inspect the runner without collection

From a repository checkout with its documented build prerequisites, including Z3 bootstrap, run:

Bash
dotnet run --project tests/Calor.Evaluation -- effect-discipline --provider mock --dry-run --sample 1

The terminal ends with Dry run - no API calls made. This preview neither generates code nor records an empirical result. Removing these safeguards can invoke a paid provider; this page authorizes no collection. Study decisions remain on the evidence status page.