Effect Discipline
Effect Discipline Runner
Reader question: What tasks does this runner execute, and what does its score establish?
This outcome-based runner is separate from the historical eight-metric dashboard and the PP-W-rows agent study. It supplies neither that study's results nor a measure of production incident prevention. Estimation-mode calculator signals are not executed task results.
The recorded outcome records a user-directed pause after two invalid/censored attempts, not a zero-effect result or a completed pilot. This product release does not authorize resuming research.
Tasks and execution
| Task category | Illustrative scenario, not a sourced incident |
|---|---|
| Flaky-test prevention | A report reads the current time instead of receiving it as an input. |
| Security boundaries | A parser makes an unexpected network request. |
| Side-effect transparency | A string utility also writes a log. |
| Cache safety | A price function reads a changing exchange rate. |
The runner generates or retrieves cached code for each language, compiles it, and executes the task's test cases. Optional C# analyzers and pattern warnings provide additional diagnostics. Tests observe only their encoded behaviors, not every possible hidden effect.
Scoring rule
The scorer
uses 50% test-pass fraction + 50% all-tests-pass indicator. Compilation
failure scores zero; no tests also scores zero. The indicator is one only
when a nonempty test set all passes. Both terms derive from the same tests:
the field called “bug prevention” is not an independent determinism proof.
Effect annotations and [Pure] earn no bonus; maintainability and pattern
warnings do not enter the score.
Equal scores when both languages pass are a consequence of that rule, not an observed tie. A passing sample does not prove purity or general cache safety. See effect coverage and verification limits.
Run an explicit-input example
With .NET 10 and Calor installed, save
this complete illustrative program as Price.calr in an empty directory:
§M{m001:PriceExample}
§F{f001:Main:pub} () -> void
§E{cw}
§B{unitPrice:i32} 10
§B{quantity:i32} 3
§B{rate:i32} 2
§P (* (* unitPrice quantity) rate)
calor run Price.calr
Expected output:
60
The rate is explicit rather than fetched. This is an arithmetic example, not a financial model, memoization proof, or measured benchmark result.
Inspect the runner without collection
From a repository checkout with its documented build prerequisites, including Z3 bootstrap, run:
dotnet run --project tests/Calor.Evaluation -- effect-discipline --provider mock --dry-run --sample 1
The terminal ends with Dry run - no API calls made. This preview neither
generates code nor records an empirical result. Removing these safeguards can
invoke a paid provider; this page authorizes no collection. Study decisions
remain on the evidence status page.