v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Results

Static Benchmark Snapshot

Evaluated across 217 programs with 8 metrics

Updated: c2a8816

Recorded source revision c2a8816d declares compiler v0.21.0; this is a pinned static-calculator run, not an agent-productivity measurement.

Corpus: tests/TestData/Benchmarks — 217 programs.

Method: Deterministic static-analysis calculators, 30 repetitions, 8 metrics.

The reports record the source revision before the release version bump, not an independently attested measurement binary. The 30 repetitions are deterministic, not independent samples. Source pairs are not all behaviorally equivalent; the reported intervals establish no measured language, agent-productivity, correctness, or safety advantage. See issue #1276.

Legacy Direction-Normalized Composite
1.32x
Above 1 favors Calor
Source Pairs
217
8 deterministic calculators
How to read this run: source pairs are not all behaviorally equivalent, so these ratios cannot establish a language advantage. The exporter reports singleton category intervals and pooled repeated-program intervals; neither establishes sampling uncertainty. These deterministic calculators are not coding-agent trials and are separate from the v0.12 release gates. Benchmark integrity follow-up.

Agent Refactoring Benchmark

Historical Claude Code study of the bundled refactoring tasks (rename, extract, inline, move, add contracts, change signature), not a measurement of current-release productivity.

Calor19/20 tasks
95%
C#19/20 tasks
95%
View category breakdown
Rename Symbol
2/3
3/3
Extract Method
4/4
4/4
Inline Function
3/3
3/3
Move Method
3/3
3/3
Add Contract
4/4
3/4
Change Signature
3/3
3/3
Majority voting (2/3 runs) with compilation + Z3 verification

Historical majority-vote refactoring study (2 of 3 runs). Corpus: tests/E2E/agent-tasks; recorded source declares compiler v0.2.3. Exact measurement binary not independently recorded.

580e189

Static Scores by Metric

Alphabetical metric order. Ratios are direction-normalized: above 1 favors Calor and below 1 favors C#. Lower-is-better metrics invert their raw score ratio. Neither means a language is better.

Comprehension static score

1.84x direction-normalized

Static structural and semantic signals; not observed reader comprehension.

Correctness static score

1.29x direction-normalized

Static correctness signals; not a measured production defect rate.

Edit Precision static score

1.36x direction-normalized

Heuristic targetability and change isolation; not observed agent editing success.

Error Detection static score

1.49x direction-normalized

Static detection signals; not observed bug-finding performance.

Generation Accuracy static score

1.02x direction-normalized

Compilation and structural signals; not generation from live agent prompts.

Information Density static score

0.97x direction-normalized

Counted semantic elements per token under the calculator rules.

Refactoring Stability static score

1.38x direction-normalized

Reference preservation under modeled transformations; not a live refactoring trial.

Token Economics static score

1.42x direction-normalized

Composite of token, character, and line ratios; not a raw-token saving.

Per-Program Static Scores

Values are direction-normalized ratios: above 1 favors Calor, below 1 favors C#, and 1 is neutral. Lower-is-better metrics invert their raw score ratio. The legacy composite combines metric ratios; Token Economics combines token, character, and line ratios, not raw-token savings.

Filter by level:
Status
Abs1
1.311.681.601.311.671.000.911.301.04
AbsoluteContracts2
1.563.501.581.392.141.140.821.370.55
AbstractClass2
1.121.151.001.330.861.000.571.351.73
Adapter2
1.551.571.601.661.831.001.051.522.13
AggregateStats2
1.111.091.001.311.001.000.631.321.52
AreaOfCircle1
1.502.271.331.312.141.001.041.301.59
ArrayContracts3
1.471.441.801.352.141.031.021.401.60
ArraySlice2
1.512.781.501.432.141.000.791.301.14
ArraySum2
1.191.480.831.231.001.001.231.321.40
AsyncChain3
1.161.171.001.431.001.000.851.351.47
AsyncEffect3
1.271.361.001.071.331.001.271.471.69
AsyncErrorHandling3
1.391.830.911.481.331.001.591.471.49
AsyncLoop3
1.391.961.001.431.001.001.481.321.91
AsyncPure3
1.040.731.001.031.001.000.791.321.45
AsyncReturn3
1.141.291.001.311.001.000.851.301.33
AutoProps2
1.181.701.001.331.001.001.051.351.01
Average2
1.381.881.001.311.571.001.531.301.44
BankAccount3
1.652.811.501.652.141.001.241.371.45
BasicTryCatch2
1.141.261.081.311.001.000.801.351.34
BclCoverage2
1.391.411.001.431.141.140.881.472.65
BinarySearch3
1.191.320.861.351.001.000.661.401.95
BinaryTree4
1.471.050.921.360.861.000.821.843.92
BitContracts3
1.402.471.601.391.831.140.611.370.74
BitSet3
1.502.711.801.312.501.000.691.440.60
BMICalculator2
1.452.091.901.312.501.000.631.300.91
BreadthFirstSearch3
1.341.450.861.351.001.001.151.402.55
BubbleSort2
1.442.361.171.271.831.001.121.371.40
BuggyContracts2
1.341.431.251.391.671.141.081.461.32
Builder3
1.251.061.201.360.861.000.981.492.07
Calculator2
1.091.291.001.351.001.000.521.371.21
Calendar3
1.452.301.901.312.501.000.661.350.61
CancellableTask3
1.241.581.001.431.001.000.961.321.61
Capitalize2
1.301.821.601.311.831.000.641.300.90
CelsiusToKelvin1
1.512.271.501.312.141.001.011.301.55
ChainOfResponsibility3
1.321.331.231.361.571.000.691.431.91
CircularBuffer3
1.522.241.291.332.141.001.181.351.62
Clamp2
1.441.331.901.351.671.000.921.252.06
CollectionLib2
1.382.291.251.521.671.100.811.400.98
Command3
1.291.560.831.681.001.000.951.401.90
CompactClass2
1.402.381.601.312.001.000.811.390.71
ComposedEffects3
1.371.840.711.391.331.141.271.561.72
Composite3
1.661.761.801.682.501.001.041.492.03
Composition3
1.441.621.401.361.831.001.141.561.64
CompoundInterest2
1.462.211.501.312.141.000.931.321.30
ConstructorInit2
1.552.751.331.332.141.001.221.351.26
Contains2
1.141.441.201.311.001.000.661.321.16
ContractedDivide3
1.310.871.801.311.671.000.691.191.96
CorrectEffects2
1.311.721.001.391.331.140.881.591.45
CountOccurrences2
1.351.901.601.311.671.001.051.320.94
CountVowels2
1.111.501.001.231.001.000.651.351.18
CsvParser2
1.351.971.251.311.671.001.031.321.24
CurrencyConverter2
1.512.701.801.312.501.000.751.290.68
CustomException2
1.351.451.601.351.571.000.871.431.53
DatabaseEffect3
1.352.211.001.391.331.141.071.501.14
DateDiff2
1.573.131.901.312.501.000.641.290.77
Decorator3
1.632.071.801.712.501.000.931.431.60
DelayedResult3
1.411.831.401.312.171.000.931.391.24
DepthFirstSearch3
1.331.551.201.311.001.001.021.352.19
Deque3
1.582.331.501.332.141.001.151.321.85
DictOps2
1.442.351.331.431.671.001.211.271.26
DigitCount2
1.532.641.901.312.501.000.901.440.52
Dijkstra4
1.712.541.801.342.501.001.421.321.78
DisjointSet3
1.291.241.201.521.001.000.861.372.13
DistanceCalculator3
1.191.431.501.311.671.000.521.340.74
DivisionContracts3
1.542.351.801.392.141.140.731.461.28
EditDistance3
1.211.471.001.231.001.001.021.341.65
EmailValidator2
1.081.551.001.311.001.000.531.300.97
Encapsulation3
1.181.071.001.361.001.000.681.371.98
EnumMatch2
1.322.071.201.311.001.001.661.301.03
EnumType3
1.262.091.401.311.001.000.971.321.00
EnumWithMethods3
1.222.091.401.311.001.000.791.300.87
ErrorPropagation2
1.261.881.171.311.571.000.831.301.04
ExceptionChain2
1.401.791.361.311.431.001.471.351.45
ExhaustiveMatch2
1.302.131.171.310.861.001.481.321.14
ExpressionBodied1
1.141.691.001.311.001.000.851.271.02
Factorial2
1.141.311.201.311.001.000.791.371.15
Factory3
1.631.991.501.352.141.000.971.892.21
Fibonacci2
1.141.411.401.311.001.000.751.370.86
FileEffects3
1.392.151.001.391.331.101.051.501.61
Filter2
1.242.061.201.431.001.000.551.271.41
FizzBuzz2
1.231.181.401.271.331.000.261.561.85
Flyweight3
1.571.881.501.362.501.001.321.461.54
FormatHeader2
1.141.291.001.311.001.000.661.441.43
GCD2
1.101.241.001.311.001.000.721.371.13
GenericClass2
1.421.681.501.331.431.001.181.321.89
GenericConstraints3
1.302.351.401.431.001.000.861.271.13
GenericFunction2
1.252.321.001.431.001.000.881.271.12
GradeCalculator2
1.281.681.701.311.831.000.601.300.86
Graph3
1.642.651.501.612.501.001.031.321.48
GroupBy2
1.492.391.401.431.831.000.851.271.73
GuardClause2
1.332.001.331.311.571.000.651.351.40
GuardMatch2
1.311.861.401.311.001.001.321.301.33
HashMap3
1.621.941.501.282.501.001.351.431.95
HelloWorld1
1.402.611.001.311.331.000.941.391.57
HiddenNetworkEffect3
1.501.891.001.431.141.141.651.602.18
Hypotenuse2
1.472.691.501.312.141.000.891.300.95
Inheritance3
1.291.441.401.481.001.001.281.371.33
InlineSigs2
1.121.741.201.311.001.000.591.270.86
InsertionSort2
1.242.040.831.231.001.001.041.351.43
InterfaceImpl3
1.551.811.401.681.831.001.161.591.92
InterpolationSearch3
1.311.551.141.311.831.000.861.371.39
Inventory3
1.542.661.501.332.141.000.941.321.43
IsAlpha1
1.111.911.201.311.001.000.371.270.82
IsEven1
1.141.211.201.311.001.000.521.301.61
IsOdd1
1.121.211.001.311.001.000.521.301.62
IsPrime3
1.221.181.601.351.221.000.631.301.45
Iterator2
1.411.751.401.341.831.001.201.321.42
Knapsack4
1.582.661.581.232.501.001.101.351.20
LCS3
1.221.591.171.231.001.001.031.301.45
LeapYear2
1.181.361.401.351.001.000.411.401.49
LinearSearch2
1.101.570.861.231.001.000.781.321.04
LinkedList4
1.601.001.401.360.861.001.071.674.48
LinqPipeline2
1.191.621.171.311.001.000.871.271.26
ListInvariant3
1.603.581.801.392.141.140.851.370.52
ListOps2
1.342.061.331.311.671.001.081.271.00
LogPipeline2
1.151.581.001.311.001.000.531.421.33
Map2
1.402.951.001.431.001.000.871.271.67
MathLib2
1.392.471.601.391.671.100.741.370.77
MathOperations3
1.320.821.601.351.111.000.911.372.40
MatrixMultiply3
1.421.951.331.341.571.001.301.301.57
MaxHeap3
1.481.631.171.361.571.001.281.432.44
MaxTwo1
1.261.701.401.311.671.000.761.300.95
MaxValue2
1.101.331.001.230.671.001.201.231.12
Mediator3
1.301.400.831.360.861.001.121.372.46
MergeSort3
1.521.441.331.351.831.000.971.372.88
MethodOverloading3
1.081.621.001.311.001.000.571.270.90
MethodOverriding3
1.251.561.001.431.001.001.551.321.13
MinStack3
1.311.391.171.610.861.001.011.322.12
MinTwo1
1.261.701.401.311.671.000.761.300.95
MissingEffects2
1.381.951.001.391.331.140.921.561.72
MixedContracts3
1.461.551.501.392.141.141.011.491.47
MixedSyntax2
1.402.111.331.312.141.000.791.321.23
ModuloContracts3
1.643.641.901.392.501.140.681.400.43
MultipleCatch2
1.141.271.091.351.001.000.731.371.28
NamedConfig2
1.121.251.001.311.001.000.621.421.32
NestedMatch2
1.311.861.201.311.001.001.431.301.35
NetworkEffect3
1.362.171.001.391.331.141.201.531.15
NullCheck2
1.101.660.801.310.861.000.701.271.16
Observer3
1.291.401.201.361.001.000.861.402.10
OptionalIds2
1.091.621.001.311.001.000.591.270.94
OptionType2
1.482.431.501.411.571.101.241.431.16
OverflowSafe3
1.511.461.801.352.141.031.291.401.63
OverflowUnsafe3
1.611.621.801.312.501.032.001.251.36
PairLogger2
1.111.421.001.311.001.000.471.421.23
Palindrome2
1.181.491.001.311.001.001.011.351.26
ParallelTasks3
1.292.161.001.431.001.001.061.301.38
ParseAndDouble2
1.141.291.001.311.001.000.661.441.42
PasswordValidator2
1.111.491.001.311.001.000.691.321.10
PhoneBook2
1.522.381.501.332.501.000.911.271.27
Polymorphism3
1.251.191.401.681.001.000.821.431.46
Power2
1.261.061.601.351.221.000.891.271.70
PriorityQueue3
1.541.711.331.361.571.001.051.492.77
Properties2
1.391.831.331.481.571.101.041.401.39
PropertyAccess1
1.831.631.601.481.831.104.251.431.34
PropertyMatch2
1.371.791.201.651.001.001.481.371.46
ProvableContracts2
1.482.101.581.392.141.140.981.491.01
Proxy3
1.301.411.201.681.141.000.881.531.59
PureComputation2
1.342.101.501.391.671.140.751.370.83
PureFunctions1
1.361.501.701.431.831.140.741.580.99
Queue3
1.301.221.171.360.861.001.091.402.29
QuickSort4
1.571.441.141.271.571.000.901.563.69
RangeContracts3
1.502.451.581.392.501.140.731.460.77
RangeMatch2
1.542.321.701.311.831.001.781.301.11
RecordType3
1.442.801.501.711.671.100.741.430.56
Reduce2
1.232.001.201.311.001.000.791.271.24
ResultType2
1.241.530.921.361.001.001.021.371.71
ReturnMapped2
1.131.251.001.311.001.000.641.421.45
ReverseArray2
1.532.871.331.311.831.001.591.321.02
ReverseString2
1.221.400.831.311.001.001.101.321.80
ScoreBoard2
1.572.371.801.332.501.000.951.271.34
SealedClass2
1.241.461.001.331.001.000.831.351.96
SearchContracts3
1.572.921.801.392.501.140.771.400.62
SelectionSort2
1.442.231.171.231.831.001.271.351.40
SetOps2
1.482.221.501.431.671.001.201.351.48
ShoppingCart4
1.711.901.191.282.141.001.061.933.23
Sign1
1.101.381.001.311.001.000.611.301.23
SimpleAsync3
1.331.961.001.311.331.001.341.391.34
SimpleClass2
1.632.361.801.652.501.000.881.431.40
SimpleMatch2
1.372.151.401.311.001.001.871.300.90
Singleton2
1.171.310.771.361.001.000.831.371.71
Sort2
1.161.580.831.231.001.000.781.301.55
SortedList2
1.371.651.251.591.671.000.971.301.50
SortingContracts3
1.563.071.801.392.501.140.551.370.68
Stack3
1.331.151.401.360.861.000.991.402.46
State3
1.442.051.401.361.001.002.071.371.24
StateEffect3
1.392.001.001.411.331.141.661.501.10
StaticMembers3
1.261.371.001.431.001.101.341.321.48
Strategy3
1.261.331.001.501.001.000.751.402.12
StringContracts3
1.371.451.291.312.141.031.011.321.42
StringLib2
1.322.331.251.391.671.100.651.370.80
StringUtils2
1.081.290.831.351.001.000.511.371.30
SumDigits2
1.552.871.901.312.501.000.951.440.42
SumRange2
1.311.481.171.231.571.000.951.341.71
SwitchExpression2
1.561.931.701.311.831.002.221.351.11
TaxCalculator3
1.422.281.581.312.501.000.651.300.73
TemperatureConverter1
1.111.281.001.311.001.000.631.351.29
TemperatureRange2
1.121.141.001.311.001.000.691.321.47
TemplateMethod3
1.000.371.100.940.860.760.731.831.44
TernaryChain2
1.191.611.401.311.001.001.021.270.89
ThreeWayMerge2
1.091.151.001.311.001.000.521.421.33
Timer2
1.532.571.501.362.141.001.191.401.11
TimerEffect3
1.593.131.901.392.501.100.691.400.66
TodoList2
1.542.061.501.592.141.001.231.321.44
TowerOfHanoi3
1.502.331.581.352.141.000.891.491.18
Trie4
1.621.801.391.362.501.001.111.462.31
TruncateString2
1.261.851.071.311.571.000.931.301.04
TryFinally2
1.181.481.001.311.001.001.011.351.28
TupleMatch2
1.241.911.401.311.001.001.071.300.91
TypeAlias3
1.452.561.171.411.381.101.381.301.34
TypeMatch2
1.411.751.401.351.001.002.141.371.27
UnitConverter2
1.352.371.601.311.831.000.691.270.69
UrlParser2
1.161.551.401.311.001.000.751.270.96
Visitor3
1.040.421.000.981.000.760.552.131.50
VoidSequence2
1.323.031.001.351.001.000.601.471.14
VotingSystem2
1.401.951.601.331.831.000.951.301.20
WildcardMatch2
1.201.841.401.311.001.000.931.300.80
Zip2
1.282.431.001.311.001.000.841.271.36

Showing 217 of 217 programs.


Read the Results Carefully

The static snapshot above covers 217 programs and reports a legacy composite direction-normalized ratio of 1.32x. This is not a measured language advantage. Comprehension is 1.84x, error detection is 1.49x, token economics is 1.42x, and information density is 0.97x—the category whose normalized ratio favors C#. Seven category ratios favor Calor and one favors C#. The Calor parser accepted 217 of 217 inputs; the Roslyn syntax parser accepted 217 of 217 C# inputs. These are parse checks, not evidence that the pairs compute the same result.

The snapshot was recorded on 2026-09-15 from source revision c2a8816d. Its Directory.Build.props declares compiler 0.21.0, before the release version bump; the JSON's version: "1.0" identifies the data format, not the compiler release. The corpus is tests/TestData/Benchmarks. The exact compiler binary used for the measurement was not independently recorded. These are fixed-source results, rerun for the 0.22.0 release. They are not a live measure of coding-agent performance. The separate agent-study snapshots below were not rerun.

These measurements need important caveats:

  • Source pairs are not all behaviorally equivalent. For example, the C# CsvParser implements parsing, while its Calor counterpart contains only helper functions. The 0.19 rerun fixes an invalid newline literal in the C# fixture, but does not establish equal work across the corpus. The ratios therefore cannot establish a language advantage. See the benchmark integrity follow-up.
  • The metrics are deterministic static analyses. The exporter reports both zero-width intervals from singleton category aggregates and pooled intervals from repeated program scores. Thirty repetitions are not independent samples. Neither interval establishes sampling uncertainty over the corpus or a measured language, agent-productivity, correctness, or safety advantage.
  • These C#-versus-Calor micro-benchmarks are not the v0.12 release gates. PP-A1 covered adoption readiness, while PP-W5 measured toolchain tax and concluded only that no large tax was detected—not that parity was proven.

Token economics is a composite of token, character, and line ratios. Calor still pays a § token premium on small programs even though the current composite favors Calor.


Separate Agent Studies

The static dashboard does not measure effect rows or live coding-agent behavior. Those questions use separate, pre-registered agent studies. The research evidence status explains why the proposed follow-on adopter study stopped unadjudicated.

Redesigned PP-W-rows: paused, not complete

Current status, 2026-09-11: paused by the user. The central pause handoff records two consumed invalid/censored attempts and 442 untouched scheduled identities out of 444. Neither the pilot nor confirmation has completed. These interrupted attempts do not supply an effect-row benefit, negative, or null result. They must not be erased, replaced, or described as an entirely unrun redesigned study.

The single total experiment ceiling is USD1,000. Two reservations of USD25.52 each remain held: USD51.04 total. Actual provider charges are unknown; the held amount is not measured spend or reconciled charges. The conditionally authorized permanent-retention exception has not been applied. Its implementation and final reviews remain unfinished in the accounting issue and paused draft checkpoint.

There is no permission to resume. A new explicit instruction and the remaining implementation, review, and operator-readiness gates are required. This product release neither resumes collection nor changes the protocol, consumed attempts, unknown costs, or existing ceiling. Confirmation would still need its own prospective design and affordability decision; it is not automatic.

Earlier decision, 2026-09-10: USD250 total, pilot only (superseded ceiling). The recorded user decision and historical authorization receipt accepted the registered stopping rules and publication of negative or null results under a conservative pilot-only interpretation. Its dated BUDGET_NOT_RUN assessment described an earlier pre-execution hold, not the current pause after consumed attempts. Historical planning estimates were not cost lower bounds. Neither that hold nor today's pause is a pilot null or the conditional stage-2 UNDERPOWERED-CARRIED outcome.

Historical M0 disposition, 2026-09-09: unchanged. The M0 decision records formal M0 as UNADJUDICATED, with a separate maintainer administrative stop. It deferred this narrower proposal pending a separate explicit decision. Later PP-W financial decisions do not re-arm M0, supply its missing human or participant approvals, or establish that effect rows help, fail, or cannot be studied.

Local buildability passed on 2026-09-10. The reviewed gate closure records Exit A under the frozen v0.18 release criteria. This deterministic engineering check is not an agent observation. Subsequent engineering work does not turn the two interrupted attempts into a completed study or benefit result.

The proposed methodology explains the old design defect: the unwanted behavior was already visibly forbidden, so avoiding it did not require effect enforcement. The redesigned question concerns violations hidden behind an abstraction while visible tests pass. Historical zero escapes do not answer that question.

The historical benefit ledger records the old registration as UNDERPOWERED, epochRun: false, and legA/legB: null. Null means uncollected for that historical registration, not zero violations. It is not the record of the two consumed redesigned attempts. This publication changes neither ledger nor scientific gate.

Later agent-result publication must follow actual adjudication by the scientific owners. The current pause is not a completed pilot, unaffordable confirmation verdict, or confirmatory result.

Historical PP-E1: effect-row toolchain cost

The 2026-08-27 PP-E1 report records two distinct legs, not one pooled sample:

  • A seeded-fixture check caught 10 of 10 registered mutations at compiler revision 121c6681, with frozen unmutated diagnostic baselines reproduced. These are hand-seeded cases, not agent trials or a population detection rate.
  • Four ordinary programming tasks ran five times on each of two compiler configurations, for 40 valid green runs, with no dropped task pairs or censored runs in the archived analysis. The median paired output-token ratio was 1.1835. Its one-sided 95% two-level cluster-bootstrap lower bound was 0.9012, below the registered 1.0 bound gate, and the point estimate stayed below the 1.35 failure threshold.

The epoch pins record e1-rows-parity-001, started 2026-08-27, claude-opus-4-8, Claude Code 2.1.243, and the v0.14.3 control at 63316987 versus the 0.15.0 release-build treatment at b775acb4. These are different compiler versions, not the redesigned one-compiler policy contrast.

The combined registered verdict was HIT. The tasks measured compiler feedback cost on ordinary code; because they contained no callbacks, they did not measure the safety benefit of effect rows.

Historical PP-W-rows: practice rounds, not confirmation

The 2026-09-01 practice report describes historical dry runs used for sizing, not a confirmatory verdict. Both used claude-opus-4-8 with permissive v0.14.3-pre-rows control 283ec9f9 versus strict v0.15.0 treatment 3bb2601e.

ArchiveDate and agentWhat can be read from it
w-rows-dry-001 pins2026-08-28; Claude Code 2.1.248Usage-limit HTTP 429 runs were invalid. The report uses different “valid” and “readable” counts; this page does not pool them into a denominator.
w-rows-dry-002 pins2026-09-01; Claude Code 2.1.252Three tasks, three runs per arm: 18 valid built runs, no exclusions. Zero observed escapes in either arm. This is a finite practice sample, not proof of a zero population escape rate.

The second dry-run analysis reports a descriptive output-tokens-to-green ratio of 1.1175, with a one-sided 95% cluster-bootstrap lower bound of 0.8634, over its two cost-eligible task pairs (12 runs). It is not a result over all 18 runs or a measurement of general development effort.

Its old-design sizing required nine runs per task/arm, while the retired ceiling afforded three, with calculated power 0.48 under that registration. These are historical sizing assumptions, not the redesigned sample size or funding. The benefit ledger therefore remains UNDERPOWERED: the claim is neither supported nor refuted. The redesigned defect analysis also means buying more of the old tasks would not isolate the intended benefit.

The ledger must be read with care. Its top-level achievablePower field is 0.868, but the recorded reason and power curve use 0.868 for the smallest unaffordable design, nine runs per cell. The largest affordable design was three runs per cell at 0.48 power. The public interpretation therefore does not present 0.868 as affordable power.

The practice result exposed a fixture-design defect: the undesirable hidden effect was already forbidden by the visible task specification. The 2026-09-08 redesign separates visible requirements from a held-out effect-observing test and adds eight task-qualification rules. At the 2026-09-09 M0 closeout, no candidate had established both R1 (a convenient effect-hiding API) and R4 (the laundering solution passes visible tests). The local buildability gate subsequently passed on 2026-09-10, as recorded above. The two later consumed attempts and the user-directed pause are separate from these practice results.

On 2026-09-09, M0 recorded these missing prerequisites and the lack of a suitable isolated evaluator. Formal M0 status is UNADJUDICATED; the maintainer action was an administrative stop. At that closeout, no redesigned pilot or confirmatory collection was authorized. The old w-rows-001 epoch remains unrun, not a null result; this does not erase the later redesigned attempts.

Keep these statements separate:

  • The 1.32x dashboard is a deterministic source/structure comparison.
  • The 0.15 proof point did not detect a large toolchain tax.
  • The historical PP-W-rows registration did not run a confirmatory epoch.
  • The redesigned study is paused after two invalid/censored attempts, with 442 untouched identities; neither stage has completed and no benefit or null result exists.
  • The redesign and M0 records contain no protected three-arm adopter comparison, independent methods adjudication, or safety/economic advantage result.

See the artifact inventory, the proposed redesign, and the M0 evidence status for source links and execution states.


When to Use Calor

These scores are not sufficient evidence for choosing a language. Consider Calor when you want to experiment with:

  • Explicit contracts and the documented limits of optional static verification
  • Explicit IDs and structural markers for agent-editing experiments
  • Structured source representations, while measuring costs on your own tasks

Use C# when:

  • Ecosystem libraries are needed — Leverage existing tooling
  • Human readability is priority — Familiar syntax for human developers

Methodology

The current static runner is automated and can regenerate a new artifact from the checked-out source. That is not the same as exactly reproducing every historical artifact: several omit the measurement binary, working-tree state, source revision, model, or runtime configuration. Proposed protocols have not run.

The published dashboard framework:

  1. Compiles paired Calor and C# programs.
  2. Applies deterministic metric calculators to source, syntax trees, and compilation results.
  3. Aggregates per-program direction-normalized ratios. Values above 1 favor Calor, including lower-is-better metrics whose raw score ratios are inverted.
  4. Repeats the deterministic evaluation 30 times for the reported statistical artifact. Greater than 1.0 favors Calor and less than 1.0 favors C#; neither is a measured language advantage.

LLM-based task runners exist separately and are not folded into this dashboard result.

See Methodology for full details on the evaluation framework.


Running Benchmarks

To regenerate benchmark results:

Bash
# Run the evaluation framework
dotnet run --project tests/Calor.Evaluation -c Release -- run -f website -o website/public/data/benchmark-results.json --statistical --runs 30

# Start the website to view results
cd website && npm run dev

Results are automatically loaded from benchmark-results.json and displayed in the dashboard above.