Results
Static Benchmark Snapshot
Evaluated across 217 programs with 8 metrics
Recorded source revision c2a8816d declares compiler v0.21.0; this is a pinned static-calculator run, not an agent-productivity measurement.
Corpus: tests/TestData/Benchmarks — 217 programs.
Method: Deterministic static-analysis calculators, 30 repetitions, 8 metrics.
The reports record the source revision before the release version bump, not an independently attested measurement binary. The 30 repetitions are deterministic, not independent samples. Source pairs are not all behaviorally equivalent; the reported intervals establish no measured language, agent-productivity, correctness, or safety advantage. See issue #1276.
Agent Refactoring Benchmark
Historical Claude Code study of the bundled refactoring tasks (rename, extract, inline, move, add contracts, change signature), not a measurement of current-release productivity.
View category breakdown
Historical majority-vote refactoring study (2 of 3 runs). Corpus: tests/E2E/agent-tasks; recorded source declares compiler v0.2.3. Exact measurement binary not independently recorded.
Static Scores by Metric
Alphabetical metric order. Ratios are direction-normalized: above 1 favors Calor and below 1 favors C#. Lower-is-better metrics invert their raw score ratio. Neither means a language is better.
Comprehension static score
1.84x direction-normalizedStatic structural and semantic signals; not observed reader comprehension.
Correctness static score
1.29x direction-normalizedStatic correctness signals; not a measured production defect rate.
Edit Precision static score
1.36x direction-normalizedHeuristic targetability and change isolation; not observed agent editing success.
Error Detection static score
1.49x direction-normalizedStatic detection signals; not observed bug-finding performance.
Generation Accuracy static score
1.02x direction-normalizedCompilation and structural signals; not generation from live agent prompts.
Information Density static score
0.97x direction-normalizedCounted semantic elements per token under the calculator rules.
Refactoring Stability static score
1.38x direction-normalizedReference preservation under modeled transformations; not a live refactoring trial.
Token Economics static score
1.42x direction-normalizedComposite of token, character, and line ratios; not a raw-token saving.
Per-Program Static Scores
Values are direction-normalized ratios: above 1 favors Calor, below 1 favors C#, and 1 is neutral. Lower-is-better metrics invert their raw score ratio. The legacy composite combines metric ratios; Token Economics combines token, character, and line ratios, not raw-token savings.
| Status | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Abs | 1 | 1.31 | 1.68 | 1.60 | 1.31 | 1.67 | 1.00 | 0.91 | 1.30 | 1.04 | |
| AbsoluteContracts | 2 | 1.56 | 3.50 | 1.58 | 1.39 | 2.14 | 1.14 | 0.82 | 1.37 | 0.55 | |
| AbstractClass | 2 | 1.12 | 1.15 | 1.00 | 1.33 | 0.86 | 1.00 | 0.57 | 1.35 | 1.73 | |
| Adapter | 2 | 1.55 | 1.57 | 1.60 | 1.66 | 1.83 | 1.00 | 1.05 | 1.52 | 2.13 | |
| AggregateStats | 2 | 1.11 | 1.09 | 1.00 | 1.31 | 1.00 | 1.00 | 0.63 | 1.32 | 1.52 | |
| AreaOfCircle | 1 | 1.50 | 2.27 | 1.33 | 1.31 | 2.14 | 1.00 | 1.04 | 1.30 | 1.59 | |
| ArrayContracts | 3 | 1.47 | 1.44 | 1.80 | 1.35 | 2.14 | 1.03 | 1.02 | 1.40 | 1.60 | |
| ArraySlice | 2 | 1.51 | 2.78 | 1.50 | 1.43 | 2.14 | 1.00 | 0.79 | 1.30 | 1.14 | |
| ArraySum | 2 | 1.19 | 1.48 | 0.83 | 1.23 | 1.00 | 1.00 | 1.23 | 1.32 | 1.40 | |
| AsyncChain | 3 | 1.16 | 1.17 | 1.00 | 1.43 | 1.00 | 1.00 | 0.85 | 1.35 | 1.47 | |
| AsyncEffect | 3 | 1.27 | 1.36 | 1.00 | 1.07 | 1.33 | 1.00 | 1.27 | 1.47 | 1.69 | |
| AsyncErrorHandling | 3 | 1.39 | 1.83 | 0.91 | 1.48 | 1.33 | 1.00 | 1.59 | 1.47 | 1.49 | |
| AsyncLoop | 3 | 1.39 | 1.96 | 1.00 | 1.43 | 1.00 | 1.00 | 1.48 | 1.32 | 1.91 | |
| AsyncPure | 3 | 1.04 | 0.73 | 1.00 | 1.03 | 1.00 | 1.00 | 0.79 | 1.32 | 1.45 | |
| AsyncReturn | 3 | 1.14 | 1.29 | 1.00 | 1.31 | 1.00 | 1.00 | 0.85 | 1.30 | 1.33 | |
| AutoProps | 2 | 1.18 | 1.70 | 1.00 | 1.33 | 1.00 | 1.00 | 1.05 | 1.35 | 1.01 | |
| Average | 2 | 1.38 | 1.88 | 1.00 | 1.31 | 1.57 | 1.00 | 1.53 | 1.30 | 1.44 | |
| BankAccount | 3 | 1.65 | 2.81 | 1.50 | 1.65 | 2.14 | 1.00 | 1.24 | 1.37 | 1.45 | |
| BasicTryCatch | 2 | 1.14 | 1.26 | 1.08 | 1.31 | 1.00 | 1.00 | 0.80 | 1.35 | 1.34 | |
| BclCoverage | 2 | 1.39 | 1.41 | 1.00 | 1.43 | 1.14 | 1.14 | 0.88 | 1.47 | 2.65 | |
| BinarySearch | 3 | 1.19 | 1.32 | 0.86 | 1.35 | 1.00 | 1.00 | 0.66 | 1.40 | 1.95 | |
| BinaryTree | 4 | 1.47 | 1.05 | 0.92 | 1.36 | 0.86 | 1.00 | 0.82 | 1.84 | 3.92 | |
| BitContracts | 3 | 1.40 | 2.47 | 1.60 | 1.39 | 1.83 | 1.14 | 0.61 | 1.37 | 0.74 | |
| BitSet | 3 | 1.50 | 2.71 | 1.80 | 1.31 | 2.50 | 1.00 | 0.69 | 1.44 | 0.60 | |
| BMICalculator | 2 | 1.45 | 2.09 | 1.90 | 1.31 | 2.50 | 1.00 | 0.63 | 1.30 | 0.91 | |
| BreadthFirstSearch | 3 | 1.34 | 1.45 | 0.86 | 1.35 | 1.00 | 1.00 | 1.15 | 1.40 | 2.55 | |
| BubbleSort | 2 | 1.44 | 2.36 | 1.17 | 1.27 | 1.83 | 1.00 | 1.12 | 1.37 | 1.40 | |
| BuggyContracts | 2 | 1.34 | 1.43 | 1.25 | 1.39 | 1.67 | 1.14 | 1.08 | 1.46 | 1.32 | |
| Builder | 3 | 1.25 | 1.06 | 1.20 | 1.36 | 0.86 | 1.00 | 0.98 | 1.49 | 2.07 | |
| Calculator | 2 | 1.09 | 1.29 | 1.00 | 1.35 | 1.00 | 1.00 | 0.52 | 1.37 | 1.21 | |
| Calendar | 3 | 1.45 | 2.30 | 1.90 | 1.31 | 2.50 | 1.00 | 0.66 | 1.35 | 0.61 | |
| CancellableTask | 3 | 1.24 | 1.58 | 1.00 | 1.43 | 1.00 | 1.00 | 0.96 | 1.32 | 1.61 | |
| Capitalize | 2 | 1.30 | 1.82 | 1.60 | 1.31 | 1.83 | 1.00 | 0.64 | 1.30 | 0.90 | |
| CelsiusToKelvin | 1 | 1.51 | 2.27 | 1.50 | 1.31 | 2.14 | 1.00 | 1.01 | 1.30 | 1.55 | |
| ChainOfResponsibility | 3 | 1.32 | 1.33 | 1.23 | 1.36 | 1.57 | 1.00 | 0.69 | 1.43 | 1.91 | |
| CircularBuffer | 3 | 1.52 | 2.24 | 1.29 | 1.33 | 2.14 | 1.00 | 1.18 | 1.35 | 1.62 | |
| Clamp | 2 | 1.44 | 1.33 | 1.90 | 1.35 | 1.67 | 1.00 | 0.92 | 1.25 | 2.06 | |
| CollectionLib | 2 | 1.38 | 2.29 | 1.25 | 1.52 | 1.67 | 1.10 | 0.81 | 1.40 | 0.98 | |
| Command | 3 | 1.29 | 1.56 | 0.83 | 1.68 | 1.00 | 1.00 | 0.95 | 1.40 | 1.90 | |
| CompactClass | 2 | 1.40 | 2.38 | 1.60 | 1.31 | 2.00 | 1.00 | 0.81 | 1.39 | 0.71 | |
| ComposedEffects | 3 | 1.37 | 1.84 | 0.71 | 1.39 | 1.33 | 1.14 | 1.27 | 1.56 | 1.72 | |
| Composite | 3 | 1.66 | 1.76 | 1.80 | 1.68 | 2.50 | 1.00 | 1.04 | 1.49 | 2.03 | |
| Composition | 3 | 1.44 | 1.62 | 1.40 | 1.36 | 1.83 | 1.00 | 1.14 | 1.56 | 1.64 | |
| CompoundInterest | 2 | 1.46 | 2.21 | 1.50 | 1.31 | 2.14 | 1.00 | 0.93 | 1.32 | 1.30 | |
| ConstructorInit | 2 | 1.55 | 2.75 | 1.33 | 1.33 | 2.14 | 1.00 | 1.22 | 1.35 | 1.26 | |
| Contains | 2 | 1.14 | 1.44 | 1.20 | 1.31 | 1.00 | 1.00 | 0.66 | 1.32 | 1.16 | |
| ContractedDivide | 3 | 1.31 | 0.87 | 1.80 | 1.31 | 1.67 | 1.00 | 0.69 | 1.19 | 1.96 | |
| CorrectEffects | 2 | 1.31 | 1.72 | 1.00 | 1.39 | 1.33 | 1.14 | 0.88 | 1.59 | 1.45 | |
| CountOccurrences | 2 | 1.35 | 1.90 | 1.60 | 1.31 | 1.67 | 1.00 | 1.05 | 1.32 | 0.94 | |
| CountVowels | 2 | 1.11 | 1.50 | 1.00 | 1.23 | 1.00 | 1.00 | 0.65 | 1.35 | 1.18 | |
| CsvParser | 2 | 1.35 | 1.97 | 1.25 | 1.31 | 1.67 | 1.00 | 1.03 | 1.32 | 1.24 | |
| CurrencyConverter | 2 | 1.51 | 2.70 | 1.80 | 1.31 | 2.50 | 1.00 | 0.75 | 1.29 | 0.68 | |
| CustomException | 2 | 1.35 | 1.45 | 1.60 | 1.35 | 1.57 | 1.00 | 0.87 | 1.43 | 1.53 | |
| DatabaseEffect | 3 | 1.35 | 2.21 | 1.00 | 1.39 | 1.33 | 1.14 | 1.07 | 1.50 | 1.14 | |
| DateDiff | 2 | 1.57 | 3.13 | 1.90 | 1.31 | 2.50 | 1.00 | 0.64 | 1.29 | 0.77 | |
| Decorator | 3 | 1.63 | 2.07 | 1.80 | 1.71 | 2.50 | 1.00 | 0.93 | 1.43 | 1.60 | |
| DelayedResult | 3 | 1.41 | 1.83 | 1.40 | 1.31 | 2.17 | 1.00 | 0.93 | 1.39 | 1.24 | |
| DepthFirstSearch | 3 | 1.33 | 1.55 | 1.20 | 1.31 | 1.00 | 1.00 | 1.02 | 1.35 | 2.19 | |
| Deque | 3 | 1.58 | 2.33 | 1.50 | 1.33 | 2.14 | 1.00 | 1.15 | 1.32 | 1.85 | |
| DictOps | 2 | 1.44 | 2.35 | 1.33 | 1.43 | 1.67 | 1.00 | 1.21 | 1.27 | 1.26 | |
| DigitCount | 2 | 1.53 | 2.64 | 1.90 | 1.31 | 2.50 | 1.00 | 0.90 | 1.44 | 0.52 | |
| Dijkstra | 4 | 1.71 | 2.54 | 1.80 | 1.34 | 2.50 | 1.00 | 1.42 | 1.32 | 1.78 | |
| DisjointSet | 3 | 1.29 | 1.24 | 1.20 | 1.52 | 1.00 | 1.00 | 0.86 | 1.37 | 2.13 | |
| DistanceCalculator | 3 | 1.19 | 1.43 | 1.50 | 1.31 | 1.67 | 1.00 | 0.52 | 1.34 | 0.74 | |
| DivisionContracts | 3 | 1.54 | 2.35 | 1.80 | 1.39 | 2.14 | 1.14 | 0.73 | 1.46 | 1.28 | |
| EditDistance | 3 | 1.21 | 1.47 | 1.00 | 1.23 | 1.00 | 1.00 | 1.02 | 1.34 | 1.65 | |
| EmailValidator | 2 | 1.08 | 1.55 | 1.00 | 1.31 | 1.00 | 1.00 | 0.53 | 1.30 | 0.97 | |
| Encapsulation | 3 | 1.18 | 1.07 | 1.00 | 1.36 | 1.00 | 1.00 | 0.68 | 1.37 | 1.98 | |
| EnumMatch | 2 | 1.32 | 2.07 | 1.20 | 1.31 | 1.00 | 1.00 | 1.66 | 1.30 | 1.03 | |
| EnumType | 3 | 1.26 | 2.09 | 1.40 | 1.31 | 1.00 | 1.00 | 0.97 | 1.32 | 1.00 | |
| EnumWithMethods | 3 | 1.22 | 2.09 | 1.40 | 1.31 | 1.00 | 1.00 | 0.79 | 1.30 | 0.87 | |
| ErrorPropagation | 2 | 1.26 | 1.88 | 1.17 | 1.31 | 1.57 | 1.00 | 0.83 | 1.30 | 1.04 | |
| ExceptionChain | 2 | 1.40 | 1.79 | 1.36 | 1.31 | 1.43 | 1.00 | 1.47 | 1.35 | 1.45 | |
| ExhaustiveMatch | 2 | 1.30 | 2.13 | 1.17 | 1.31 | 0.86 | 1.00 | 1.48 | 1.32 | 1.14 | |
| ExpressionBodied | 1 | 1.14 | 1.69 | 1.00 | 1.31 | 1.00 | 1.00 | 0.85 | 1.27 | 1.02 | |
| Factorial | 2 | 1.14 | 1.31 | 1.20 | 1.31 | 1.00 | 1.00 | 0.79 | 1.37 | 1.15 | |
| Factory | 3 | 1.63 | 1.99 | 1.50 | 1.35 | 2.14 | 1.00 | 0.97 | 1.89 | 2.21 | |
| Fibonacci | 2 | 1.14 | 1.41 | 1.40 | 1.31 | 1.00 | 1.00 | 0.75 | 1.37 | 0.86 | |
| FileEffects | 3 | 1.39 | 2.15 | 1.00 | 1.39 | 1.33 | 1.10 | 1.05 | 1.50 | 1.61 | |
| Filter | 2 | 1.24 | 2.06 | 1.20 | 1.43 | 1.00 | 1.00 | 0.55 | 1.27 | 1.41 | |
| FizzBuzz | 2 | 1.23 | 1.18 | 1.40 | 1.27 | 1.33 | 1.00 | 0.26 | 1.56 | 1.85 | |
| Flyweight | 3 | 1.57 | 1.88 | 1.50 | 1.36 | 2.50 | 1.00 | 1.32 | 1.46 | 1.54 | |
| FormatHeader | 2 | 1.14 | 1.29 | 1.00 | 1.31 | 1.00 | 1.00 | 0.66 | 1.44 | 1.43 | |
| GCD | 2 | 1.10 | 1.24 | 1.00 | 1.31 | 1.00 | 1.00 | 0.72 | 1.37 | 1.13 | |
| GenericClass | 2 | 1.42 | 1.68 | 1.50 | 1.33 | 1.43 | 1.00 | 1.18 | 1.32 | 1.89 | |
| GenericConstraints | 3 | 1.30 | 2.35 | 1.40 | 1.43 | 1.00 | 1.00 | 0.86 | 1.27 | 1.13 | |
| GenericFunction | 2 | 1.25 | 2.32 | 1.00 | 1.43 | 1.00 | 1.00 | 0.88 | 1.27 | 1.12 | |
| GradeCalculator | 2 | 1.28 | 1.68 | 1.70 | 1.31 | 1.83 | 1.00 | 0.60 | 1.30 | 0.86 | |
| Graph | 3 | 1.64 | 2.65 | 1.50 | 1.61 | 2.50 | 1.00 | 1.03 | 1.32 | 1.48 | |
| GroupBy | 2 | 1.49 | 2.39 | 1.40 | 1.43 | 1.83 | 1.00 | 0.85 | 1.27 | 1.73 | |
| GuardClause | 2 | 1.33 | 2.00 | 1.33 | 1.31 | 1.57 | 1.00 | 0.65 | 1.35 | 1.40 | |
| GuardMatch | 2 | 1.31 | 1.86 | 1.40 | 1.31 | 1.00 | 1.00 | 1.32 | 1.30 | 1.33 | |
| HashMap | 3 | 1.62 | 1.94 | 1.50 | 1.28 | 2.50 | 1.00 | 1.35 | 1.43 | 1.95 | |
| HelloWorld | 1 | 1.40 | 2.61 | 1.00 | 1.31 | 1.33 | 1.00 | 0.94 | 1.39 | 1.57 | |
| HiddenNetworkEffect | 3 | 1.50 | 1.89 | 1.00 | 1.43 | 1.14 | 1.14 | 1.65 | 1.60 | 2.18 | |
| Hypotenuse | 2 | 1.47 | 2.69 | 1.50 | 1.31 | 2.14 | 1.00 | 0.89 | 1.30 | 0.95 | |
| Inheritance | 3 | 1.29 | 1.44 | 1.40 | 1.48 | 1.00 | 1.00 | 1.28 | 1.37 | 1.33 | |
| InlineSigs | 2 | 1.12 | 1.74 | 1.20 | 1.31 | 1.00 | 1.00 | 0.59 | 1.27 | 0.86 | |
| InsertionSort | 2 | 1.24 | 2.04 | 0.83 | 1.23 | 1.00 | 1.00 | 1.04 | 1.35 | 1.43 | |
| InterfaceImpl | 3 | 1.55 | 1.81 | 1.40 | 1.68 | 1.83 | 1.00 | 1.16 | 1.59 | 1.92 | |
| InterpolationSearch | 3 | 1.31 | 1.55 | 1.14 | 1.31 | 1.83 | 1.00 | 0.86 | 1.37 | 1.39 | |
| Inventory | 3 | 1.54 | 2.66 | 1.50 | 1.33 | 2.14 | 1.00 | 0.94 | 1.32 | 1.43 | |
| IsAlpha | 1 | 1.11 | 1.91 | 1.20 | 1.31 | 1.00 | 1.00 | 0.37 | 1.27 | 0.82 | |
| IsEven | 1 | 1.14 | 1.21 | 1.20 | 1.31 | 1.00 | 1.00 | 0.52 | 1.30 | 1.61 | |
| IsOdd | 1 | 1.12 | 1.21 | 1.00 | 1.31 | 1.00 | 1.00 | 0.52 | 1.30 | 1.62 | |
| IsPrime | 3 | 1.22 | 1.18 | 1.60 | 1.35 | 1.22 | 1.00 | 0.63 | 1.30 | 1.45 | |
| Iterator | 2 | 1.41 | 1.75 | 1.40 | 1.34 | 1.83 | 1.00 | 1.20 | 1.32 | 1.42 | |
| Knapsack | 4 | 1.58 | 2.66 | 1.58 | 1.23 | 2.50 | 1.00 | 1.10 | 1.35 | 1.20 | |
| LCS | 3 | 1.22 | 1.59 | 1.17 | 1.23 | 1.00 | 1.00 | 1.03 | 1.30 | 1.45 | |
| LeapYear | 2 | 1.18 | 1.36 | 1.40 | 1.35 | 1.00 | 1.00 | 0.41 | 1.40 | 1.49 | |
| LinearSearch | 2 | 1.10 | 1.57 | 0.86 | 1.23 | 1.00 | 1.00 | 0.78 | 1.32 | 1.04 | |
| LinkedList | 4 | 1.60 | 1.00 | 1.40 | 1.36 | 0.86 | 1.00 | 1.07 | 1.67 | 4.48 | |
| LinqPipeline | 2 | 1.19 | 1.62 | 1.17 | 1.31 | 1.00 | 1.00 | 0.87 | 1.27 | 1.26 | |
| ListInvariant | 3 | 1.60 | 3.58 | 1.80 | 1.39 | 2.14 | 1.14 | 0.85 | 1.37 | 0.52 | |
| ListOps | 2 | 1.34 | 2.06 | 1.33 | 1.31 | 1.67 | 1.00 | 1.08 | 1.27 | 1.00 | |
| LogPipeline | 2 | 1.15 | 1.58 | 1.00 | 1.31 | 1.00 | 1.00 | 0.53 | 1.42 | 1.33 | |
| Map | 2 | 1.40 | 2.95 | 1.00 | 1.43 | 1.00 | 1.00 | 0.87 | 1.27 | 1.67 | |
| MathLib | 2 | 1.39 | 2.47 | 1.60 | 1.39 | 1.67 | 1.10 | 0.74 | 1.37 | 0.77 | |
| MathOperations | 3 | 1.32 | 0.82 | 1.60 | 1.35 | 1.11 | 1.00 | 0.91 | 1.37 | 2.40 | |
| MatrixMultiply | 3 | 1.42 | 1.95 | 1.33 | 1.34 | 1.57 | 1.00 | 1.30 | 1.30 | 1.57 | |
| MaxHeap | 3 | 1.48 | 1.63 | 1.17 | 1.36 | 1.57 | 1.00 | 1.28 | 1.43 | 2.44 | |
| MaxTwo | 1 | 1.26 | 1.70 | 1.40 | 1.31 | 1.67 | 1.00 | 0.76 | 1.30 | 0.95 | |
| MaxValue | 2 | 1.10 | 1.33 | 1.00 | 1.23 | 0.67 | 1.00 | 1.20 | 1.23 | 1.12 | |
| Mediator | 3 | 1.30 | 1.40 | 0.83 | 1.36 | 0.86 | 1.00 | 1.12 | 1.37 | 2.46 | |
| MergeSort | 3 | 1.52 | 1.44 | 1.33 | 1.35 | 1.83 | 1.00 | 0.97 | 1.37 | 2.88 | |
| MethodOverloading | 3 | 1.08 | 1.62 | 1.00 | 1.31 | 1.00 | 1.00 | 0.57 | 1.27 | 0.90 | |
| MethodOverriding | 3 | 1.25 | 1.56 | 1.00 | 1.43 | 1.00 | 1.00 | 1.55 | 1.32 | 1.13 | |
| MinStack | 3 | 1.31 | 1.39 | 1.17 | 1.61 | 0.86 | 1.00 | 1.01 | 1.32 | 2.12 | |
| MinTwo | 1 | 1.26 | 1.70 | 1.40 | 1.31 | 1.67 | 1.00 | 0.76 | 1.30 | 0.95 | |
| MissingEffects | 2 | 1.38 | 1.95 | 1.00 | 1.39 | 1.33 | 1.14 | 0.92 | 1.56 | 1.72 | |
| MixedContracts | 3 | 1.46 | 1.55 | 1.50 | 1.39 | 2.14 | 1.14 | 1.01 | 1.49 | 1.47 | |
| MixedSyntax | 2 | 1.40 | 2.11 | 1.33 | 1.31 | 2.14 | 1.00 | 0.79 | 1.32 | 1.23 | |
| ModuloContracts | 3 | 1.64 | 3.64 | 1.90 | 1.39 | 2.50 | 1.14 | 0.68 | 1.40 | 0.43 | |
| MultipleCatch | 2 | 1.14 | 1.27 | 1.09 | 1.35 | 1.00 | 1.00 | 0.73 | 1.37 | 1.28 | |
| NamedConfig | 2 | 1.12 | 1.25 | 1.00 | 1.31 | 1.00 | 1.00 | 0.62 | 1.42 | 1.32 | |
| NestedMatch | 2 | 1.31 | 1.86 | 1.20 | 1.31 | 1.00 | 1.00 | 1.43 | 1.30 | 1.35 | |
| NetworkEffect | 3 | 1.36 | 2.17 | 1.00 | 1.39 | 1.33 | 1.14 | 1.20 | 1.53 | 1.15 | |
| NullCheck | 2 | 1.10 | 1.66 | 0.80 | 1.31 | 0.86 | 1.00 | 0.70 | 1.27 | 1.16 | |
| Observer | 3 | 1.29 | 1.40 | 1.20 | 1.36 | 1.00 | 1.00 | 0.86 | 1.40 | 2.10 | |
| OptionalIds | 2 | 1.09 | 1.62 | 1.00 | 1.31 | 1.00 | 1.00 | 0.59 | 1.27 | 0.94 | |
| OptionType | 2 | 1.48 | 2.43 | 1.50 | 1.41 | 1.57 | 1.10 | 1.24 | 1.43 | 1.16 | |
| OverflowSafe | 3 | 1.51 | 1.46 | 1.80 | 1.35 | 2.14 | 1.03 | 1.29 | 1.40 | 1.63 | |
| OverflowUnsafe | 3 | 1.61 | 1.62 | 1.80 | 1.31 | 2.50 | 1.03 | 2.00 | 1.25 | 1.36 | |
| PairLogger | 2 | 1.11 | 1.42 | 1.00 | 1.31 | 1.00 | 1.00 | 0.47 | 1.42 | 1.23 | |
| Palindrome | 2 | 1.18 | 1.49 | 1.00 | 1.31 | 1.00 | 1.00 | 1.01 | 1.35 | 1.26 | |
| ParallelTasks | 3 | 1.29 | 2.16 | 1.00 | 1.43 | 1.00 | 1.00 | 1.06 | 1.30 | 1.38 | |
| ParseAndDouble | 2 | 1.14 | 1.29 | 1.00 | 1.31 | 1.00 | 1.00 | 0.66 | 1.44 | 1.42 | |
| PasswordValidator | 2 | 1.11 | 1.49 | 1.00 | 1.31 | 1.00 | 1.00 | 0.69 | 1.32 | 1.10 | |
| PhoneBook | 2 | 1.52 | 2.38 | 1.50 | 1.33 | 2.50 | 1.00 | 0.91 | 1.27 | 1.27 | |
| Polymorphism | 3 | 1.25 | 1.19 | 1.40 | 1.68 | 1.00 | 1.00 | 0.82 | 1.43 | 1.46 | |
| Power | 2 | 1.26 | 1.06 | 1.60 | 1.35 | 1.22 | 1.00 | 0.89 | 1.27 | 1.70 | |
| PriorityQueue | 3 | 1.54 | 1.71 | 1.33 | 1.36 | 1.57 | 1.00 | 1.05 | 1.49 | 2.77 | |
| Properties | 2 | 1.39 | 1.83 | 1.33 | 1.48 | 1.57 | 1.10 | 1.04 | 1.40 | 1.39 | |
| PropertyAccess | 1 | 1.83 | 1.63 | 1.60 | 1.48 | 1.83 | 1.10 | 4.25 | 1.43 | 1.34 | |
| PropertyMatch | 2 | 1.37 | 1.79 | 1.20 | 1.65 | 1.00 | 1.00 | 1.48 | 1.37 | 1.46 | |
| ProvableContracts | 2 | 1.48 | 2.10 | 1.58 | 1.39 | 2.14 | 1.14 | 0.98 | 1.49 | 1.01 | |
| Proxy | 3 | 1.30 | 1.41 | 1.20 | 1.68 | 1.14 | 1.00 | 0.88 | 1.53 | 1.59 | |
| PureComputation | 2 | 1.34 | 2.10 | 1.50 | 1.39 | 1.67 | 1.14 | 0.75 | 1.37 | 0.83 | |
| PureFunctions | 1 | 1.36 | 1.50 | 1.70 | 1.43 | 1.83 | 1.14 | 0.74 | 1.58 | 0.99 | |
| Queue | 3 | 1.30 | 1.22 | 1.17 | 1.36 | 0.86 | 1.00 | 1.09 | 1.40 | 2.29 | |
| QuickSort | 4 | 1.57 | 1.44 | 1.14 | 1.27 | 1.57 | 1.00 | 0.90 | 1.56 | 3.69 | |
| RangeContracts | 3 | 1.50 | 2.45 | 1.58 | 1.39 | 2.50 | 1.14 | 0.73 | 1.46 | 0.77 | |
| RangeMatch | 2 | 1.54 | 2.32 | 1.70 | 1.31 | 1.83 | 1.00 | 1.78 | 1.30 | 1.11 | |
| RecordType | 3 | 1.44 | 2.80 | 1.50 | 1.71 | 1.67 | 1.10 | 0.74 | 1.43 | 0.56 | |
| Reduce | 2 | 1.23 | 2.00 | 1.20 | 1.31 | 1.00 | 1.00 | 0.79 | 1.27 | 1.24 | |
| ResultType | 2 | 1.24 | 1.53 | 0.92 | 1.36 | 1.00 | 1.00 | 1.02 | 1.37 | 1.71 | |
| ReturnMapped | 2 | 1.13 | 1.25 | 1.00 | 1.31 | 1.00 | 1.00 | 0.64 | 1.42 | 1.45 | |
| ReverseArray | 2 | 1.53 | 2.87 | 1.33 | 1.31 | 1.83 | 1.00 | 1.59 | 1.32 | 1.02 | |
| ReverseString | 2 | 1.22 | 1.40 | 0.83 | 1.31 | 1.00 | 1.00 | 1.10 | 1.32 | 1.80 | |
| ScoreBoard | 2 | 1.57 | 2.37 | 1.80 | 1.33 | 2.50 | 1.00 | 0.95 | 1.27 | 1.34 | |
| SealedClass | 2 | 1.24 | 1.46 | 1.00 | 1.33 | 1.00 | 1.00 | 0.83 | 1.35 | 1.96 | |
| SearchContracts | 3 | 1.57 | 2.92 | 1.80 | 1.39 | 2.50 | 1.14 | 0.77 | 1.40 | 0.62 | |
| SelectionSort | 2 | 1.44 | 2.23 | 1.17 | 1.23 | 1.83 | 1.00 | 1.27 | 1.35 | 1.40 | |
| SetOps | 2 | 1.48 | 2.22 | 1.50 | 1.43 | 1.67 | 1.00 | 1.20 | 1.35 | 1.48 | |
| ShoppingCart | 4 | 1.71 | 1.90 | 1.19 | 1.28 | 2.14 | 1.00 | 1.06 | 1.93 | 3.23 | |
| Sign | 1 | 1.10 | 1.38 | 1.00 | 1.31 | 1.00 | 1.00 | 0.61 | 1.30 | 1.23 | |
| SimpleAsync | 3 | 1.33 | 1.96 | 1.00 | 1.31 | 1.33 | 1.00 | 1.34 | 1.39 | 1.34 | |
| SimpleClass | 2 | 1.63 | 2.36 | 1.80 | 1.65 | 2.50 | 1.00 | 0.88 | 1.43 | 1.40 | |
| SimpleMatch | 2 | 1.37 | 2.15 | 1.40 | 1.31 | 1.00 | 1.00 | 1.87 | 1.30 | 0.90 | |
| Singleton | 2 | 1.17 | 1.31 | 0.77 | 1.36 | 1.00 | 1.00 | 0.83 | 1.37 | 1.71 | |
| Sort | 2 | 1.16 | 1.58 | 0.83 | 1.23 | 1.00 | 1.00 | 0.78 | 1.30 | 1.55 | |
| SortedList | 2 | 1.37 | 1.65 | 1.25 | 1.59 | 1.67 | 1.00 | 0.97 | 1.30 | 1.50 | |
| SortingContracts | 3 | 1.56 | 3.07 | 1.80 | 1.39 | 2.50 | 1.14 | 0.55 | 1.37 | 0.68 | |
| Stack | 3 | 1.33 | 1.15 | 1.40 | 1.36 | 0.86 | 1.00 | 0.99 | 1.40 | 2.46 | |
| State | 3 | 1.44 | 2.05 | 1.40 | 1.36 | 1.00 | 1.00 | 2.07 | 1.37 | 1.24 | |
| StateEffect | 3 | 1.39 | 2.00 | 1.00 | 1.41 | 1.33 | 1.14 | 1.66 | 1.50 | 1.10 | |
| StaticMembers | 3 | 1.26 | 1.37 | 1.00 | 1.43 | 1.00 | 1.10 | 1.34 | 1.32 | 1.48 | |
| Strategy | 3 | 1.26 | 1.33 | 1.00 | 1.50 | 1.00 | 1.00 | 0.75 | 1.40 | 2.12 | |
| StringContracts | 3 | 1.37 | 1.45 | 1.29 | 1.31 | 2.14 | 1.03 | 1.01 | 1.32 | 1.42 | |
| StringLib | 2 | 1.32 | 2.33 | 1.25 | 1.39 | 1.67 | 1.10 | 0.65 | 1.37 | 0.80 | |
| StringUtils | 2 | 1.08 | 1.29 | 0.83 | 1.35 | 1.00 | 1.00 | 0.51 | 1.37 | 1.30 | |
| SumDigits | 2 | 1.55 | 2.87 | 1.90 | 1.31 | 2.50 | 1.00 | 0.95 | 1.44 | 0.42 | |
| SumRange | 2 | 1.31 | 1.48 | 1.17 | 1.23 | 1.57 | 1.00 | 0.95 | 1.34 | 1.71 | |
| SwitchExpression | 2 | 1.56 | 1.93 | 1.70 | 1.31 | 1.83 | 1.00 | 2.22 | 1.35 | 1.11 | |
| TaxCalculator | 3 | 1.42 | 2.28 | 1.58 | 1.31 | 2.50 | 1.00 | 0.65 | 1.30 | 0.73 | |
| TemperatureConverter | 1 | 1.11 | 1.28 | 1.00 | 1.31 | 1.00 | 1.00 | 0.63 | 1.35 | 1.29 | |
| TemperatureRange | 2 | 1.12 | 1.14 | 1.00 | 1.31 | 1.00 | 1.00 | 0.69 | 1.32 | 1.47 | |
| TemplateMethod | 3 | 1.00 | 0.37 | 1.10 | 0.94 | 0.86 | 0.76 | 0.73 | 1.83 | 1.44 | |
| TernaryChain | 2 | 1.19 | 1.61 | 1.40 | 1.31 | 1.00 | 1.00 | 1.02 | 1.27 | 0.89 | |
| ThreeWayMerge | 2 | 1.09 | 1.15 | 1.00 | 1.31 | 1.00 | 1.00 | 0.52 | 1.42 | 1.33 | |
| Timer | 2 | 1.53 | 2.57 | 1.50 | 1.36 | 2.14 | 1.00 | 1.19 | 1.40 | 1.11 | |
| TimerEffect | 3 | 1.59 | 3.13 | 1.90 | 1.39 | 2.50 | 1.10 | 0.69 | 1.40 | 0.66 | |
| TodoList | 2 | 1.54 | 2.06 | 1.50 | 1.59 | 2.14 | 1.00 | 1.23 | 1.32 | 1.44 | |
| TowerOfHanoi | 3 | 1.50 | 2.33 | 1.58 | 1.35 | 2.14 | 1.00 | 0.89 | 1.49 | 1.18 | |
| Trie | 4 | 1.62 | 1.80 | 1.39 | 1.36 | 2.50 | 1.00 | 1.11 | 1.46 | 2.31 | |
| TruncateString | 2 | 1.26 | 1.85 | 1.07 | 1.31 | 1.57 | 1.00 | 0.93 | 1.30 | 1.04 | |
| TryFinally | 2 | 1.18 | 1.48 | 1.00 | 1.31 | 1.00 | 1.00 | 1.01 | 1.35 | 1.28 | |
| TupleMatch | 2 | 1.24 | 1.91 | 1.40 | 1.31 | 1.00 | 1.00 | 1.07 | 1.30 | 0.91 | |
| TypeAlias | 3 | 1.45 | 2.56 | 1.17 | 1.41 | 1.38 | 1.10 | 1.38 | 1.30 | 1.34 | |
| TypeMatch | 2 | 1.41 | 1.75 | 1.40 | 1.35 | 1.00 | 1.00 | 2.14 | 1.37 | 1.27 | |
| UnitConverter | 2 | 1.35 | 2.37 | 1.60 | 1.31 | 1.83 | 1.00 | 0.69 | 1.27 | 0.69 | |
| UrlParser | 2 | 1.16 | 1.55 | 1.40 | 1.31 | 1.00 | 1.00 | 0.75 | 1.27 | 0.96 | |
| Visitor | 3 | 1.04 | 0.42 | 1.00 | 0.98 | 1.00 | 0.76 | 0.55 | 2.13 | 1.50 | |
| VoidSequence | 2 | 1.32 | 3.03 | 1.00 | 1.35 | 1.00 | 1.00 | 0.60 | 1.47 | 1.14 | |
| VotingSystem | 2 | 1.40 | 1.95 | 1.60 | 1.33 | 1.83 | 1.00 | 0.95 | 1.30 | 1.20 | |
| WildcardMatch | 2 | 1.20 | 1.84 | 1.40 | 1.31 | 1.00 | 1.00 | 0.93 | 1.30 | 0.80 | |
| Zip | 2 | 1.28 | 2.43 | 1.00 | 1.31 | 1.00 | 1.00 | 0.84 | 1.27 | 1.36 |
Showing 217 of 217 programs.
Read the Results Carefully
The static snapshot above covers 217 programs and reports a legacy composite direction-normalized ratio of 1.32x. This is not a measured language advantage. Comprehension is 1.84x, error detection is 1.49x, token economics is 1.42x, and information density is 0.97x—the category whose normalized ratio favors C#. Seven category ratios favor Calor and one favors C#. The Calor parser accepted 217 of 217 inputs; the Roslyn syntax parser accepted 217 of 217 C# inputs. These are parse checks, not evidence that the pairs compute the same result.
The snapshot was recorded on 2026-09-15 from source revision c2a8816d.
Its Directory.Build.props declares compiler 0.21.0, before the release
version bump; the JSON's version: "1.0" identifies the data format, not the
compiler release. The corpus is tests/TestData/Benchmarks. The exact compiler
binary used for the measurement was not independently recorded. These are
fixed-source results, rerun for the 0.22.0 release. They are not a live measure of
coding-agent performance. The separate agent-study snapshots below were not rerun.
These measurements need important caveats:
- Source pairs are not all behaviorally equivalent. For example, the C#
CsvParserimplements parsing, while its Calor counterpart contains only helper functions. The 0.19 rerun fixes an invalid newline literal in the C# fixture, but does not establish equal work across the corpus. The ratios therefore cannot establish a language advantage. See the benchmark integrity follow-up. - The metrics are deterministic static analyses. The exporter reports both zero-width intervals from singleton category aggregates and pooled intervals from repeated program scores. Thirty repetitions are not independent samples. Neither interval establishes sampling uncertainty over the corpus or a measured language, agent-productivity, correctness, or safety advantage.
- These C#-versus-Calor micro-benchmarks are not the v0.12 release gates. PP-A1 covered adoption readiness, while PP-W5 measured toolchain tax and concluded only that no large tax was detected—not that parity was proven.
Token economics is a composite of token, character, and line ratios. Calor
still pays a § token premium on small programs even though the current
composite favors Calor.
Separate Agent Studies
The static dashboard does not measure effect rows or live coding-agent behavior. Those questions use separate, pre-registered agent studies. The research evidence status explains why the proposed follow-on adopter study stopped unadjudicated.
Redesigned PP-W-rows: paused, not complete
Current status, 2026-09-11: paused by the user. The central pause handoff records two consumed invalid/censored attempts and 442 untouched scheduled identities out of 444. Neither the pilot nor confirmation has completed. These interrupted attempts do not supply an effect-row benefit, negative, or null result. They must not be erased, replaced, or described as an entirely unrun redesigned study.
The single total experiment ceiling is USD1,000. Two reservations of USD25.52 each remain held: USD51.04 total. Actual provider charges are unknown; the held amount is not measured spend or reconciled charges. The conditionally authorized permanent-retention exception has not been applied. Its implementation and final reviews remain unfinished in the accounting issue and paused draft checkpoint.
There is no permission to resume. A new explicit instruction and the remaining implementation, review, and operator-readiness gates are required. This product release neither resumes collection nor changes the protocol, consumed attempts, unknown costs, or existing ceiling. Confirmation would still need its own prospective design and affordability decision; it is not automatic.
Earlier decision, 2026-09-10: USD250 total, pilot only (superseded ceiling). The recorded user decision and historical authorization receipt accepted the registered stopping rules and publication of negative or null results under a conservative pilot-only interpretation. Its dated BUDGET_NOT_RUN assessment described an earlier pre-execution hold, not the current pause after consumed attempts. Historical planning estimates were not cost lower bounds. Neither that hold nor today's pause is a pilot null or the conditional stage-2 UNDERPOWERED-CARRIED outcome.
Historical M0 disposition, 2026-09-09: unchanged. The M0 decision records formal M0 as UNADJUDICATED, with a separate maintainer administrative stop. It deferred this narrower proposal pending a separate explicit decision. Later PP-W financial decisions do not re-arm M0, supply its missing human or participant approvals, or establish that effect rows help, fail, or cannot be studied.
Local buildability passed on 2026-09-10. The reviewed gate closure records Exit A under the frozen v0.18 release criteria. This deterministic engineering check is not an agent observation. Subsequent engineering work does not turn the two interrupted attempts into a completed study or benefit result.
The proposed methodology explains the old design defect: the unwanted behavior was already visibly forbidden, so avoiding it did not require effect enforcement. The redesigned question concerns violations hidden behind an abstraction while visible tests pass. Historical zero escapes do not answer that question.
The historical benefit ledger
records the old registration as UNDERPOWERED, epochRun: false, and
legA/legB: null. Null means uncollected for that historical registration,
not zero violations. It is not the record of the two consumed redesigned
attempts. This publication changes neither ledger nor scientific gate.
Later agent-result publication must follow actual adjudication by the scientific owners. The current pause is not a completed pilot, unaffordable confirmation verdict, or confirmatory result.
Historical PP-E1: effect-row toolchain cost
The 2026-08-27 PP-E1 report records two distinct legs, not one pooled sample:
- A seeded-fixture check caught 10 of 10 registered mutations at compiler
revision
121c6681, with frozen unmutated diagnostic baselines reproduced. These are hand-seeded cases, not agent trials or a population detection rate. - Four ordinary programming tasks ran five times on each of two compiler configurations, for 40 valid green runs, with no dropped task pairs or censored runs in the archived analysis. The median paired output-token ratio was 1.1835. Its one-sided 95% two-level cluster-bootstrap lower bound was 0.9012, below the registered 1.0 bound gate, and the point estimate stayed below the 1.35 failure threshold.
The epoch pins
record e1-rows-parity-001, started 2026-08-27, claude-opus-4-8,
Claude Code 2.1.243, and the v0.14.3 control at 63316987 versus the
0.15.0 release-build treatment at b775acb4. These are different compiler
versions, not the redesigned one-compiler policy contrast.
The combined registered verdict was HIT. The tasks measured compiler feedback cost on ordinary code; because they contained no callbacks, they did not measure the safety benefit of effect rows.
Historical PP-W-rows: practice rounds, not confirmation
The 2026-09-01 practice report
describes historical dry runs used for sizing, not a confirmatory verdict.
Both used claude-opus-4-8 with permissive v0.14.3-pre-rows control
283ec9f9 versus strict v0.15.0 treatment 3bb2601e.
| Archive | Date and agent | What can be read from it |
|---|---|---|
w-rows-dry-001 pins | 2026-08-28; Claude Code 2.1.248 | Usage-limit HTTP 429 runs were invalid. The report uses different “valid” and “readable” counts; this page does not pool them into a denominator. |
w-rows-dry-002 pins | 2026-09-01; Claude Code 2.1.252 | Three tasks, three runs per arm: 18 valid built runs, no exclusions. Zero observed escapes in either arm. This is a finite practice sample, not proof of a zero population escape rate. |
The second dry-run analysis reports a descriptive output-tokens-to-green ratio of 1.1175, with a one-sided 95% cluster-bootstrap lower bound of 0.8634, over its two cost-eligible task pairs (12 runs). It is not a result over all 18 runs or a measurement of general development effort.
Its old-design sizing required nine runs per task/arm, while the retired ceiling afforded three, with calculated power 0.48 under that registration. These are historical sizing assumptions, not the redesigned sample size or funding. The benefit ledger therefore remains UNDERPOWERED: the claim is neither supported nor refuted. The redesigned defect analysis also means buying more of the old tasks would not isolate the intended benefit.
The ledger must be read with care. Its top-level achievablePower field is
0.868, but the recorded reason and power curve use 0.868 for the smallest
unaffordable design, nine runs per cell. The largest affordable design was
three runs per cell at 0.48 power. The public interpretation therefore does
not present 0.868 as affordable power.
The practice result exposed a fixture-design defect: the undesirable hidden effect was already forbidden by the visible task specification. The 2026-09-08 redesign separates visible requirements from a held-out effect-observing test and adds eight task-qualification rules. At the 2026-09-09 M0 closeout, no candidate had established both R1 (a convenient effect-hiding API) and R4 (the laundering solution passes visible tests). The local buildability gate subsequently passed on 2026-09-10, as recorded above. The two later consumed attempts and the user-directed pause are separate from these practice results.
On 2026-09-09, M0 recorded these missing prerequisites and the lack of a
suitable isolated evaluator. Formal M0 status is UNADJUDICATED; the
maintainer action was an administrative stop. At that closeout, no redesigned
pilot or confirmatory collection was authorized. The old w-rows-001 epoch
remains unrun, not a null result; this does not erase the later redesigned attempts.
Keep these statements separate:
- The 1.32x dashboard is a deterministic source/structure comparison.
- The 0.15 proof point did not detect a large toolchain tax.
- The historical PP-W-rows registration did not run a confirmatory epoch.
- The redesigned study is paused after two invalid/censored attempts, with 442 untouched identities; neither stage has completed and no benefit or null result exists.
- The redesign and M0 records contain no protected three-arm adopter comparison, independent methods adjudication, or safety/economic advantage result.
See the artifact inventory, the proposed redesign, and the M0 evidence status for source links and execution states.
When to Use Calor
These scores are not sufficient evidence for choosing a language. Consider Calor when you want to experiment with:
- Explicit contracts and the documented limits of optional static verification
- Explicit IDs and structural markers for agent-editing experiments
- Structured source representations, while measuring costs on your own tasks
Use C# when:
- Ecosystem libraries are needed — Leverage existing tooling
- Human readability is priority — Familiar syntax for human developers
Methodology
The current static runner is automated and can regenerate a new artifact from the checked-out source. That is not the same as exactly reproducing every historical artifact: several omit the measurement binary, working-tree state, source revision, model, or runtime configuration. Proposed protocols have not run.
The published dashboard framework:
- Compiles paired Calor and C# programs.
- Applies deterministic metric calculators to source, syntax trees, and compilation results.
- Aggregates per-program direction-normalized ratios. Values above 1 favor Calor, including lower-is-better metrics whose raw score ratios are inverted.
- Repeats the deterministic evaluation 30 times for the reported statistical artifact. Greater than 1.0 favors Calor and less than 1.0 favors C#; neither is a measured language advantage.
LLM-based task runners exist separately and are not folded into this dashboard result.
See Methodology for full details on the evaluation framework.
Running Benchmarks
To regenerate benchmark results:
# Run the evaluation framework
dotnet run --project tests/Calor.Evaluation -c Release -- run -f website -o website/public/data/benchmark-results.json --statistical --runs 30
# Start the website to view results
cd website && npm run dev
Results are automatically loaded from benchmark-results.json and displayed in the dashboard above.