v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Benchmarking

The static dashboard, recorded on 2026-09-15 for v0.22.0, covers 217 source pairs across eight deterministic metrics, with 30 repetitions. Its recorded source revision c2a8816d declares compiler 0.21.0, before the release version bump; the executed binary was not independently recorded. It reports a 1.32x legacy direction-normalized composite. The pairs are not all behaviorally equivalent, and the reported intervals do not establish independent sampling uncertainty. These scores cannot establish a measured language, agent-productivity, correctness, or safety advantage. Read the snapshot provenance and limits. The research evidence status records why the 0.20 planning tranche stopped without a scientific verdict or software release.

The redesigned PP-W-rows status is separate: the user paused the study on 2026-09-11 after two invalid/censored attempts. Neither the pilot nor confirmation has completed, and no benefit or null result exists. Historical practice rounds and static scores do not fill that gap. This product release does not authorize resuming research.


The Metrics

Published Metrics

The names below identify calculator rules, not observed agent abilities:

Static metricCalculator input or rule
ComprehensionStructural and semantic signals
CorrectnessStatic correctness signals
Edit PrecisionHeuristic targetability and change isolation
Error DetectionExplicit detection signals
Generation AccuracyCompilation and structural signals
Information DensityCounted semantic elements per token
Refactoring StabilityReference preservation under modeled transformations
Token EconomicsComposite of token, character, and line ratios

The repository also contains separate LLM and estimation-mode suites for task completion, safety, and effect discipline. Their results are not included in the eight-metric headline or the 1.32x composite.


Direction-Normalized Metric Ratios

Static metricDirection-normalized ratio
Comprehension1.84x
Correctness1.29x
Edit Precision1.36x
Error Detection1.49x
Generation Accuracy1.02x
Information Density0.97x
Refactoring Stability1.38x
Token Economics1.42x

Each metric is normalized so ratios above 1 favor Calor and ratios below 1 favor C#, including lower-is-better metrics whose raw score ratios are inverted. They are not measurements of language quality. The token economics composite must not be read as a raw-token saving. Use the dated source, method, and limitations when interpreting any row.


Agent Task Benchmark

The Agent Task Benchmark reports a separate historical prompt-to-code artifact from 2026-02-16. Its recorded summary is 77/89 across 17 categories, while its 18 category entries sum to 78/89. The exact model and runtime configuration are unknown. Its mixed syntax-pattern, compilation, and optional contract checks are not a systematic behavioral test, and its 80% gate is project-defined rather than calibrated reliability. In the historical run, all 89 tasks used Calor-to-C# transpilation and text-pattern scripts; zero enabled contract-verdict checks and zero behavioral executions. The snapshot is not part of the dashboard composite and does not establish an independent adopter handoff or a C#/protected-C#/Calor comparison.


Learn More