Benchmarking
The static dashboard, recorded on 2026-09-15 for v0.22.0, covers 217 source pairs
across eight deterministic metrics, with 30 repetitions. Its recorded
source revision c2a8816d declares compiler 0.21.0, before the release version
bump; the executed binary was
not independently recorded. It reports a 1.32x legacy direction-normalized composite.
The pairs are not all behaviorally equivalent, and the reported intervals do not
establish independent sampling uncertainty. These scores cannot establish a
measured language, agent-productivity, correctness, or safety advantage. Read the
snapshot provenance and limits.
The research evidence status records why
the 0.20 planning tranche stopped without a scientific verdict or software
release.
The redesigned PP-W-rows status is separate: the user paused the study on 2026-09-11 after two invalid/censored attempts. Neither the pilot nor confirmation has completed, and no benefit or null result exists. Historical practice rounds and static scores do not fill that gap. This product release does not authorize resuming research.
The Metrics
Published Metrics
The names below identify calculator rules, not observed agent abilities:
| Static metric | Calculator input or rule |
|---|---|
| Comprehension | Structural and semantic signals |
| Correctness | Static correctness signals |
| Edit Precision | Heuristic targetability and change isolation |
| Error Detection | Explicit detection signals |
| Generation Accuracy | Compilation and structural signals |
| Information Density | Counted semantic elements per token |
| Refactoring Stability | Reference preservation under modeled transformations |
| Token Economics | Composite of token, character, and line ratios |
The repository also contains separate LLM and estimation-mode suites for task completion, safety, and effect discipline. Their results are not included in the eight-metric headline or the 1.32x composite.
Direction-Normalized Metric Ratios
| Static metric | Direction-normalized ratio |
|---|---|
| Comprehension | 1.84x |
| Correctness | 1.29x |
| Edit Precision | 1.36x |
| Error Detection | 1.49x |
| Generation Accuracy | 1.02x |
| Information Density | 0.97x |
| Refactoring Stability | 1.38x |
| Token Economics | 1.42x |
Each metric is normalized so ratios above 1 favor Calor and ratios below 1 favor C#, including lower-is-better metrics whose raw score ratios are inverted. They are not measurements of language quality. The token economics composite must not be read as a raw-token saving. Use the dated source, method, and limitations when interpreting any row.
Agent Task Benchmark
The Agent Task Benchmark reports a separate historical prompt-to-code artifact from 2026-02-16. Its recorded summary is 77/89 across 17 categories, while its 18 category entries sum to 78/89. The exact model and runtime configuration are unknown. Its mixed syntax-pattern, compilation, and optional contract checks are not a systematic behavioral test, and its 80% gate is project-defined rather than calibrated reliability. In the historical run, all 89 tasks used Calor-to-C# transpilation and text-pattern scripts; zero enabled contract-verdict checks and zero behavioral executions. The snapshot is not part of the dashboard composite and does not establish an independent adopter handoff or a C#/protected-C#/Calor comparison.
Learn More
- Methodology - How benchmarks work
- Effect Rows Study - Registered method; research paused, no completed pilot
- Results - Detailed results table
- Agent Tasks - Claude code generation benchmark
- Research Evidence Status - Current study decisions and missing evidence
- Individual Metrics - Deep dive into each metric