v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Methodology

What Calor's benchmark runners calculate, which artifacts actually ran, and which questions remain proposals.

The proposed Effect Rows Study has a separate two-stage protocol. The user paused it on 2026-09-11 after two invalid/censored attempts; neither stage has completed. It is not this static evaluation framework or the outcome-based Effect Discipline runner.


Evaluation Approach

The repository contains several separate benchmark runners. They do not form one comprehensive assessment, and their results must not be combined unless a protocol explicitly registers that combination.

The published static dashboard is narrower than the full framework: it uses 217 paired programs and eight deterministic metrics across 30 repetitions. Optional LLM suites are separate artifacts and do not contribute to the 1.32x headline.

Recorded on 2026-09-15 for release v0.22.0, its source is c2a8816d, which declares compiler 0.21.0 before the release version bump. The exact executed measurement binary was not independently recorded. The corpus is tests/TestData/Benchmarks. See the snapshot provenance and limits.

Published artifact inventory

This table separates checked-in observations from runner capabilities and proposed methods. "Unknown" means the artifact does not independently record the field.

ArtifactDate and sourceDenominatorExecution and configurationHeadline use
Eight-metric static dashboard2026-09-15; source c2a8816d, declaring compiler 0.21.0 before the v0.22.0 bump217 source pairsDeterministic calculators repeated 30 times. No LLM. No independently attested measurement binary.Supplies the legacy 1.32x composite. It establishes no measured language, agent-productivity, correctness, or safety advantage.
Correctness estimation mode inside the static dashboardSame artifact and 217 pairs217 source pairsHistorical heuristic scoring because no LLM provider was configured. It does not execute generated programs as an agent study.Included as one of the eight legacy dashboard metrics.
Historical Safety estimation2026-02-17; artifact revision e02684e1, recording source 59799d140 programsHistorical heuristic estimate; no agent execution. Current provider-backed runners do not silently fall back to it.Supplies the separate historical 1.59x Safety headline; not part of the 1.32x dashboard.
Historical Effect Discipline estimation2026-02-17; artifact revision e02684e1, recording source 59799d140 programsHistorical structural estimate; no agent execution. Current provider-backed runners do not silently fall back to it.Supplies a separate historical 1.00x tie; not part of the 1.32x dashboard.
Checked-in 50-task LLM artifact2026-02-21; source revision absent50 tasksProvider claude; model is null. All calls recorded authentication failures, producing zero valid generations and zero compilation/test observations.Does not contribute to any headline. Its numerical 1.00 ratio is only 0 divided against 0 by the exporter and is not a parity result.
Safety provider artifact2026-02-21; source revision absent30 tasks, 60 language callsProvider claude; model is null. All 60 calls recorded authentication failures, so there are no valid generation, compilation, test, violation-detection, or error-quality observations.Does not support the separate historical 1.59x estimation headline or any other headline.
Effect Discipline provider artifact2026-02-21; source revision absent40 tasks, 80 language callsProvider claude; model is null. All 80 calls recorded authentication failures, so there are no valid generation, compilation, test, or bug-prevention observations.Does not contribute to any headline.
Agent task snapshot2026-02-16; source 107462e, declaring compiler 0.2.389 tasks. The recorded summary says 77/89 across 17 categories; the 18 category entries total 78/89.Historical single-run Claude Code CLI snapshot. Exact model, CLI version, executed compiler binary, working-tree state, and host configuration are unknown. The runner skipped calor init and its lifecycle hooks. All 89 tasks required Calor-to-C# transpilation and an all-grep text-pattern script. Zero enabled contract-verdict checks and zero behavioral executions.Preserves the unreconciled 86.5% summary headline and project-defined 80% gate. The entries imply 87.6%. Neither rate is calibrated reliability, part of the static composite, an independent adopter handoff, or a C#/protected-C#/Calor comparison. The old generator was replaced by a fail-closed archival stub.
Agent refactoring snapshot2026-02-15; source 580e189, declaring compiler 0.2.320 tasks per language, three runs per taskclaude-code, exact model unknown; majority vote requires two of three runs.Supplies the separate public 95% snapshot headline for each language; not part of the static composite.
PP-E1 effect-row proof point2026-08-27; artifact revision 3bb2601e3ff83597ddf2f27dcc334b6399ab97ea, measured release commit 72d060f400b5787ba4ec11cb4dd8860c95173da310 seeded mutations plus 40 ordinary-task runsclaude-opus-4-8, Claude Code 2.1.243; v0.14.3 control versus v0.15.0 treatment, five runs per arm. Verdict HIT for its bounded toolchain-cost claim.Separate proof point; no contribution to 1.32x.
PP-W-rows practice and sizing2026-09-01; artifact revision 82a7c653cbf1ea2f6231e38cc328c74a34e6589b. Its measuredCommit value 3bb2601e0cbd93fc25fdaaf2a0ea5183b8a2dd6a does not resolve; the pins record arm commits 283ec9f9964ddd5b21da15b646a0dd77d53de99e and 3bb2601e3ff83597ddf2f27dcc334b6399ab97ea.28 valid practice runs; registered main epoch had zero runsclaude-opus-4-8, Claude Code 2.1.248 and 2.1.252; v0.14.3 pre-rows/permissive control versus v0.15.0 strict treatment, three runs per arm. The main w-rows-001 epoch never ran; verdict UNDERPOWERED.No effect-row safety-benefit headline.
Redesigned PP-W-rows protocol2026-09-08; revision 16880d006db760d7b47d829fd4282b033c3158a0Registered method, not collected observationsLocal buildability subsequently passed on 2026-09-10. This document records the design rules; later attempts are tracked separately below.No agent-benefit result follows from registration or buildability.
Paused redesigned PP-W-rows study2026-09-11 user-directed pause; draft checkpointTwo consumed invalid/censored attempts; 442 untouched identities out of 444Single USD1,000 total ceiling. Two reservations hold USD51.04 total; actual charges are unknown, not measured spend. Permanent-retention exception not applied. No permission to resume.Neither pilot nor confirmation completed; no benefit or null result. This is not the historical ledger's unrun epoch.

The Results page explains the PP-W chronology and the Effect Rows Study records the redesign gates.

Statistical Rigor

All metrics support statistical analysis mode:

  • Multiple runs (default: n=30) for variance measurement
  • The exporter labels two different calculations as 95% confidence intervals: singleton category aggregates produce zero-width intervals, while pooled program scores include all 30 deterministic repetitions. Neither establishes sampling uncertainty over the program corpus; repeated observations are not independent samples.
  • Cohen's d effect size calculations
  • Paired t-tests for statistical significance (p < 0.05)
Bash
# Run with statistical analysis
dotnet run --project tests/Calor.Evaluation -c Release -- run --statistical --runs 30

Test Corpus

Published Snapshot Scale

  • 217 programs across 14 manifest categories
  • Each program has paired Calor (.calr) and C# (.cs) implementations
  • Complexity levels 1-4 in the published corpus: 14 level-1, 112 level-2, 84 level-3, and 7 level-4 programs

Categories

Manifest categoryCount
TokenEconomics52
DataStructures15
DesignPatterns17
DomainProblems20
ComplexAlgorithms15
AsyncConcurrency10
CollectionsLINQ11
CompactSyntax5
ContractVerification15
ErrorHandling10
EffectSoundness10
InteropEffectCoverage5
OOPGenerics21
PatternMatching11
Total217

Intended Requirements

For a meaningful comparison, both implementations must:

  1. Compile successfully
  2. Produce identical output for the same inputs
  3. Have the same logical structure

The published source pairs are not all behaviorally equivalent. CsvParser is a concrete counterexample: C# implements parsing, while Calor supplies only helper functions. The current scores do not establish a language advantage. Pair equivalence and interval calculations are tracked in issue #1276.


The Metrics

1. Token Economics

Measures: Tokens required to represent equivalent logic.

Method:

  • Simple tokenization (split on whitespace and punctuation)
  • Character count (excluding whitespace)
  • Line count
  • Composite ratio

Interpretation: Lower is better (less context window usage).


2. Generation Accuracy

Measures: Ability to generate valid code.

Factors:

  • Compilation success (50%)
  • Structural completeness (30%)
  • Error count (20%)

Structural completeness:

  • Calor: Module, functions, bodies present
  • C#: Namespace, class, methods present

3. Comprehension

Measures: How easily an agent can understand code structure.

Calor factors:

  • Module declarations (§M{)
  • Function declarations (§F{)
  • Input/output annotations (§I{, §O{)
  • Effect declarations (§E{)
  • Contracts (§Q, §S)
  • Indented body lines (the structural-completeness signal)

C# factors:

  • Namespace declarations
  • Class declarations
  • Documentation comments (///)
  • Type annotations
  • Contract patterns

Scoring: Weighted sum of factors present, normalized to 0-1.


4. Edit Precision

Measures: Ability to target specific code elements accurately.

Approach: Synthesizes edit-task descriptions and scores targetability with source-pattern heuristics; it does not execute an agent or apply the edits:

  • Change loop bounds by ID
  • Add preconditions to functions
  • Rename functions (ID vs name-based)
  • Change function signatures
  • Modify return types

Calor advantages:

  • Explicit IDs can enable precise targeting (§F{f001:)
  • Indentation defines block boundaries
  • ID-based references don't break on rename

C# challenges:

  • Name collisions reduce targeting accuracy
  • Brace nesting creates ambiguity
  • Cascading changes needed for renames

Scoring: 40% structural analysis + 60% heuristic simulated-task score. The current implementation still contains legacy §V/§LOOP detectors, so current §B/§L syntax does not receive those particular bonuses. Treat this as a static representation metric, not observed editing accuracy.


5. Error Detection

Measures: Bug detection capability through explicit contracts.

Calor factors:

  • Preconditions (§Q) - +0.25
  • Postconditions (§S) - +0.20
  • Invariants - the current calculator still detects legacy §INV, not current §IV; this is a known measurement limitation
  • Effect declarations - +0.10

C# factors:

  • Debug.Assert statements
  • Contract.Requires / Contract.Ensures
  • Null checks
  • Exception handling

6. Information Density

Measures: Semantic content per token.

Semantic elements counted:

  • Calor: Modules, functions, variables, type annotations, contracts, effects, control flow, expressions
  • C#: Namespaces, classes, methods, variables, type annotations, control flow, expressions

Density formula: Total semantic elements / token count


7. Refactoring Stability

Measures: How well explicit IDs are preserved in modeled code transformations.

Scenarios tested:

ScenarioWhat's Measured
Rename functionDoes ID survive when name changes?
Extract methodIs new ID assigned, original preserved?
Move functionDoes ID survive cross-module move?
Change signatureDo callers update correctly?
Inline variableDo references resolve correctly?

Scoring weights:

  • ID preservation: 30%
  • Reference validity after edit: 25%
  • Minimal diff size: 20%
  • Semantic equivalence: 25%

LLM Runner Capabilities

The repository implements optional LLM runners. A runner's presence is not evidence that its proposed protocol ran successfully.

8. Task Completion

Measures: LLM code generation success using language-neutral prompts.

Implemented method:

  • Read task manifests
  • Request code from a configured provider
  • Compile generated code and run task tests
  • Score compilation and test outcomes

The scoring configuration is manifest-driven. See the 50-Task LLM Runner section below for the checked-in artifact's actual weights.


9. Safety

Measures: Contract enforcement effectiveness and error quality for catching bugs.

Current implementation: Requires a configured provider and task manifest. The checked-in 1.59x figure came from a historical estimation path; it is not a current provider fallback or an observed bug-detection rate.

What it catches:

  • Division by zero
  • Array bounds violations
  • Integer overflow
  • Null dereferences
  • Invalid argument values

10. Effect Discipline

Measures: Side effect management quality and bug prevention.

Current implementation: Requires a configured provider and task manifest. The separate structural estimate is a historical artifact, not a current provider fallback. Neither is the PP-W-rows study or evidence that effect rows prevent more bugs in agent-produced code.

What it catches:

  • Flaky tests (non-determinism from hidden state)
  • Security violations (unauthorized I/O)
  • Side effect transparency issues
  • Cache safety problems (memoization correctness)

Calor-Only Metrics

11. Interop Effect Coverage

Measures: BCL methods covered by effect manifests.

The BCL effect manifest tracks which .NET methods have effects, enabling verification even when calling external code. This has no C# equivalent since C# lacks an effect system.


These calculators describe source features and runner outputs. They do not explain or establish an overall Calor advantage.


50-Task LLM Runner

The task-completion runner has a 50-task manifest across four categories. The checked-in website/public/data/llm-results.json is not a successful benchmark: every provider call failed authentication, the recorded model is unknown, and no generated program compiled or ran.

Design Philosophy: Neutral Prompts

The benchmark uses the same functional requirements for both languages without syntax hints:

Plain Text
Write a public function named Factorial that computes the factorial
of an integer n. The function must only accept non-negative values
of n. The result is always at least 1.

This is the intended comparison. It is not an observed result from the checked-in artifact.

How It Works

  1. Load 50 programming tasks across four categories.
  2. Submit the configured prompt for each language.
  3. Compile generated code and run the task tests.
  4. Compute scores from compilation and test outcomes.

The current JSON records usesNeutralPrompt: false for its tasks, so it cannot support the neutral-prompt claim previously made on this page.

Task Categories

The checked-in failed artifact used these categories:

CategoryCountExamples
basic-algorithms15Factorial, Fibonacci, IsPrime, GCD, Power
contracts10SafeDivide, Clamp, SafeModulo, NormalizeScore
data-structures10Sum, Max, Min, Average, Median
logic15BoolToInt, LogicalAnd, IsMultipleOf, SameSign

The separate task-manifest-neutral.json uses a safety category instead of contracts. It was not the manifest serialized into the checked-in artifact.

Scoring Formula

Plain Text
Score = compiled × compilation weight
      + test pass rate × test weight
      + contract result × contract weight

The checked-in artifact records either 30% compilation, 50% tests, and 20% contracts, or 30% compilation, 40% tests, and 30% contracts. Contract weight is applied only to Calor results. If Calor compiles without an explicit contract result, the runner awards half of the configured contract weight. The neutral manifest also defaults to a 20% contract weight when it omits the field. A future language-neutral study must register and justify its scoring before execution rather than describe this implementation as a 40/60 comparison.

Checked-in artifact status

FieldRecorded value
Timestamp2026-02-21
Source revisionNot recorded
Provider / modelclaude / unknown
Tasks50
Valid generations0
Compilation and test observations0
Exported ratio1.00 from equal zero scores; not interpretable as parity

Running Locally

Bash
# Run LLM benchmark with neutral prompts
dotnet run --project tests/Calor.Evaluation -- llm-tasks \
  --manifest tests/Calor.Evaluation/Tasks/task-manifest-neutral.json \
  --verbose

# Refresh cache (re-run all tasks)
dotnet run --project tests/Calor.Evaluation -- llm-tasks \
  --manifest tests/Calor.Evaluation/Tasks/task-manifest-neutral.json \
  --refresh-cache

Runner controls

  • Caching: Results cached to avoid redundant API calls
  • Budget caps: A configured run can enforce a spend limit
  • Estimated cost: The runner can display an estimate before API calls

Proposed LLM Evaluation

Multi-model validation, comprehension questions, repeated runs, and cross-model agreement are proposed methodology. No checked-in artifact on this page records an executed Claude 3.5 Sonnet/GPT-4o cross-validation.

Proposed comprehension questions

LLMs answer questions about code understanding:

  1. What is the main purpose of this code?
  2. What are the input parameters and their constraints?
  3. What does this function return?
  4. What side effects does this code have?
  5. What would happen if [edge case]?

Proposed scoring

  • Correctness of answers (0-1) graded against ground truth
  • Tokens used to formulate answer
  • Consistency across multiple runs
  • Cross-model agreement (higher = more reliable)

Running the Evaluation

Basic Run

Bash
# JSON output (default)
dotnet run --project tests/Calor.Evaluation -- run --output report.json

# Markdown output
dotnet run --project tests/Calor.Evaluation -- run --format markdown --output report.md

# Website dashboard format
dotnet run --project tests/Calor.Evaluation -- run --format website --output results.json

# HTML dashboard
dotnet run --project tests/Calor.Evaluation -- run --format html --output dashboard.html

Statistical Analysis

Bash
# Run with 30 statistical samples
dotnet run --project tests/Calor.Evaluation -- run --statistical --runs 30

# Custom run count
dotnet run --project tests/Calor.Evaluation -- run --statistical --runs 50

Specific Metrics

Bash
# Run only specific categories
dotnet run --project tests/Calor.Evaluation -- run \
  --category Comprehension \
  --category EditPrecision \
  --category RefactoringStability

Interpreting Results

Direction Normalization

For each metric:

  • Higher is better (Comprehension, Error Detection, etc.): ratio = Calor/C#.
  • Lower is better (Token Economics): ratio = C#/Calor.

After normalization, a ratio above 1 favors Calor and a ratio below 1 favors C#. The ratio does not mean that Calor's raw score is higher.

Statistical Significance

With statistical mode enabled:

  • p < 0.05: Result is statistically significant
  • Cohen's d: Effect size interpretation
    • d < 0.2: Negligible
    • d < 0.5: Small
    • d < 0.8: Medium
    • d ≥ 0.8: Large

The Tradeoff

The legacy static snapshot scores Calor higher on several structural calculators and C# higher on information density. Because the pairs are not all behaviorally equivalent and the metrics are not agent outcomes, this does not establish better agent reasoning, safer refactoring, or a language advantage.


CI/CD Integration

Automated Benchmarks

Repository workflows may run selected benchmark and regression commands. Workflow availability does not establish a completed weekly multi-model LLM study, PP-W pilot, or confirmation. The paused PP-W attempts are recorded separately in the inventory above.

Regression Detection

CI checks implementation and artifact invariants defined by the current workflow. It does not independently adjudicate scientific validity.


Limitations

  1. Corpus size: 217 paired programs still may not cover all patterns
  2. C# baseline: Other languages might perform differently
  3. Static analysis: Most headline metrics do not capture runtime behavior
  4. Artifact provenance: Several historical artifacts omit an exact model, runtime configuration, measurement binary, or source revision
  5. LLM variance: Model updates can affect evaluation scores
  6. No adopter comparison: There is no protected three-arm adopter study, independent methods adjudication, or measured safety/economic advantage

Next