Methodology
What Calor's benchmark runners calculate, which artifacts actually ran, and which questions remain proposals.
The proposed Effect Rows Study has a separate two-stage protocol. The user paused it on 2026-09-11 after two invalid/censored attempts; neither stage has completed. It is not this static evaluation framework or the outcome-based Effect Discipline runner.
Evaluation Approach
The repository contains several separate benchmark runners. They do not form one comprehensive assessment, and their results must not be combined unless a protocol explicitly registers that combination.
The published static dashboard is narrower than the full framework: it uses 217 paired programs and eight deterministic metrics across 30 repetitions. Optional LLM suites are separate artifacts and do not contribute to the 1.32x headline.
Recorded on 2026-09-15 for release v0.22.0, its source is c2a8816d, which
declares compiler 0.21.0 before the release version bump. The exact executed
measurement binary was not independently recorded. The corpus is
tests/TestData/Benchmarks. See the snapshot provenance and limits.
Published artifact inventory
This table separates checked-in observations from runner capabilities and proposed methods. "Unknown" means the artifact does not independently record the field.
| Artifact | Date and source | Denominator | Execution and configuration | Headline use |
|---|---|---|---|---|
| Eight-metric static dashboard | 2026-09-15; source c2a8816d, declaring compiler 0.21.0 before the v0.22.0 bump | 217 source pairs | Deterministic calculators repeated 30 times. No LLM. No independently attested measurement binary. | Supplies the legacy 1.32x composite. It establishes no measured language, agent-productivity, correctness, or safety advantage. |
| Correctness estimation mode inside the static dashboard | Same artifact and 217 pairs | 217 source pairs | Historical heuristic scoring because no LLM provider was configured. It does not execute generated programs as an agent study. | Included as one of the eight legacy dashboard metrics. |
| Historical Safety estimation | 2026-02-17; artifact revision e02684e1, recording source 59799d1 | 40 programs | Historical heuristic estimate; no agent execution. Current provider-backed runners do not silently fall back to it. | Supplies the separate historical 1.59x Safety headline; not part of the 1.32x dashboard. |
| Historical Effect Discipline estimation | 2026-02-17; artifact revision e02684e1, recording source 59799d1 | 40 programs | Historical structural estimate; no agent execution. Current provider-backed runners do not silently fall back to it. | Supplies a separate historical 1.00x tie; not part of the 1.32x dashboard. |
| Checked-in 50-task LLM artifact | 2026-02-21; source revision absent | 50 tasks | Provider claude; model is null. All calls recorded authentication failures, producing zero valid generations and zero compilation/test observations. | Does not contribute to any headline. Its numerical 1.00 ratio is only 0 divided against 0 by the exporter and is not a parity result. |
| Safety provider artifact | 2026-02-21; source revision absent | 30 tasks, 60 language calls | Provider claude; model is null. All 60 calls recorded authentication failures, so there are no valid generation, compilation, test, violation-detection, or error-quality observations. | Does not support the separate historical 1.59x estimation headline or any other headline. |
| Effect Discipline provider artifact | 2026-02-21; source revision absent | 40 tasks, 80 language calls | Provider claude; model is null. All 80 calls recorded authentication failures, so there are no valid generation, compilation, test, or bug-prevention observations. | Does not contribute to any headline. |
| Agent task snapshot | 2026-02-16; source 107462e, declaring compiler 0.2.3 | 89 tasks. The recorded summary says 77/89 across 17 categories; the 18 category entries total 78/89. | Historical single-run Claude Code CLI snapshot. Exact model, CLI version, executed compiler binary, working-tree state, and host configuration are unknown. The runner skipped calor init and its lifecycle hooks. All 89 tasks required Calor-to-C# transpilation and an all-grep text-pattern script. Zero enabled contract-verdict checks and zero behavioral executions. | Preserves the unreconciled 86.5% summary headline and project-defined 80% gate. The entries imply 87.6%. Neither rate is calibrated reliability, part of the static composite, an independent adopter handoff, or a C#/protected-C#/Calor comparison. The old generator was replaced by a fail-closed archival stub. |
| Agent refactoring snapshot | 2026-02-15; source 580e189, declaring compiler 0.2.3 | 20 tasks per language, three runs per task | claude-code, exact model unknown; majority vote requires two of three runs. | Supplies the separate public 95% snapshot headline for each language; not part of the static composite. |
| PP-E1 effect-row proof point | 2026-08-27; artifact revision 3bb2601e3ff83597ddf2f27dcc334b6399ab97ea, measured release commit 72d060f400b5787ba4ec11cb4dd8860c95173da3 | 10 seeded mutations plus 40 ordinary-task runs | claude-opus-4-8, Claude Code 2.1.243; v0.14.3 control versus v0.15.0 treatment, five runs per arm. Verdict HIT for its bounded toolchain-cost claim. | Separate proof point; no contribution to 1.32x. |
| PP-W-rows practice and sizing | 2026-09-01; artifact revision 82a7c653cbf1ea2f6231e38cc328c74a34e6589b. Its measuredCommit value 3bb2601e0cbd93fc25fdaaf2a0ea5183b8a2dd6a does not resolve; the pins record arm commits 283ec9f9964ddd5b21da15b646a0dd77d53de99e and 3bb2601e3ff83597ddf2f27dcc334b6399ab97ea. | 28 valid practice runs; registered main epoch had zero runs | claude-opus-4-8, Claude Code 2.1.248 and 2.1.252; v0.14.3 pre-rows/permissive control versus v0.15.0 strict treatment, three runs per arm. The main w-rows-001 epoch never ran; verdict UNDERPOWERED. | No effect-row safety-benefit headline. |
| Redesigned PP-W-rows protocol | 2026-09-08; revision 16880d006db760d7b47d829fd4282b033c3158a0 | Registered method, not collected observations | Local buildability subsequently passed on 2026-09-10. This document records the design rules; later attempts are tracked separately below. | No agent-benefit result follows from registration or buildability. |
| Paused redesigned PP-W-rows study | 2026-09-11 user-directed pause; draft checkpoint | Two consumed invalid/censored attempts; 442 untouched identities out of 444 | Single USD1,000 total ceiling. Two reservations hold USD51.04 total; actual charges are unknown, not measured spend. Permanent-retention exception not applied. No permission to resume. | Neither pilot nor confirmation completed; no benefit or null result. This is not the historical ledger's unrun epoch. |
The Results page explains the PP-W chronology and the Effect Rows Study records the redesign gates.
Statistical Rigor
All metrics support statistical analysis mode:
- Multiple runs (default: n=30) for variance measurement
- The exporter labels two different calculations as 95% confidence intervals: singleton category aggregates produce zero-width intervals, while pooled program scores include all 30 deterministic repetitions. Neither establishes sampling uncertainty over the program corpus; repeated observations are not independent samples.
- Cohen's d effect size calculations
- Paired t-tests for statistical significance (p < 0.05)
# Run with statistical analysis
dotnet run --project tests/Calor.Evaluation -c Release -- run --statistical --runs 30
Test Corpus
Published Snapshot Scale
- 217 programs across 14 manifest categories
- Each program has paired Calor (.calr) and C# (.cs) implementations
- Complexity levels 1-4 in the published corpus: 14 level-1, 112 level-2, 84 level-3, and 7 level-4 programs
Categories
| Manifest category | Count |
|---|---|
| TokenEconomics | 52 |
| DataStructures | 15 |
| DesignPatterns | 17 |
| DomainProblems | 20 |
| ComplexAlgorithms | 15 |
| AsyncConcurrency | 10 |
| CollectionsLINQ | 11 |
| CompactSyntax | 5 |
| ContractVerification | 15 |
| ErrorHandling | 10 |
| EffectSoundness | 10 |
| InteropEffectCoverage | 5 |
| OOPGenerics | 21 |
| PatternMatching | 11 |
| Total | 217 |
Intended Requirements
For a meaningful comparison, both implementations must:
- Compile successfully
- Produce identical output for the same inputs
- Have the same logical structure
The published source pairs are not all behaviorally equivalent. CsvParser
is a concrete counterexample: C# implements parsing, while Calor supplies only
helper functions. The current scores do not establish a language advantage.
Pair equivalence and interval calculations are tracked in
issue #1276.
The Metrics
1. Token Economics
Measures: Tokens required to represent equivalent logic.
Method:
- Simple tokenization (split on whitespace and punctuation)
- Character count (excluding whitespace)
- Line count
- Composite ratio
Interpretation: Lower is better (less context window usage).
2. Generation Accuracy
Measures: Ability to generate valid code.
Factors:
- Compilation success (50%)
- Structural completeness (30%)
- Error count (20%)
Structural completeness:
- Calor: Module, functions, bodies present
- C#: Namespace, class, methods present
3. Comprehension
Measures: How easily an agent can understand code structure.
Calor factors:
- Module declarations (
§M{) - Function declarations (
§F{) - Input/output annotations (
§I{,§O{) - Effect declarations (
§E{) - Contracts (
§Q,§S) - Indented body lines (the structural-completeness signal)
C# factors:
- Namespace declarations
- Class declarations
- Documentation comments (
///) - Type annotations
- Contract patterns
Scoring: Weighted sum of factors present, normalized to 0-1.
4. Edit Precision
Measures: Ability to target specific code elements accurately.
Approach: Synthesizes edit-task descriptions and scores targetability with source-pattern heuristics; it does not execute an agent or apply the edits:
- Change loop bounds by ID
- Add preconditions to functions
- Rename functions (ID vs name-based)
- Change function signatures
- Modify return types
Calor advantages:
- Explicit IDs can enable precise targeting (
§F{f001:) - Indentation defines block boundaries
- ID-based references don't break on rename
C# challenges:
- Name collisions reduce targeting accuracy
- Brace nesting creates ambiguity
- Cascading changes needed for renames
Scoring: 40% structural analysis + 60% heuristic simulated-task score.
The current implementation still contains legacy §V/§LOOP detectors, so
current §B/§L syntax does not receive those particular bonuses. Treat this
as a static representation metric, not observed editing accuracy.
5. Error Detection
Measures: Bug detection capability through explicit contracts.
Calor factors:
- Preconditions (
§Q) - +0.25 - Postconditions (
§S) - +0.20 - Invariants - the current calculator still detects legacy
§INV, not current§IV; this is a known measurement limitation - Effect declarations - +0.10
C# factors:
Debug.AssertstatementsContract.Requires/Contract.Ensures- Null checks
- Exception handling
6. Information Density
Measures: Semantic content per token.
Semantic elements counted:
- Calor: Modules, functions, variables, type annotations, contracts, effects, control flow, expressions
- C#: Namespaces, classes, methods, variables, type annotations, control flow, expressions
Density formula: Total semantic elements / token count
7. Refactoring Stability
Measures: How well explicit IDs are preserved in modeled code transformations.
Scenarios tested:
| Scenario | What's Measured |
|---|---|
| Rename function | Does ID survive when name changes? |
| Extract method | Is new ID assigned, original preserved? |
| Move function | Does ID survive cross-module move? |
| Change signature | Do callers update correctly? |
| Inline variable | Do references resolve correctly? |
Scoring weights:
- ID preservation: 30%
- Reference validity after edit: 25%
- Minimal diff size: 20%
- Semantic equivalence: 25%
LLM Runner Capabilities
The repository implements optional LLM runners. A runner's presence is not evidence that its proposed protocol ran successfully.
8. Task Completion
Measures: LLM code generation success using language-neutral prompts.
Implemented method:
- Read task manifests
- Request code from a configured provider
- Compile generated code and run task tests
- Score compilation and test outcomes
The scoring configuration is manifest-driven. See the 50-Task LLM Runner section below for the checked-in artifact's actual weights.
9. Safety
Measures: Contract enforcement effectiveness and error quality for catching bugs.
Current implementation: Requires a configured provider and task manifest. The checked-in 1.59x figure came from a historical estimation path; it is not a current provider fallback or an observed bug-detection rate.
What it catches:
- Division by zero
- Array bounds violations
- Integer overflow
- Null dereferences
- Invalid argument values
10. Effect Discipline
Measures: Side effect management quality and bug prevention.
Current implementation: Requires a configured provider and task manifest. The separate structural estimate is a historical artifact, not a current provider fallback. Neither is the PP-W-rows study or evidence that effect rows prevent more bugs in agent-produced code.
What it catches:
- Flaky tests (non-determinism from hidden state)
- Security violations (unauthorized I/O)
- Side effect transparency issues
- Cache safety problems (memoization correctness)
Calor-Only Metrics
11. Interop Effect Coverage
Measures: BCL methods covered by effect manifests.
The BCL effect manifest tracks which .NET methods have effects, enabling verification even when calling external code. This has no C# equivalent since C# lacks an effect system.
These calculators describe source features and runner outputs. They do not explain or establish an overall Calor advantage.
50-Task LLM Runner
The task-completion runner has a 50-task manifest across four categories. The
checked-in website/public/data/llm-results.json is not a successful benchmark:
every provider call failed authentication, the recorded model is unknown, and
no generated program compiled or ran.
Design Philosophy: Neutral Prompts
The benchmark uses the same functional requirements for both languages without syntax hints:
Write a public function named Factorial that computes the factorial
of an integer n. The function must only accept non-negative values
of n. The result is always at least 1.
This is the intended comparison. It is not an observed result from the checked-in artifact.
How It Works
- Load 50 programming tasks across four categories.
- Submit the configured prompt for each language.
- Compile generated code and run the task tests.
- Compute scores from compilation and test outcomes.
The current JSON records usesNeutralPrompt: false for its tasks, so it cannot
support the neutral-prompt claim previously made on this page.
Task Categories
The checked-in failed artifact used these categories:
| Category | Count | Examples |
|---|---|---|
| basic-algorithms | 15 | Factorial, Fibonacci, IsPrime, GCD, Power |
| contracts | 10 | SafeDivide, Clamp, SafeModulo, NormalizeScore |
| data-structures | 10 | Sum, Max, Min, Average, Median |
| logic | 15 | BoolToInt, LogicalAnd, IsMultipleOf, SameSign |
The separate task-manifest-neutral.json uses a safety category instead of
contracts. It was not the manifest serialized into the checked-in artifact.
Scoring Formula
Score = compiled × compilation weight
+ test pass rate × test weight
+ contract result × contract weight
The checked-in artifact records either 30% compilation, 50% tests, and 20% contracts, or 30% compilation, 40% tests, and 30% contracts. Contract weight is applied only to Calor results. If Calor compiles without an explicit contract result, the runner awards half of the configured contract weight. The neutral manifest also defaults to a 20% contract weight when it omits the field. A future language-neutral study must register and justify its scoring before execution rather than describe this implementation as a 40/60 comparison.
Checked-in artifact status
| Field | Recorded value |
|---|---|
| Timestamp | 2026-02-21 |
| Source revision | Not recorded |
| Provider / model | claude / unknown |
| Tasks | 50 |
| Valid generations | 0 |
| Compilation and test observations | 0 |
| Exported ratio | 1.00 from equal zero scores; not interpretable as parity |
Running Locally
# Run LLM benchmark with neutral prompts
dotnet run --project tests/Calor.Evaluation -- llm-tasks \
--manifest tests/Calor.Evaluation/Tasks/task-manifest-neutral.json \
--verbose
# Refresh cache (re-run all tasks)
dotnet run --project tests/Calor.Evaluation -- llm-tasks \
--manifest tests/Calor.Evaluation/Tasks/task-manifest-neutral.json \
--refresh-cache
Runner controls
- Caching: Results cached to avoid redundant API calls
- Budget caps: A configured run can enforce a spend limit
- Estimated cost: The runner can display an estimate before API calls
Proposed LLM Evaluation
Multi-model validation, comprehension questions, repeated runs, and cross-model agreement are proposed methodology. No checked-in artifact on this page records an executed Claude 3.5 Sonnet/GPT-4o cross-validation.
Proposed comprehension questions
LLMs answer questions about code understanding:
- What is the main purpose of this code?
- What are the input parameters and their constraints?
- What does this function return?
- What side effects does this code have?
- What would happen if [edge case]?
Proposed scoring
- Correctness of answers (0-1) graded against ground truth
- Tokens used to formulate answer
- Consistency across multiple runs
- Cross-model agreement (higher = more reliable)
Running the Evaluation
Basic Run
# JSON output (default)
dotnet run --project tests/Calor.Evaluation -- run --output report.json
# Markdown output
dotnet run --project tests/Calor.Evaluation -- run --format markdown --output report.md
# Website dashboard format
dotnet run --project tests/Calor.Evaluation -- run --format website --output results.json
# HTML dashboard
dotnet run --project tests/Calor.Evaluation -- run --format html --output dashboard.html
Statistical Analysis
# Run with 30 statistical samples
dotnet run --project tests/Calor.Evaluation -- run --statistical --runs 30
# Custom run count
dotnet run --project tests/Calor.Evaluation -- run --statistical --runs 50
Specific Metrics
# Run only specific categories
dotnet run --project tests/Calor.Evaluation -- run \
--category Comprehension \
--category EditPrecision \
--category RefactoringStability
Interpreting Results
Direction Normalization
For each metric:
- Higher is better (Comprehension, Error Detection, etc.): ratio = Calor/C#.
- Lower is better (Token Economics): ratio = C#/Calor.
After normalization, a ratio above 1 favors Calor and a ratio below 1 favors C#. The ratio does not mean that Calor's raw score is higher.
Statistical Significance
With statistical mode enabled:
- p < 0.05: Result is statistically significant
- Cohen's d: Effect size interpretation
- d < 0.2: Negligible
- d < 0.5: Small
- d < 0.8: Medium
- d ≥ 0.8: Large
The Tradeoff
The legacy static snapshot scores Calor higher on several structural calculators and C# higher on information density. Because the pairs are not all behaviorally equivalent and the metrics are not agent outcomes, this does not establish better agent reasoning, safer refactoring, or a language advantage.
CI/CD Integration
Automated Benchmarks
Repository workflows may run selected benchmark and regression commands. Workflow availability does not establish a completed weekly multi-model LLM study, PP-W pilot, or confirmation. The paused PP-W attempts are recorded separately in the inventory above.
Regression Detection
CI checks implementation and artifact invariants defined by the current workflow. It does not independently adjudicate scientific validity.
Limitations
- Corpus size: 217 paired programs still may not cover all patterns
- C# baseline: Other languages might perform differently
- Static analysis: Most headline metrics do not capture runtime behavior
- Artifact provenance: Several historical artifacts omit an exact model, runtime configuration, measurement binary, or source revision
- LLM variance: Model updates can affect evaluation scores
- No adopter comparison: There is no protected three-arm adopter study, independent methods adjudication, or measured safety/economic advantage
Next
- Results - Detailed results table
- Individual Metrics - Deep dive into each metric