v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Agent Tasks

Agent Task Benchmark

This page reports one historical prompt-to-code snapshot. On 2026-02-16, a Claude Code CLI run attempted 89 Calor tasks after receiving a generated CLAUDE.md syntax reference. The artifact's recorded summary reports 77 passes and 12 failures across 17 categories, or 86.5%.

The same artifact contains 18 category entries that sum to 78 passes and 11 failures (78/89), or 87.6%. That inconsistency has not been reconciled, so the dashboard preserves the recorded 77/89 headline and shows the category-entry totals separately. Its source revision is 107462e, which declares compiler v0.2.3. The exact Claude model, Claude Code version, executed compiler binary, working-tree state, and host configuration were not recorded in the public artifact.

This is a dated, single-run artifact, not a current model-wide success rate, production reliability measurement, or part of the historical eight-metric composite. It does not establish an independent adopter handoff or a comparative C#/protected-C#/Calor result. See the research evidence status for those unfilled evidence requirements.

This snapshot is also not the redesigned effect-rows study. See its registered methodology and paused research status. Two invalid/censored attempts were consumed before the pause; neither the pilot nor confirmation has completed. This historical dashboard is not a substitute for a benefit or null result.

Historical Single-Run Task Snapshot

Dated task-check results, not a current model reliability measurement

Last run:
Commit: 107462e

Historical single-run Claude task snapshot. Corpus: tests/E2E/agent-tasks. Recorded source 107462e declares compiler v0.2.3. The exact Claude model, Claude Code version, measurement binary, working-tree state, and host configuration were not independently recorded.

Unreconciled artifact totals: the recorded summary reports 77/89 passes, 17 categories, and 86.5%. The 18 category entries sum to 78/89, or 87.6%. The cards preserve the recorded summary; the category list renders the entries.
Recorded Summary Rate
86.5%
Above the project-defined 80% gate
Recorded Summary Passes
77
of 89 total
Recorded Summary Failures
12
Recorded Summary Categories
17
18 entries; 10 at 100%

Recorded Summary

Passed
Failed
77 passed12 failed

Results by Category

Advanced Contracts— Quantifiers, implications, custom messages
5/5100%
Basic Syntax— Simple functions and types
4/4100%
Calor Idioms— Common Calor patterns
2/2100%
Contract Writing— Preconditions and postconditions
5/5100%
Enums— Enum definitions and extensions
3/3100%
Generics— Generic functions and classes
3/3100%
GitHub Projects— Real-world project tasks
4/4100%
Logic Implementation— Algorithm implementation with contracts
4/4100%
OOP Features— Classes, interfaces, properties
6/6100%
Type System— Option, Result, type inference
6/6100%
Control Flow— Loops, conditionals, iteration
7/888%
Collections— List, Dictionary, HashSet operations
5/683%
Effects System— Effect declarations and tracking
5/683%
Pattern Matching— Switch expressions and patterns
5/683%
String Operations— String manipulation functions
4/580%
Refactoring— Code transformation tasks
6/875%
Async Functions— Async/await patterns
2/450%
Lambdas & Delegates— Lambda expressions and delegates
2/450%

Perfect Score Categories (10)

advanced contractsbasic syntaxcalor idiomscontract writingenumsgenericsgithub projectslogic implementationoop featurestype system

Areas for Improvement

async functions — Complex await/ConfigureAwait syntax
lambdas delegates — Block lambda syntax specifics

Methodology

Each task provides a natural language prompt describing code to generate. Claude reads CLAUDE.md syntax reference and generates Calor code. Success requires passing custom verification scripts.

Mode: Single run (no majority voting)
Project-defined gate: 80% pass rate required for success

This gate is not a calibrated production reliability or adoption threshold. All 89 tasks required Calor-to-C# transpilation, and all 89 used text-pattern scripts. There were 0 enabled contract-verdict checks and 0 behavioral executions. The generated C# was not built or executed. The runner skipped Calor agent lifecycle hooks.


How It Works

For each task, the historical runner:

  1. Provides a natural language prompt describing what code to generate
  2. Creates a generated CLAUDE.md with a Calor syntax reference
  3. Invokes the Claude Code CLI once and asks it to write a .calr file
  4. Runs the task's configured checks

The runner deliberately skipped calor init, including its agent lifecycle hooks. The snapshot therefore did not test Calor's full generate-compile- diagnose-retry workflow.

The historical verifier matrix was:

Verifier levelHistorical coverageWhat it checkedWhat it did not establish
Syntax-pattern script89 of 89 tasksA custom shell script used grep to require text or Calor markersSemantic correctness or runtime behavior
Calor transpilation89 of 89 tasksThe Calor compiler accepted the .calr file and produced a .g.cs fileWhether the generated C# built or produced correct output
Contract verdict0 of 89 tasksNo task enabled the runner's calor verify checkAny proven or disproven contract evidence
Behavioral execution0 of 89 tasksNo script ran the generated program against expected outputsFunctional correctness for representative inputs

A task passed when transpilation and its text-pattern script passed. The aggregate must not be read as 77 C# builds, contract proofs, or independently observed behavioral successes. There were zero enabled contract-verdict checks and zero behavioral executions.


Artifact Category Entries

These are all 18 entries in the published JSON, not the unreconciled 17-category summary field.

Category entryRecorded result
Advanced Contracts5/5
Async Functions2/4
Basic Syntax4/4
Calor Idioms2/2
Collections5/6
Contract Writing5/5
Control Flow7/8
Effects System5/6
Enums3/3
Generics3/3
GitHub Projects4/4
Lambdas & Delegates2/4
Logic Implementation4/4
OOP Features6/6
Pattern Matching5/6
Refactoring6/8
String Operations4/5
Type System6/6
Category-entry total78/89

Reproduction Status

No documented command can reproduce this historical snapshot because the exact model and runtime configuration are unknown and the task corpus has changed. The checked-in generate-benchmark.sh is not a provenance-preserving refresh path: the unsafe fallback implementation has been replaced with a fail-closed archival stub that cannot overwrite the dated public artifact.


What the Snapshot Supports

The bounded observation is that one unspecified Claude model, under the historical runner and generated syntax reference, met the configured checks for the artifact's recorded 77/89 tasks. This can identify documentation or fixture areas worth investigating.

A failed task can come from documentation, prompting, model behavior, compiler behavior, the fixture, or the verifier. A passing task shows only that its configured checks accepted that output. Neither outcome isolates documentation quality by itself.


Pass Rate Threshold

The runner labels an aggregate at or above 80% as success. This is a project-defined reporting gate. The repository does not record a calibration study tying 80% to production reliability, adoption readiness, or an acceptable user failure rate.


Adding New Tests

These steps extend the current harness; they do not amend or reproduce the 2026-02-16 artifact:

  1. Create a directory in tests/E2E/agent-tasks/tasks/{category}/{number}_{name}/
  2. Add task.json with prompt and metadata
  3. Add verify.sh with verification logic
  4. Run tests to validate

See Adding Benchmarks for full details.