Agent Tasks
Agent Task Benchmark
This page reports one historical prompt-to-code snapshot. On 2026-02-16, a
Claude Code CLI run attempted 89 Calor tasks after receiving a generated
CLAUDE.md syntax reference. The artifact's recorded summary reports 77 passes
and 12 failures across 17 categories, or 86.5%.
The same artifact contains 18 category entries that sum to 78 passes and 11
failures (78/89), or 87.6%. That inconsistency has not been reconciled, so the dashboard
preserves the recorded 77/89 headline and shows the category-entry totals
separately. Its source revision is
107462e, which
declares compiler v0.2.3. The exact Claude model, Claude Code version, executed
compiler binary, working-tree state, and host configuration were not recorded
in the public artifact.
This is a dated, single-run artifact, not a current model-wide success rate, production reliability measurement, or part of the historical eight-metric composite. It does not establish an independent adopter handoff or a comparative C#/protected-C#/Calor result. See the research evidence status for those unfilled evidence requirements.
This snapshot is also not the redesigned effect-rows study. See its registered methodology and paused research status. Two invalid/censored attempts were consumed before the pause; neither the pilot nor confirmation has completed. This historical dashboard is not a substitute for a benefit or null result.
Historical Single-Run Task Snapshot
Dated task-check results, not a current model reliability measurement
107462eHistorical single-run Claude task snapshot. Corpus: tests/E2E/agent-tasks. Recorded source 107462e declares compiler v0.2.3. The exact Claude model, Claude Code version, measurement binary, working-tree state, and host configuration were not independently recorded.
Recorded Summary
Results by Category
Perfect Score Categories (10)
Areas for Improvement
Methodology
Each task provides a natural language prompt describing code to generate. Claude reads CLAUDE.md syntax reference and generates Calor code. Success requires passing custom verification scripts.
This gate is not a calibrated production reliability or adoption threshold. All 89 tasks required Calor-to-C# transpilation, and all 89 used text-pattern scripts. There were 0 enabled contract-verdict checks and 0 behavioral executions. The generated C# was not built or executed. The runner skipped Calor agent lifecycle hooks.
How It Works
For each task, the historical runner:
- Provides a natural language prompt describing what code to generate
- Creates a generated
CLAUDE.mdwith a Calor syntax reference - Invokes the Claude Code CLI once and asks it to write a
.calrfile - Runs the task's configured checks
The runner deliberately skipped calor init, including its agent lifecycle
hooks. The snapshot therefore did not test Calor's full generate-compile-
diagnose-retry workflow.
The historical verifier matrix was:
| Verifier level | Historical coverage | What it checked | What it did not establish |
|---|---|---|---|
| Syntax-pattern script | 89 of 89 tasks | A custom shell script used grep to require text or Calor markers | Semantic correctness or runtime behavior |
| Calor transpilation | 89 of 89 tasks | The Calor compiler accepted the .calr file and produced a .g.cs file | Whether the generated C# built or produced correct output |
| Contract verdict | 0 of 89 tasks | No task enabled the runner's calor verify check | Any proven or disproven contract evidence |
| Behavioral execution | 0 of 89 tasks | No script ran the generated program against expected outputs | Functional correctness for representative inputs |
A task passed when transpilation and its text-pattern script passed. The aggregate must not be read as 77 C# builds, contract proofs, or independently observed behavioral successes. There were zero enabled contract-verdict checks and zero behavioral executions.
Artifact Category Entries
These are all 18 entries in the published JSON, not the unreconciled 17-category summary field.
| Category entry | Recorded result |
|---|---|
| Advanced Contracts | 5/5 |
| Async Functions | 2/4 |
| Basic Syntax | 4/4 |
| Calor Idioms | 2/2 |
| Collections | 5/6 |
| Contract Writing | 5/5 |
| Control Flow | 7/8 |
| Effects System | 5/6 |
| Enums | 3/3 |
| Generics | 3/3 |
| GitHub Projects | 4/4 |
| Lambdas & Delegates | 2/4 |
| Logic Implementation | 4/4 |
| OOP Features | 6/6 |
| Pattern Matching | 5/6 |
| Refactoring | 6/8 |
| String Operations | 4/5 |
| Type System | 6/6 |
| Category-entry total | 78/89 |
Reproduction Status
No documented command can reproduce this historical snapshot because the exact
model and runtime configuration are unknown and the task corpus has changed.
The checked-in generate-benchmark.sh is not a provenance-preserving refresh
path: the unsafe fallback implementation has been replaced with a fail-closed
archival stub that cannot overwrite the dated public artifact.
What the Snapshot Supports
The bounded observation is that one unspecified Claude model, under the historical runner and generated syntax reference, met the configured checks for the artifact's recorded 77/89 tasks. This can identify documentation or fixture areas worth investigating.
A failed task can come from documentation, prompting, model behavior, compiler behavior, the fixture, or the verifier. A passing task shows only that its configured checks accepted that output. Neither outcome isolates documentation quality by itself.
Pass Rate Threshold
The runner labels an aggregate at or above 80% as success. This is a project-defined reporting gate. The repository does not record a calibration study tying 80% to production reliability, adoption readiness, or an acceptable user failure rate.
Adding New Tests
These steps extend the current harness; they do not amend or reproduce the 2026-02-16 artifact:
- Create a directory in
tests/E2E/agent-tasks/tasks/{category}/{number}_{name}/ - Add
task.jsonwith prompt and metadata - Add
verify.shwith verification logic - Run tests to validate
See Adding Benchmarks for full details.