v0.22.0—Bounded nullability checks and practical .NET migration guidance.See what's new

Documentation

Changelog

All notable changes to Calor, organized by release.


[Unreleased]

[0.22.0] - 2026-09-15

Added

  • Nullability checks at three receiving boundaries. Supported possibly-null values now stop compilation at initialization (Calor0272), native return (Calor0273), and resolved method input (Calor0274). Checks cover scalar strings, supported arrays, whitelisted immediate string generic payloads, and identity-proven reference types. Type/effect opt-outs and transpile-only do not disable these active binding checks.
  • A practical nullability migration guide. Complete examples explain nullable references versus runtime Option<T>, explicit defaults, coalescing throws, typed patterns, .NET annotations, and the boundaries that remain unchecked. See Nullability and .NET Interop.
  • Reviewed conversion evidence. The Stage B comparison retains all 289 configured input identities, diagnostic details, fallback outcomes, and attempt histories. Successful fallback C# is not counted as native Calor execution or proof of null safety.

Fixed

  • Nullable-reference metadata survives supported consumers. BCL/native and delegate returns, locals, members, casts, conditional results, and coalescing preserve the supported annotations. Nullable references do not implicitly convert to or unwrap runtime Option<T>.
  • Method arguments use the selected parameter mapping. Named arguments and supported call forms report the relevant nullability diagnostic instead of losing annotation differences during overload selection.
  • Guard facts account for captured mutations. Callable aliases, loop paths, branches, exception filters, and shadowed captured fields no longer restore the stale non-null facts covered by these repairs. Unused lambdas do not execute their writes. Loop analysis also preserves hard diagnostics.
  • Array allocation preserves declared annotations. Sized and rectangular allocations retain a non-null container and their declared element types. This restores the converted FizzBuzz Range example without accepting explicitly nullable elements at non-null receiving boundaries.
  • C# conversion preserves explicit behavior choices. Nullable context and null-forgiving expressions are retained. Calls with an unsupported same-name current-class overload stay opaque. Conversion does not invent defaults, throws, or implicit unwraps.

Known limits

This is bounded receiving-boundary enforcement, not whole-program null safety. General assignment/rebinding, member writes, and constructor inputs remain outside its scope. Generic outer-container nullability and arbitrary generic payloads are not actively checked. Array annotations do not prove every slot initialized; newly allocated reference slots can still contain null.

Some valid C# early-return guard patterns need explicit manual consumption after conversion. The measured MediatR ObjectDetails case is documented with an executable migration; its rejection is not a claim of a null bug in the original C#. The converter does not choose that migration automatically.

D3/D12/D14 safeguards, which conservatively demote proofs and retain runtime guards, remain unchanged. #875 remains open. The separate PP-W-rows experiment remains paused; this release neither resumes it nor establishes a benefit result.

[0.21.0] - 2026-09-11

This product-only release fixes cache correctness and makes diagnostic and conversion evidence easier to interpret. It does not activate the planned nullability checks or require completion of the paused PP-W-rows experiment.

Fixed

  • Cached compilation respects effective type-checking settings. A result produced with type checking disabled cannot bypass a later default invocation. Root and watch cache identities include the effective settings; old policy-blind cache entries are invalidated.
  • Editor diagnostics distinguish analysis-only findings from build blockers. Binder provenance survives ordinary diagnostics and code actions. Existing compilation errors still reject programs under supported type/effect opt-outs.
  • Legacy-semantics guidance no longer promises an unavailable mode. Unsupported major versions still fail with Calor0701. The hint calls for manual migration and distinguishes binder diagnostics from CLI rejection. Corrected historical notes make clear that legacy Info mode is unavailable.
  • Releases no longer start paid model benchmarks. Paid evaluation and agent refactoring require explicit manual opt-in. Automatic benchmark runs use 30 deterministic repetitions; they do not resume the paused experiment.

Added

  • An executable binder diagnostic catalog. Error-capable reporting sites have explicit routing and ownership. Structured receiving shapes prepare for later activation without enabling all uses of a diagnostic code.
  • Structured round-trip failure evidence. Reports retain diagnostic codes, source spans, original inputs, candidate identity, and later recovery outcomes. Fixed paired test attempts remain visible, and a later pass cannot erase an earlier failure. Comparisons and classification records remain unadjudicated until reviewed; successful fallback C# is not nullability acceptance.

Known limits

Calor0272, Calor0273, and Calor0274 remain analysis-only binder findings. Nullable-reference typing, safe-consumption repairs, call mapping, and staged enforcement are unfinished 0.22 work, not features claimed by this release. Existing D3/D12/D14 safeguards and runtime guards remain unchanged.

The PP-W-rows experiment is paused, not complete. Two invalid/censored attempts are preserved; neither the pilot nor confirmation has completed. This release supplies no effect-row benefit or null result and does not authorize collection or apply the pending accounting amendment. The separate research handoff records the remaining work.

Correction - 2026-09-11: nullability diagnostics are not CLI enforcement

The v0.14.0-v0.14.2 entries below overstated nullability enforcement. Calor0272 (binding initialization), Calor0273 (return), and Calor0274 (argument) can be emitted as binder/editor errors, but those diagnostics do not establish production CLI rejection. Program.Compile filters them out through BindingDiagnosticPolicy; the language server binds directly. Other compiler passes may independently reject the same program.

This exclusion appears in the tagged v0.14.0-v0.14.2 sources and remains at the checked main revision 080ed5a7. It is a longstanding routing gap, not a regression introduced by this correction. See the source and release ancestry. The original release text is retained below with dated annotations.

The promised legacy Info mode is not available. The current compiler refuses older semantics majors with Calor0701; changing the declaration alone is not a migration. Its hint now calls for manual migration and distinguishes binder diagnostics from CLI rejection. The hint correction changes neither version compatibility nor which binder diagnostics block compilation, and adds no migration tool.

Full nullability enforcement in the bounded 0.22 plan (#1082) is planned, not shipped; this release contains routing groundwork only. The plan covers specified initialization, return, and resolved method-input boundaries, not every assignment or possible null value. The D3/D12/D14 safeguards, which demote conditional proofs and retain their runtime guards, remain in place.

[0.20.0] - 2026-09-10

This release improves compiler fidelity for character literals and contextual returns, then makes the public evidence boundary explicit. The v0.20 planning tranche ended UNADJUDICATED with an administrative stop: it did not produce an adopter handoff or a comparative agent study.

Fixed

  • Single-quoted character literals now preserve their meaning. Lexing, parsing, migration, emission, attributes, enum constants, numeric conversions, contracts, and round trips now agree on ordinary and escaped character values. Escaped lone surrogate code units round-trip unchanged. Malformed escapes and supplementary Unicode scalars that require two UTF-16 code units are rejected.
  • Result and Option returns keep their expected contextual types. Nested returns no longer lose payloads or choose the wrong wrapper while moving through C# to Calor and back.
  • Arrow conditionals preserve the enclosing indentation boundary. Multiline conditions and indented arrow bodies no longer consume statements belonging to the surrounding block.
  • Character predicates participate in contract verification. Native character types are recognized by the contract pipeline instead of being treated as unsupported numeric expressions.
  • The website's first path is executable and accessible. The homepage, workflow guide, conceptual references, examples, comparison labels, contrast, motion, and large-media behavior were repaired.

Added

  • A reviewed PP-W-rows buildability instrument. The local buildability gate now has held-out runtime checks and tamper-resistant fixtures. Passing that gate establishes instrument buildability only; no agent collection or effect-row benefit result followed.
  • Public M0 evidence and benchmark inventories. The website now separates static calculators, failed provider artifacts, historical snapshots, registered protocols, practice data, and unrun studies.
  • Artifact-derived claim regressions. Tests pin release surfaces, benchmark provenance, verifier levels, denominators, configurations, and evidence links so public claims cannot silently drift from checked-in records.

Changed

  • The historical agent-task result is bounded to what it checked. The site preserves the artifact's 77/89 summary across 17 categories while exposing that its 18 category entries total 78/89. All 89 tasks used Calor-to-C# transpilation and text-pattern scripts; zero enabled contract-verdict checks or behavioral executions. Its 80% threshold is a project gate, not calibrated production reliability.
  • The unsafe agent benchmark generator now fails closed. It can no longer overwrite the dated public artifact with fallback counts and a fresh timestamp.
  • Adoption and benchmark guidance no longer turns missing evidence into a positive claim. The docs distinguish compiler behavior from workflow evidence and label unknown model, runtime, and source provenance directly.

Known limits

The M0 result remains UNADJUDICATED. No qualifying adopter, independent handoff, or protected C#/ordinary C#/Calor comparison was completed. The downstream study remained unauthorized, and the redesigned PP-W-rows protocol has no pilot or confirmatory runs. The static benchmark pairs are not all behaviorally equivalent. See the research evidence status for the formal record.

[0.19.0] - 2026-09-09

This release closes 16 language-audit findings and 15 website-audit findings. The focus is correctness: a successful compilation or proof must not hide a changed result, a missing runtime check, or a reordered side effect.

Fixed

  • The C# CSV benchmark compiles again. A malformed newline literal was incorrectly penalizing the C# input. Both report formats were regenerated; the broader pair-equivalence and interval limitations are tracked in issue #1276.
  • Contract checks survive parameter mutation. Postcondition proofs no longer assume that a parameter still has its entry value after the body changes it. Contract simplification also respects NaN, numeric types, and evaluation order.
  • Effect checking includes assignment targets. Indexed writes and executable receivers, getters, and setters contribute their effects. Nested implementations retain inherited interface contracts in the covered inheritance cases.
  • Native code preserves expression meaning and scope. C# emission retains expression and pattern grouping and gives sibling match arms separate scopes. Unsupported expression-match block arms are rejected instead of losing statements.
  • Migration preserves observable execution in the repaired cases. Short-circuit operands stay conditional; eager operands keep their evaluation order; tuple right-hand sides are captured before destination writes. Dictionary initializers retain their operations and types, loop conditions are reevaluated, and LINQ grouping retains element selectors and deferred execution.
  • Website navigation and accessibility work consistently. Fixes cover benchmark sorting and row expansion, heading anchors and browser history, keyboard-accessible drawers, active-page navigation, theme persistence, contrast, and copy controls. Benchmark timestamps now identify their timezone.

Added

  • Behavioral regression coverage through the production compiler. Generated cases compare results, exceptions, output, state, and evaluation traces. Negative controls confirm that the oracles detect deliberately changed behavior.
  • Frontend coverage floors, native code-generation mutation targets, and specification drift checks. These complement runtime regressions rather than treating emitted text or AST counts as proof of correctness.
  • A complete first-run guide and local documentation search. The website now includes clearer setup instructions, cross-document search, canonical routes, sitemap and robots metadata, and clearer code labels.

Changed

  • Integer overflow follows the documented module policy in production. Under the default checked policy, code that previously wrapped accidentally can now throw OverflowException. Use the documented unchecked policy when wrapping is intended; disabling contract checks does not disable overflow checks.
  • Published claims distinguish guarantees from evidence. Documentation separates runtime contract modes, optional static proofs, effect-checking limits, and static benchmark scores. Decorative website media is optional.

Known limits

Some migration compositions are preserved as explicitly reported C# interop or rejected, rather than translated into native Calor. These fixes do not establish whole-language soundness, complete cross-file effect checking, or an advantage for coding agents over C#. Earlier roadmap proposals outside these two audits are not claimed complete by this release.

The pre-existing optional Tier 2 corpus-workflow defects remain tracked in issue #1241. Historical measurements and their recorded misses remain unchanged.

[0.18.0] - 2026-09-08

0.17 was about reach — how many of the 364 converted modules the effect checker manages to get to. It got to 324. 0.18 asks the next question: once it gets there, is what it says true? Two answers came back, and both were uncomfortable.

The compiler was reporting errors against code it had written itself. And a claim we published in 0.17 — that six ways of hiding an effect had been closed — had never actually been measured. It has been now.

Benchmark Results (Statistical: 30 runs)

  • Overall Advantage: 1.32 (Calor leads)
  • Metrics: Calor wins 7 categories, C# wins 1
  • Highlights:
    • Comprehension: 1.84x (Calor)
    • ErrorDetection: 1.49x (Calor)
    • TokenEconomics: 1.42x (Calor)
    • RefactoringStability: 1.38x (Calor)
    • InformationDensity: 0.98x (C#)
  • Programs Tested: 217
  • Note: re-run for this release, unlike 0.17's, which were carried forward. All eight ratios reproduce the previous figures exactly. That means they are reproducible, not that they are stable: these metrics are computed over a fixed set of programs, so a deterministic calculation agreeing thirty times is not evidence of anything. Worth settling before a future release treats exact reproduction as a result.

Fixed

  • The C# → Calor converter now declares the effects the code it writes actually performs. Every Calor declaration carries an effect row — the §E{...} line saying what the body may do: alloc for allocating memory, mut for mutating something, cw for writing to the console. The converter wrote that row using its own copy of the analysis, kept beside the code that emits the Calor. The two copies drifted. The converter's copy never looked inside a foreach loop's collection, a using statement's resource, or a lambda body, so it wrote "this function is pure" over functions that allocate. Compiling the result then reported an error — "uses effect 'alloc' but does not declare it" — against code the converter had just produced.

    The row is now computed by the compiler's own effect analysis, the same pass that later checks it, run over the Calor text the converter is about to write. One analysis, one answer, so the row and the check cannot disagree. Across the three real projects we convert as a test — MediatR, Serilog and FluentValidation, 364 files — this takes those errors from 219 in 53 files to zero.

  • A property's get can now do the ordinary things getters do. A property accessor has nowhere to write down its effects, so the compiler gives it a fixed allowance instead. The get had no allowance at all — the compiler hands one to set and to init and simply missed get — so a getter was checked as though it had promised to do nothing, and any effect in one was an error you could not fix. A getter that returns a fresh list allocates. A getter that computes its value once and remembers it mutates. Neither could be written.

    Why widening the allowance is safe is the part worth stating: the allowance is not what keeps effects honest. Reading a property charges the getter's effects to the code doing the reading, which still has to declare them. Fixing this exposed a second hole and closed it — reading your own property through this. charged nothing, while the identical read through another variable charged in full.

  • --permissive-effects now waives only what the compiler cannot determine. The flag exists for code the compiler can't fully analyse, and it used to downgrade every "uses an effect it does not declare" error to a warning. That covered two different situations: we cannot tell and we know this is wrong. It now waives only the first. Cost, measured before the change: about 21 diagnostics across 14 modules move from warning to error, against 328 the flag still suppresses.

    This is also what uncovered the converter defect above. The flag had been hiding it, on exactly the code it was built for.

Verified

  • The six ways of hiding an effect: twelve shapes checked, twelve closed. 0.17 said it had closed a set of holes where an effectful function could be smuggled into a context declared not to have effects. The instrument that was supposed to check that claim was registered and then never built, so the claim shipped unevidenced. This release built it: twelve specific shapes, five that used to slip through and seven that were already handled and must stay that way. All twelve behave as the table says. The seven controls matter as much as the five fixes — a result of "five out of five" with a regressed control would read like success and be one.

Changed

  • The test suite no longer kills the machine that runs it. Our continuous integration was being terminated part-way through at a measured 14.7 % rate, which meant roughly one in seven attempts to verify a change produced no answer. The compiler's test suite now runs as two separate processes rather than one, which halves peak memory: 9.3 GB → 5.9 GB on a 16 GB machine.

    This caps the symptom and does not fix the cause. Eight explanations were tested and refuted, and what allocates the memory is still unidentified. A single process still exhausts a 16 GB Linux machine. Said plainly here because a green pipeline is easy to read as "solved".

Measurement integrity

Three fixes that change no compiler behaviour and matter anyway, because a number nobody can check is not evidence.

  • Every measurement we publish now names a commit that exists. Each of our measurement records stamped the commit it was taken at, and every roadmap since 0.15 cited those stamps as provenance. None of them resolved. The repository squashes branches when merging, so a stamp written on a branch names a commit the merge throws away; three of five named commits existed on no remote branch at all and would not have survived routine cleanup. The numbers were never in doubt — what was gone was anyone else's ability to check them. The stamps are preserved (overwriting one would falsify the record) and a companion index now carries a commit that does resolve, stating for each whether it is the same compiler, verified by comparing source trees, or merely where the numbers landed. A test fails if any stamp stops resolving.

  • The benchmark bot can no longer swap one kind of measurement for another. The published headline figure is computed from 30 runs. An ordinary automated benchmark run produces a single-run figure, and it was overwriting the 30-run one under the same name, in a pull request that said "updated with latest metrics". A reviewer would read 1.32 → 1.28 as Calor losing ground. It is not — the two numbers are not measuring the same thing. The generator now refuses, and a test compares what the website publishes against what the changelog says.

  • A second defect in the same report did not reproduce, and we say so rather than shipping a fix for a problem that is not there.

Not in this release

Named rather than omitted, because a thing that is written down gets picked up.

  • Two effect-row features slip for the second time: function-typed values carrying their row end-to-end, and index parity with calor build. Both are real and neither is urgent. Venue: 0.19. A third slip should mean scheduling them or retiring them, not re-tiering them again.
  • Effect promises are not checked across file boundaries. A class that overrides a method from a base class in another file is not checked against that base's effect row — the compiler cannot see across the boundary at exactly the two places whose purpose is stopping effects from leaking through inheritance. This is 54 % of one diagnostic category, and it was carried for three releases as though those base classes were external library types. They are not: 60 of 61 are declared in the same project, one file over. Diagnosed here, fixed in 0.19.
  • The agent experiment did not run. Its fixtures were redesigned first, because the previous set could not measure what it claimed to: the tasks were built so that hiding an effect also broke a stated requirement, which the ordinary tests already catch. Where those coincide, effect checking is redundant by construction and no number of runs produces a result. The redesign is registered; the run is not funded yet.

[0.17.0] - 2026-09-02

Calor compiles to C#, and the way we check that claim is to convert three real open-source C# projects — MediatR, Serilog and FluentValidation, 364 files, one Calor module each — and see how much of the result the compiler can still make sense of. 0.16 got all 364 modules to parse. 0.17 is about what happens after that: a module that parses can still stop at name binding, before any effect checking runs. Sixty stopped there when this release opened. Forty do now, and the number of modules the effect checker reaches went from 304 to 324.

Added

  • The compiler now records why it stops on the code it cannot process. When we convert real C# projects to Calor, some modules fail during name binding and never reach the effect checks. Until now that was a bare count — sixty of them, cause unknown. We now record which diagnostic stopped each module, and in which file, so the biggest group can be found and fixed rather than guessed at. One diagnostic accounted for two thirds of them, which is what the next two entries went after. We also count the modules that stop for more than one reason, because fixing a single cause cannot rescue those. As shipped, 40 modules stop at binding: 17 on the largest cause, 12 and 11 on the other two, with 6 carrying more than one.
  • A demand count we had been reading from the wrong diagnostic. We track how often the compiler meets a function value it cannot identify. The point of the count is to decide whether it is worth reading effect information out of compiled .NET assemblies, for the delegates the .NET libraries hand back. That count read zero — because the check watched a diagnostic this case never emits. A different diagnostic reports it, and nobody was counting that one. Over exactly the same set of files, the count is 8,231 occurrences across 187 of the 324 modules the effect checker reaches. This does not justify building the feature. That diagnostic covers every unresolved external call, not just the narrow case we care about, so it is a ceiling on the demand rather than a measurement of it. Both the test and the saved record say so, in those words.
  • A property can now carry an effect row. §PROP{p:Handler:Func<i32>:pub} §E{cw} says what the callback stored in that property is allowed to do — the same annotation §FLD has carried since 0.15. In 0.16 writing it was a parse error, so a property was the one place a callback could be stored with nothing said about it, and calling it went unchecked. It is now read, emitted, and enforced.
  • Overload resolution now understands the conversions C# already allows. When you call a function and more than one overload has the right name, the compiler decides which one you meant. It was comparing parameter types too literally. Passing a PropertyEnricher to a parameter typed ILogEventEnricher — an interface the class implements — was reported as "no overload matches", and so was passing a generic parameter T to a parameter typed as one of T's own constraints. Both are ordinary implicit conversions, and both now resolve, for methods and for module-level functions alike. Over the conversion corpus this took the modules the effect checker reaches from 304 to 319.
  • The compiler can now find members a type inherits. Looking up a method asked the receiver type for its own declarations only — nothing about base classes, base interfaces, generic constraints, or System.Object. So IValidator.GetType() was "no such member", because an interface has no base type to walk, and so were MethodInfo.GetParameters() (declared on MethodBase), IList<T>.Add(item) (declared on ICollection<T>) and Severity.ToString() (an enum reaches ToString through System.Enum). Of the 1,248 calls into .NET library types we measure across the three projects, the compiler resolves 1,197 — up from 1,158, and all three projects improved. That is the measured denominator, not every call in the corpus: calls into the projects' own types are not in it.
  • A related fix, for members of types defined in the code being compiled. Resolution works by asking Roslyn a question in a small synthetic C# file. That file can only name types it has a reference to, so a type from the code under analysis could not be named at all, and the failure was indistinguishable from "no such member". For an inherited member the distinction is empty — IValidator.GetType is object.GetType — so the question is now re-asked of whichever type actually declares the member.

Changed

  • An effect variable the compiler cannot pin down is now charged, and that can turn a clean build into an error. When a function is written to work with any callback — §F{f:Run:pub}<eff e> (Func<i32>:f §E{e}) — the compiler works out what e stands for at each call site, from the callback you passed. If it could not work that out, it said so and then charged the caller nothing. Unknown is not "no effects", it is "any effects", so a §E{} function could call a console-writing method through such a callback and the build still passed. It is now charged as unknown, which means the caller reports Calor0410 unless it declares the effect. Hand-written Calor that compiled under 0.16 can fail under 0.17 for this reason. --permissive-effects still relaxes it, and converted code is unaffected — the converter emits no rows.

Fixed

  • Converted C# using an extension class emitted calls the C# compiler rejects. A class carrying §EXT emitted calls to module-level functions unqualified, so the generated C# failed with CS0103 — the name does not exist in this context. The same function referenced as a method group rather than called had the same problem. (#1137, #1118)
  • A var holding the result of a call lost its type. var evens = numbers.Where(...) came back untyped, so the next call on it — evens.OrderBy(...) — had no receiver type to resolve against, and the effect checker reported an unknown call it then refused. The type is now written down. (#1128)
  • A named argument matching no parameter crashed the binder. Calling §C{Only} §A[noSuchParameter] INT:1 §/C against a function with no parameter of that name threw IndexOutOfRangeException instead of reporting that no overload matches. (#1127)

Fixed before release

Defects this release's own work introduced, found by the adversarial review of it and fixed before anything shipped. No released version has them; they are listed because the review that caught them is part of how this project works.

  • A cyclic generic constraint — §WHERE T : U together with §WHERE U : T — sent the new conversion-cost search chasing the two constraints in a loop until the stack ran out. A stack overflow cannot be caught, so the process died rather than reporting anything.
  • The new overload work decided whether it could see a type well enough to judge a call by checking whether the type's name was upper-case. The canonical form of the smaller numeric types is not — FLOAT[bits=32], INT[bits=8][signed=true] — so every call carrying an i8, u8, i16, u16 or f32 argument was treated as unjudgeable and mismatches went unreported.
  • A pattern variable was not registered as a declaration, so after §IF{i1} (is o i32 value) a reference to value could resolve to a module-level function of the same name instead of the variable just introduced.
  • The converter's new type inference wrote anonymous types and tuples into a binding header. Their names contain colons, commas and spaces — the characters that separate the fields of that header — so the Calor it produced no longer parsed. Both are left untyped again.
  • The parser accepted an effect row on a property; the emitter dropped it. calor fix, the indent migrator and the C# converter all round-trip through that emitter, so each silently deleted a hand-written row.

Corrected

  • A resolution figure published in 0.15 was wrong, and the compiler was never the reason. The 0.15 notes reported gate 6 at 817/1,248 = 65.5%. The harness was asking the compiler to find the reduced form of an extension method — for items.Where(pred) it asked for Enumerable.Where(Func<T,bool>), a one-argument overload that does not exist and never did. 341 of the 431 unresolved cases were that question. Asked correctly, the same 0.15-era compiler resolves 1,158/1,248 = 92.8%. This release's improvement is measured from 92.8%, not from 65.5%, and the old figure is printed here beside it rather than quietly replaced.

Proof points and gates

  • PP-R1 leg 1: MISS. Before measuring anything, we registered the rule that the overload work had to recover at least half the largest group of binding failures, and never fewer than 10 modules. The group turned out to be 40, so the target was 20. That work recovered 15. The rule was frozen before the measurement existed, so this is a miss, published rather than re-scoped. The release total is 20 — but the other 5 came from later, different work, and a target met by something other than the thing it was registered against is not the target being met.
  • Gate 6 (metadata resolution), movement leg: SATISFIED. 1,158 → 1,197 of 1,248, and no project fell.
  • Gate 9 (conversion denominator): held. Parse failures stay at 0; ModulesEnforced is 324 against a floor of 304.

[0.16.0] - 2026-09-01

Calor 0.16 makes the project model useful to agents. The new calor_query MCP tool exposes callers, callees, change impact, and effect rows from the same persistent index used by calor query. Those CLI facets also gain machine-readable JSON and explicit partial-answer residuals.

Migration is more dependable. All 364 modules in the MediatR, Serilog, and FluentValidation conversion corpus now parse. Uncertain lambda parameter types no longer become invalid ? syntax, and one failed file no longer floods its dependents with misleading generated-C# errors.

MSBuild projects can use CalorPermissiveEffects without disabling effect checking entirely, and verification now says plainly when Z3 is unavailable. The effect-row benefit study was prepared and dry-run honestly, but the sample was underpowered. This release therefore makes no safety-benefit claim from it.

Added

  • The scorecard for the "do effect rows actually help?" experiment now exists, and it was written down before the experiment ran. We are about to pay for a set of coding runs that ask one question: given the same task, does an agent using the newer, stricter compiler sneak fewer hidden side effects past it than an agent using the older, more permissive one — and does the stricter compiler cost noticeably more work to reach a green build? A new script, bench/phase0-agent-native/ppw-analyze.py, scores the finished runs; a new ledger, bench/phase0-agent-native/effect-rows-benefit-ledger.json, records the answer. Both were built before any run happened, so the rules cannot be bent once the numbers are in: the script transcribes rules already frozen in the project's registration document, a test recomputes the ledger from scratch and compares it byte for byte, and the script refuses to score a run at all if the recording is incomplete. One rule matters more than the rest and is easy to get backwards: a "sneaked-past side effect" only counts on a workspace that actually built. Runs that never compiled go in their own separate bucket, because counting them as failures would make the stricter compiler look worse precisely for being strict. The script also refuses to declare a winner from the cheap practice round: a practice run's only job is to work out how many runs the real experiment needs, so it reports that number and stays silent about who won. Nothing here changes the compiler or any program you write — it is measurement plumbing. (v0.16 W2, gate 10)

  • We did a cheap practice round of that experiment, and it found three defects in our own measuring equipment. The practice round ran real coding tasks against both compilers. Before it could tell us anything about the compilers, it told us that the scoring script and the harness that runs the experiment disagreed in three ways, each of which would have wasted the whole paid experiment: they spelled the task names differently, so the scorer saw no tasks at all; the safety check that proves each compiler is configured as claimed writes one of two different words depending on which compiler it checked, and the scorer accepted only a third word, so it discarded every run of one arm; and the version of the test equipment was never recorded. All three are fixed, and the scorer now also catches the case where the two compilers were accidentally swapped. The practice round itself is committed to the repository — every run, including its full transcript — so anyone can rerun the scoring and get the same numbers. It is a practice round, so it declares no winner. Full write-up: docs/plans/2026-09-01-ppw-rows-dry-run.md. (v0.16 W2)

  • The benchmark's per-turn recording is now proven on real runs. The recording added earlier this release saves a full transcript of what the agent did on each turn. The practice round is the first set of runs to carry those transcripts, so the per-turn table — files read, searches, builds, edits — has real numbers behind it for the first time. (v0.16 W1/W4)

  • The practice round finished, and it says the real experiment would be too small to settle anything. A second practice round covered the three tasks the usage limit cut short, so all six now have results. The practice round's job is to work out how many runs the real experiment needs, and the answer is nine runs per task per compiler — roughly twice what the budget we fixed in advance can pay for, which affords three. Three runs would give the experiment a 48 % chance of detecting the effect even if the effect is real, so by the rule we wrote down before starting, the experiment is recorded as underpowered rather than being run anyway and reported as if it had settled the question. Two other things worth saying plainly: across all 28 usable runs, not one agent on either compiler hid a side effect where the tests could catch it — the behaviour being measured did not occur at all — and the stricter compiler cost about 12 % more work to reach a green build, below the threshold we had registered as "noticeable". Neither is a result about effect rows yet; both are results about these six tasks. The scorecard now says exactly that, in the place we promised to write the answer: underpowered — the claim is neither supported nor refuted. The next release can revisit it with tasks that actually tempt an agent into cheating. (v0.16 W2)

  • New error Calor0406: the compiler now tells you when effect checking gave up early. Effect inference runs as fixed-point loops with an iteration limit, so mutually recursive functions cannot make the compiler spin forever. Before, hitting that limit in the loop that propagates an effect up a call chain was not reported at all — the compiler stopped and declared the program fine. Now both loops report Calor0406 as an error naming which loop stopped, the limit, and the functions involved, so an unfinished analysis is never passed off as a clean build. The limit for a recursive group now grows with the size of the group, so a big but ordinary group never trips it. The project index (calor query effects) says "did not converge" for such a file instead of recording half-finished rows. (v0.16 W5, gate 11)

  • A new "edit script" test, ES-08, checks effect rows across files. The compiler's build-cache test suite replays small editing sessions and checks that a clean build and an incremental build report exactly the same diagnostics. ES-08 is the first script whose edit changes a shared helper's effect row — the side effects a function passed into it may have. It confirms that callers in untouched files get the right new errors, that the project index's effect-row entries move with the edit, and that unrelated files are left alone. This script was supposed to be registered before the effect-row feature merged in 0.15 and was not; the registration record says so.

  • CI now installs the calor command-line tool from a freshly built package on Windows, Linux and macOS and checks that it can actually prove contracts — and that when the Z3 solver is missing it says so out loud instead of quietly skipping everything. Carried over from PR #982. This check is advisory for now: like the matching check for the SDK package, it is not a required status check, so a failure is visible but blocks nothing.

  • New MSBuild property CalorPermissiveEffects. Setting it to true in your project file does what the command line's --permissive-effects already did: the compiler assumes a call it cannot resolve is harmless, so it stops reporting Calor0411 and Calor0425 (the two "I cannot tell what this does" diagnostics) and downgrades Calor0410 — "this function does something it did not declare" — from error to warning, whether the call stays in one file or crosses files. That helps while migrating old code. It does not relax the checks on effects you declared yourself: a callback whose effects do not fit its destination (Calor0424) and an override or interface implementation that does more than the method it inherits from (Calor0420, Calor0421 — doing less is fine) are still errors. The property is off by default, so nothing changes unless you turn it on — and the first build after you change it rebuilds every file.

  • AI agents can now ask the MCP server about your project's structure. A new calor_query tool answers four questions about any function or method: who calls it (callers), what it calls (callees), what a change to it would affect (impact, and with effects: true, which callers would stop compiling if its effects changed), and what effects it declares versus what it actually does (effects). The answers come from the project index calor query reads, and they match the command line word for word. If the index is stale the tool rebuilds it first (or refuses, with noBuild: true), and every answer says when it might be incomplete and why. Like the file-writing tool, it only works inside the folder the server was started on (calor mcp --root <project>), because rebuilding an index writes to disk.

  • calor query callers, callees and impact gained --json, so every query facet is now machine-readable, not just effects. The text output of every facet is unchanged, character for character. Two additions to the JSON: each answer now names the facet it answers (all of them are labelled query, so this is how a program tells them apart), and an answer that might be incomplete now carries the residual — the list of things the index could not resolve — that the text output has always printed. calor query effects --json is where you will notice both, and both are new fields: nothing was removed or renamed.

  • calor_compile (MCP) can compile a folder as one project with options.crossModule: true, checking effects across files the way calor -i a.calr -i b.calr does, and it now returns every diagnostic (warnings included) for each file. options.enforceEffects and options.requireDocs match the command line's --no-enforce-effects and --require-docs.

  • Benchmarks: a script that explains the "more turns" gap from what was archived. Our 0.15 agent benchmark showed the new compiler's agents taking a few more turns than the old one's, and we could not say why. bench/phase0-agent-native/ppe1-turn-attribution.py now recomputes every number behind that finding from the archived runs — turns, tokens, wall-clock, the diagnostics each build produced (none from the new effect checks), and two significance tests — so anyone can rerun it and get the same file byte for byte. A --transcripts mode counts what an agent did on each turn (reads, searches, builds, edits) once the next benchmark saves those transcripts. A test fails if the committed numbers drift or an archived run goes missing.

  • Benchmark: added six programming tasks that will test whether the effect rows added in 0.15 stop an AI agent from hiding a console message inside a function declared silent. Each task ships the same starting program built for the old compiler (0.14.3) and the new one (0.15.0), a plain-language description, held-out tests that catch the message if it escapes, and two worked solutions — the tempting shortcut and the honest fix. We also recorded what each compiler says about every one of those solutions, so the numbers are fixed before anyone runs the experiment. Nothing is run yet.

  • Found while building the above: on 0.15.0, when you pass a stored callback to another function and the compiler cannot tell which one you meant — you wrote this. in front of it, you kept it in a property instead of a field, or it came from a base class — the compiler stops tracking that callback's effect row. It warns instead of failing the build, and the hidden message gets through. Reported as issue #1136.

  • Also found: a derived class can no longer call the plain functions in its own module by their short name — the generated C# does not compile. The fully qualified name works around it. Reported as issue #1137.

Changed

  • Calor0600 is no longer used for effect checking that did not finish. That code belongs to the API-strictness family; the loop that borrowed it as a warning now reports Calor0406 as an error instead. (v0.16 W5)

Corrected

  • A number we published in 0.15 was wrong, and here is the right one. We keep a measurement ledger in this repository — bench/phase0-agent-native/calor0425-corpus-ledger.json, also quoted in our 0.15 planning document — and during 0.15 it said the compiler's "callback effects are unknown here" warning (Calor0425) appeared 8 times across 99 of the 364 real-world C# files we convert and check. (The figure was never in these release notes or on the website; it is being corrected here because it is a published number either way.) That 99 was not how many files the compiler actually checks — it was how many files our measurement chose to check. The measurement discarded any file that name binding complained about at all, but the real compiler keeps going: most of those complaints are internal notes it never shows you, so it checks the file anyway. Measured the way the compiler really works, it checks 256 of those files, and the warning appears 67 times across 30 of them. Nothing about the compiler changed — only the yardstick did. Both figures are published side by side so the correction is visible rather than quietly swapped in. The measurement now uses the compiler's own rule, and each of our two measurement ledgers states on its face which rule produced it, so this cannot happen silently again.

  • Internal benchmark tooling: every agent-harness run now saves the full turn-by-turn transcript, the agent's own build output, and the compiler's build fingerprint next to its result; a run without a transcript no longer counts. The harness also accepts a registered "pre-rows" control arm for the v0.16 rows experiment.

Fixed

  • calor query effects and the calor_query tool no longer claim a large group of mutually recursive functions "did not converge" when the compiler builds it fine. Effect inference iterates over each group of mutually recursive functions, with a limit on how many rounds it will run. That limit grows with the size of the group, so an ordinary big group never trips it — but the project index was pinning the limit to a flat 100 rounds, which switched the size-based growth off. The result: a big group needing more than 100 rounds to settle — a chain of 100 members or more, each calling its neighbours — compiled cleanly, yet the index reported Calor0406 for that file, discarded every effect row in it, and answered "effect inference did not converge" for any function you asked about. The index now uses exactly the same limit the compiler uses, so the two agree.

  • The compiler no longer crashes on converted code whose local bindings form a cycle (for example, a value whose type depends on itself, which the C# → Calor converter can produce from out var arguments). The effect checker now detects that it is re-entering a name it is already resolving, treats that value's type as unknown, and moves on. This also fixes two calor mcp tools — edit_preview and file_write — which run the same check and could take the whole server down on a file of that shape. There is also a limit on how far the checker will follow one chain of bindings: code that stays under the limit — which is all ordinary code, by a wide margin — is checked exactly as before, and code past it is treated as "type unknown" instead of crashing. (#1104)

  • No more misleading "generated C# failed" errors next to a real error. When one file in a project failed to compile, a clean build used to also blame every file that called into it (Calor1002, "name does not exist"), while an incremental rebuild of the same project did not — and dotnet build and calor build disagreed about it. Now both report only the real error: the compiler skips checking just the generated code that calls into the failed file, and still checks everything else, so a genuine problem in an unrelated file is not hidden. A file whose generated code was skipped this way is not written out or cached, because nothing checked it; the next build checks it. Found while building ES-08.

  • calor verify tells you when nothing was proved. If the Z3 solver cannot be loaded, the text report now says so at the top and again in the summary. Before, that notice appeared only in JSON output, so a broken install looked like "0 proved, everything skipped" with no explanation. Carried over from PR #982.

  • Converting C# to Calor now produces Calor the compiler can parse for 59 more real-world files (#903). Before, across the three sample projects we test against (MediatR, Serilog, FluentValidation), 59 of 364 converted files failed to parse: long lambdas emitted at the wrong indent (Calor0099), an empty marker interface followed by the next type (Calor0100), a method call split across lines (Calor0100), a list of objects with property initializers (Calor0100), and an else if whose condition needed a temporary (Calor0117). All 364 now parse. An empty interface followed by a sibling type is also accepted in hand-written Calor, as long as the next type starts in the same column as the interface; a type indented underneath one is still an error, because interfaces cannot contain types.

  • Converting a lambda whose parameter type could not be inferred no longer writes ? as the type (#1097). The compiler could not parse ? and abandoned the rest of that method with an internal error (Calor0932); it now treats the parameter like any other untyped lambda parameter, and when the lambda is passed to a declared Func<...> or Action<...> it reads the parameter type from there. Across the same three projects this removes 72 of 73 internal errors.

[0.15.0] - 2026-08-27

Calor 0.15 is the "Composable Effects" release. You can now write, on a callback's type, what that callback is allowed to do — the same §E{…} tag functions have always used, placed on the parameter, field, return type or lambda it describes. One function can serve both pure and impure callbacks: give it a placeholder effect (<eff e>), and the compiler resolves the placeholder at each call site from the callback you actually passed. And the compiler checks all of it: a callback that does more than its destination allows is an error (Calor0424), one the compiler cannot verify is a warning (Calor0425), and calling a callback charges its effects to the caller. So a callback cannot smuggle an effect past the six places the compiler now closes: binding a callback to a name, passing one as an argument, returning one, overriding a method, implementing an interface, and instantiating a generic placeholder. Not closed yet: callbacks reached by reflection or dynamic, event-handler +=, delegates returned by .NET libraries (warned, not verified), and async rows. The project index knows about effects too (calor query effects). The proof point we registered in advance came out HIT, and here is what it cost: on ordinary tasks with no callbacks in them, an AI agent used roughly 18 % more output tokens to get the same programs green with the stricter compiler — a real cost, and inside the limit we set beforehand. See Effects and the calor effects command.

Deferred to 0.15.x under roadmap §4.2's own rule for its SHOULD tier ("ship if they fit; defer to 0.15.x without renegotiation") — nothing overran; the release simply ships without them: review-packet reading the index (E6), the MCP query surface over the index (E7), contract outcomes in the index (E8), an affected-tests facet (E9), and a real Calor arm in the agent benchmark (M2).

Added

  • We built the scorecard that will judge the new effect-rows feature. Before any of the effect-rows work was written, we registered a test for it: ten small, deliberate mistakes (a callback allowed to print when its interface forbids it, an effect smuggled through a generic helper, a callback with its effect annotation deleted) planted in five frozen example programs, plus a "nothing else may change" check on the unmodified programs. The compiler must catch all ten with the exact diagnostic we predicted, at the exact place, and stay quiet on the originals. That scorecard now exists as a file the test suite re-checks on every build, and the deterministic half scores 10/10 with the originals clean. The other half — does the stricter compiler make an AI agent spend noticeably more tokens to get a program green? — was a paid experiment whose runner and arithmetic are checked in (and the runner refuses to spend money without an explicit confirmation).
  • We ran the cost test for the new effect-rows feature, and the scorecard reads HIT. An AI agent solved the same four small programming tasks five times each against the previous compiler (v0.14.3) and five times each against the new one — 40 runs, every one finishing green. Agents typically needed about 18 % more output tokens to finish the same small tasks with the new compiler. That is a real cost, but it is within the limit we set beforehand (the test fails only when the typical (median) increase is more than 35 % and the statistics say the increase is clearly above zero; here the lower bound is below zero). One task (the inventory one) was 51 % more expensive on its own; another was cheaper. "HIT" means we did not detect a large slowdown — with 40 runs the test cannot prove there is none, and it says so. The runs, the arithmetic and the verdict are all checked in, and the release re-checks them. The 40 finished programs the agents wrote are archived with the run, so the set of checked-in Calor files grew from 886 to 926 and the counters that watch that set were updated to match (the counts that measure the compiler itself did not move).
  • You can now ask the project index about effects. calor query effects Leaky tells you three things about a function: what it declares (§E{…}), what the compiler inferred its body actually does, and whether the two agree — with the exact diagnostic that fires when they don't. The verdict is the compiler's own, recorded when the index was built and read back rather than recomputed for the query. (It is not a prediction of what your next build will print: the index is always built with strict effect settings, so a build you run with --permissive-effects can stay silent about something the query still reports.) If the compiler had to assume something (say, because the body contains C# it can't see into), the answer says what it assumed and why. Add --json to get the answer as data. There is a "what would break" question too: calor query impact Log --effects --row fs:w walks every function that calls Log, directly or transitively, and tells you which of them would stop compiling if Log were allowed to write files. Along the way, two things got fixed: calor query callers and impact could not see a call written as its own statement (§C{Log} §A x §/C) — only calls used inside an expression — so they now see both; and the build cache's effect summary used to key functions by name, which collapsed two same-named methods in one class into one. It keys them by identity now. Both the index and the build cache rebuild once from scratch the first time you run this version; nothing you wrote changes.
  • Those callback annotations now mean something to the compiler. You could already write §I{Func<i32,i32>:transform} §E{cw} — "this callback may print" — and the compiler would faithfully remember the words. Now the annotation is part of the type. Two callbacks with the same shape but different permissions are two different types, the way i32 and str are two different types, and the compiler can tell them apart. With that comes a subtyping check with three answers rather than two: yes, no, and cannot tell. When a callback comes from outside Calor, or from a library the compiler has no information about, the honest answer is "I don't know", and saying "yes" there is how effects quietly escape. So "I don't know" is now a thing the compiler can say, and it never counts as a yes.
  • Now the compiler tells you when a callback doesn't fit. If you pass a callback that prints to something that promised not to print, that is an error: Calor0424. It names both sides — what the callback may do, what the destination allows, and exactly which extra permissions are the problem — and it suggests the two fixes: widen the destination's §E{…}, or pass a different callback. This is checked in five places: binding a callback to a name, passing one as an argument, returning one, overriding a method, and implementing an interface. When the compiler can't tell — because one side carries no annotation at all — you get Calor0425 instead, as a warning. Same five places. (The sixth place from the introduction — a generic effect placeholder — is reported as Calor0410 on the calling function instead; see the <eff e> entry below.) "I don't know" and "that's wrong" are different answers and now have different codes.
  • --permissive-effects does less than it used to, on purpose. It used to turn effect mismatches on overrides and interface implementations into warnings. It no longer does, and it will never soften a Calor0424 either. What it still does is waive the "I can't tell" cases — it silences both Calor0411 and Calor0425, and demotes Calor0410 to a warning. Waiving we don't know is honest; waiving we know it's wrong isn't. See Effects for the current list of what it does and does not waive.
  • Write one function that works for pure and impure callbacks. Say you write a Map that takes a list and a callback. If the callback is pure, Map should be pure. If the callback prints, Map prints. Before, you had to pick one and be wrong half the time. Now you write <eff e> on the function and use e as the callback's permission, and the compiler resolves e at each call site, from the callback you actually passed. Pass a pure callback and Map is pure; pass a printing one and Map prints — and if the calling function didn't declare that it prints, you get told, with a line explaining exactly where the permission came from. The name you give the placeholder doesn't matter: an interface can call it e and the implementing class f; the compiler matches them by position, the way it already matches ordinary type parameters. If the compiler can't resolve the placeholder — because the callback you passed carries no annotation — it says so (Calor0425) instead of guessing.
  • The §E{…} on a lambda is now a promise the compiler checks. You could always write one; the compiler read it and threw it away. Now it compares the annotation against what the lambda body actually does, and reports a body that does more than the lambda promised. If you don't write one, the compiler infers the permission from the body — which means a lambda you hand to something annotated is now properly checked instead of coming back "I can't tell".
  • Two more "I can't tell" cases now say so out loud. When a method overrides something from a C# base class the compiler can't see, or a class satisfies an interface through a member inherited from one, the permissions can't be verified. That used to be filed as an assumption (Calor0419); it is now a Calor0425 warning that names the row it is assuming and says plainly that it is assumed, not verified.
  • Callbacks whose type came from a §CSHARP block are checked now too. If you declared a delegate inside an interop block and used it as a parameter type, the compiler previously didn't recognise it as a callback at all and skipped the check. It asks the type checker now instead of guessing from the name.
  • Calling a callback now just works — and is charged to whoever calls it. Until now, calling a function value (a callback parameter, a lambda you bound to a name, a callback field) was refused outright under effect enforcement (Calor0418), whatever you had annotated. Now the compiler reads the callback's annotation and charges it to the calling function, the same way calling a named function charges that function's §E{…}. If transform may print and Apply calls it, Apply prints — and if Apply didn't declare that, the error tells you exactly why: Effect row: charged by invoking 'transform' (row: cw). A callback with no annotation is a callback the compiler knows nothing about, so calling it is "I can't tell" (Calor0425) and the caller is charged an unknown effect. Calor0418 survives for one thing only: calling something that provably isn't a function at all, like an i32.
  • §E{db} now covers §E{db:r}, and the same for net and env. If a function declared the broad "touches the database" effect, and the code inside it only read from the database, the compiler used to complain that the narrow effect wasn't declared. It no longer does — declaring the broad effect covers the narrow ones, which is what everyone already assumed it meant. This only ever makes more programs compile, never fewer.
  • A clearer message when an effect annotation is on something that cannot have one. Writing §E{cw} after a plain i32 — a number performs no effects — used to be silently accepted and quietly re-read as the whole function's effect declaration. It is now Calor0405, the same diagnostic you already get for an annotation on the wrong line, and it tells you both ways to fix it: remove it, or move it onto its own line if it was meant to be the function's own declaration.
  • You can now write down what a callback is allowed to do. Use the same §E{…} tag effects have always used, placed on the same line as the type it describes: §I{Func<i32,i32>:transform} §E{cw} says "this callback may write to the console". You can write one on a parameter, a return type, a §B binding, a §FLD field, a §LAM lambda and a §DEL delegate — eight places in all. A function or method can also declare a placeholder effect, written <eff e> in its type-parameter list, standing for "whatever the caller's function does". One thing changed meaning: an §E{…} written on the same line as a type used to be read as the whole function's effect declaration; now it belongs to that type. Written on its own line, as almost everyone writes it and as the docs have always recommended, nothing changes at all. Two new diagnostics come with it: Calor0405 when an effect row is on the wrong line (naming both ways to fix it), and Calor0404 when an eff placeholder is declared somewhere it cannot be, or is named after a real effect code such as cw. See Effects.
  • The test that judges this release's effects feature was written down before the feature was built. It is called PP-E1, and it lives in the project's proof-point registry (docs/plans/agent-native-gates.md, entry A-1.11). The pass mark, the exact files it runs on, and the ten deliberate bugs it injects into them were all fixed in advance, so the result could not be argued into a success later. It has four possible verdicts — pass, fail, too-noisy-to-tell, and cannot-be-judged — with the rules for each set beforehand. A later correction (entry A-1.11.1) re-measured the "control" for one file, which had been copied from a run that used a switch the test forbids; the original record is left untouched and the corrected one sits beside it.
  • CI now refuses any pull request that edits or deletes an already-frozen line in the project's proof-point registry; only additions are allowed. The append-only check that already protected docs/experiments/registry.json now also runs scripts/check-annex-freeze.py on every pull request and every push to main — using the copy of the script already on main, so a pull request cannot weaken the guard and use it in the same change.
  • §SEMVER{MAJOR.MINOR.PATCH} is a new module-level directive. It did not lex before this release: any §SEMVER line failed with Calor0006, even though the docs described it. Write it as the first line of the module body, e.g. §SEMVER{2.0.0}. Only that exact spelling is accepted — the caret and range forms the versioning page used to show are no longer documented and are rejected with Calor0702, as are the bracket form §SEMVER[…], a leading sign, and any whitespace inside the braces.

Changed

  • .NET method lookup now uses one method identity instead of a handful of loose text strings. Before, asking "what does File.ReadAllText do?" meant passing around a type name, a method name and a list of argument types as plain text, and hoping the spellings lined up. Now one record names the method — where it comes from, what it is called, what it takes, and whether it is an ordinary method, a property read or write, a constructor, or an extension method — and every lookup goes through it. Nothing about your code compiles differently. One diagnostic does change, in one situation: if you call something by a bare name — §C{u} — and the compiler cannot work out what u is at all, it used to say "you cannot invoke a function-typed value of type ?". It now says the call target is unknown, and counts it as unknown, which is the safe answer.
  • The effect checker now asks the type checker what a method call is made on, instead of hunting through the source text for a type name. Two kinds of code that used to be reported as "unknown call" now compile cleanly, because the type checker knew the answer all along: a binding with no written type whose value comes from a .NET call — §B{g} §C{System.Guid.NewGuid}, then g.ToString — and a local that shares its name with a field of a different type. If the type checker looked and could not name the type, the effect checker no longer substitutes a guess — the call is treated as unknown. Lambdas also get a proper function type internally, so "is this a function value?" is answered by the type rather than by matching text like Func<.
  • When the compiler cannot infer the type of a method call's receiver, it now records that as a fact instead of leaving a placeholder. In an editor you get an informational diagnostic (Calor0270) on the two cases you can actually fix — a variable the compiler could not infer a type for, and one whose written type it could not make sense of. Adding an explicit type to the binding clears it. Cases you cannot fix stay silent. Nothing that already inferred a type behaves differently, and command-line output — including calor effects suggest — is byte-for-byte what it was.
  • Files that declare §SEMVER{1.x} (or 0.x) are now refused with an error pointing at #1084. The compiler implements semantics version 2.0.0. An older major (§SEMVER{1.0.0}) stops the build with Calor0701, and the message tells you what to do: migrate the module and declare §SEMVER{2.0.0} after reviewing the nullability rules (Calor0272/0273/0274); see #1084. A newer major (§SEMVER{3.0.0}) also stops the build with Calor0701 ("upgrade the compiler"). A higher minor (§SEMVER{2.1.0}) compiles with a Calor0700 warning. A version that is not exactly MAJOR.MINOR.PATCH, or a second §SEMVER in the same module, is a new Calor0702 error. Files that declare nothing are unchanged.
  • Proof-based guard elision is now on by default. When Z3 proves a contract — with --verify (postconditions) or refinement verification (§PROOF obligations, refinement-type and index-bounds checks) — the generated C# no longer carries that runtime check. No --elide-proven-guards needed (the flag still works and simply restates the default). Only a clean Proven verdict with no assumptions attached qualifies; preconditions and every other verdict keep their guards, and a compile that runs no verification is unchanged. To opt out and keep every guard (the 0.13/0.14 behavior), pass --keep-proven-guards to calor compile, calor run or calor test, set ElideProvenGuards = false on CompilationOptions, pass keepProvenGuards: true to the MCP calor_compile tool, or set <CalorElideProvenGuards>false</CalorElideProvenGuards> in MSBuild. Why now: a CI test suite now checks that "Z3 says proven" and "the check never fails at runtime" agree, and reports zero disagreements across 65 contract shapes (40 of them elidable). One honest caveat: that suite runs the guarded version of the code and checks the shape of the elided version — it does not run the elided version.
  • calor effects suggest takes a variable's type from the compiler's own type information, not from the source text. A variable whose type is only known through .NET metadata resolves to the right type. A call whose receiver the compiler cannot vouch for — an inferred variable it could not type, a chain like a.b.Method, a lambda or function value, or a lowercase name that is not a known variable — is now reported with a new Calor1360 warning and left out of the suggested manifest instead of becoming a manifest entry named after the variable. In --json mode the warning is in diagnostics and data.untypedReceivers gives the count. See the calor effects command.

Fixed

  • Benchmark cost metric no longer under-counts agent output tokens (#881). The agent-loop benchmark scripts read a number from the agent's result file that covered only its final turn, so a run that delegated work to a helper agent, or resumed after its context was compacted, looked far cheaper than it was — one archived run recorded 543 tokens when the model had actually generated 30,084 (55x). Both scripts now use one shared helper that sums every model's output across the whole run, keeps the old number next to the corrected one so the fix can be audited, and warns when the two disagree.
  • The round-trip check no longer fails a pull request because of a test that MediatR's own test suite is known to flake on. The known-flake list was being dropped on the way into the check. Ignored flakes are still listed by name in the report.

[0.14.3] - 2026-08-24

Small fix — §NEW{Type} constructor expressions now correctly reflect that new Foo() is non-null. Improves downstream nullability checks without changing any observable behavior on well-annotated code.

Fixed

  • §NEW{Type} constructor expressions carry NotAnnotated, not Oblivious. A new Foo() is provably non-null by construction, so the BoundType returned by BoundNewExpression now reflects that. Applies to bare §NEW{Foo}, §NEW{Foo} §A x §/C (with constructor arguments), and generic forms like §NEW{List<int>}.

[0.14.2] - 2026-08-24

Correction (2026-09-11): The diagnostic additions below describe binder/editor findings, not production CLI rejection by Calor0272/0273/0274. Other passes may reject independently. See the dated correction above; original text follows.

Threads Calor-native call-site annotations through the binder. Pure-Calor callees now behave symmetrically with BCL calls for the two 0.14 diagnostics that used to skip them: Calor0274 on the argument side, and BoundCallExpression.Type picking up the declared return annotation. Downstream consumers see byte-identical DisplayString for pure-Calor calls; only the annotation channel changes.

Added

  • Calor0274 NullableArgumentToNonNullableParameter fires on pure-Calor call sites, not only BCL-resolved ones. A ?Foo argument into a :Foo parameter on a §C{this.Take} call now surfaces the diagnostic symmetrically with the BCL path.
  • BoundCallExpression.Type on pure-Calor calls carries the declared return-type annotation. -> ?string / -> ?Foo surfaces as Annotated; -> string / -> Foo as NotAnnotated. Previously stayed Oblivious. BCL-resolved returns retain priority.

[0.14.1] - 2026-08-24

Correction (2026-09-11): "Enforcement widens" below overstates the change: binder checks widened, but Calor0272/0273/0274 were excluded from production CLI rejection. Other passes may reject independently. Original text follows.

Widens the 0.14 nullability gate along two axes: whitelisted generic containers over string, and user-declared reference types. The core three diagnostics (Calor0272/0273/0274) fire symmetrically whenever a possibly-null value is funneled into a target that refuses null — now covering Option<T>, List<T>, IEnumerable<T> (and their read-only / interface siblings), and any §CL-declared class. Also fixes a header-pill stale-version follow-up and adds an on-demand verify-release CI workflow that installs the shipped tool from nuget.org and exercises Z3.

Added

  • Whitelisted generic-instantiation nullability. Option<string>, List<string>, IList<string>, IEnumerable<string>, IReadOnlyList<string>, ICollection<string>, IReadOnlyCollection<string> now participate. A source whose payload/element is ?string assigned into a same-shape target whose payload/element is string trips the three diagnostics. Container annotation is orthogonal — only the position-0 argument mismatch is diagnosed.
  • User-declared reference-type nullability. §B{b:Foo} a where a is :?Foo trips Calor0272; the return-site variant trips Calor0273. Fires only when the source is explicitly Annotated — Oblivious sources (the default for any Calor-native value that has not yet been annotated) do not fire, so existing §B{x:Foo} someCall patterns keep working until Calor-native call-site annotation flow lands as a follow-on.
  • On-demand verify-release CI workflow. gh workflow run verify-release.yml -f version=X.Y.Z installs the calor global tool from nuget.org on three OS/arch runners and asserts calor verify samples/Verification/proven-contracts.calr prints Proven: 14, Skipped: 0. Reusable per release from environments whose local dotnet cannot reach api.nuget.org directly.

Changed

  • Reference-expression type flows the declared nullability for user-declared reference types. A §MT{...} (?Foo:a) parameter reference now reads as Annotated Foo at its use site instead of falling through as Oblivious. Leading/trailing ? is stripped from the internal QualifiedName so short-name comparisons work without double-encoding the annotation.
  • Diagnostic messages echo the target shape. 'Option<string>' / 'List<string>' / 'Foo' labels replace the scalar 'string' boilerplate when the target is a generic instantiation or a user-declared type.

Fixed

  • Site-header pill lingered at v0.13.2 after v0.14.0. website/src/lib/version.ts was not in the release skill's version-file checklist. Now it is.
  • Performance-test synthetic fixtures no longer misfire on nullability findings. After the 0.14.0 severity flip promoted the three nullability codes to Error, the taint-analysis performance tests' blanket HasErrors assertion rejected the synthetic modules. Nullability findings on synthetic fixtures are now allowed through — the perf tests exercise timings, not nullability enforcement.

[0.14.0] - 2026-08-23

Correction (2026-09-11): "Nullability enforcement" below describes binder diagnostics, not production CLI rejection by these three codes. The promised legacy Info mode was future work, not a shipped feature. Current compilers refuse older majors with Calor0701. See the dated correction above. Original text follows.

Nullability enforcement lands for scalar :string values and their arrays. The binder now flags possibly-null values funneled into non-nullable :string targets at three sites: variable bindings (Calor0272), function returns (Calor0273), and call-site arguments (Calor0274). Severity is Error by default under the bumped SemanticsVersion.Major = 2; legacy modules keep Info once the module-level §SEMVER directive is threaded through.

Added

  • Three nullability diagnostics. Calor0272 NullableToNonNullableBinding for §B{x:string} from a possibly-null value; Calor0273 NullableReturnFromNonNullable for a non-null -> string returning a possibly-null value; Calor0274 NullableArgumentToNonNullableParameter for a possibly-null argument passed into a :string parameter.
  • Array element types participate. []string is a non-null-element array; []?string element annotations trip the same three codes when the container annotation is orthogonal to the element mismatch.
  • Declared nullability flows through parameters, fields, properties, lambda parameters, and foreach loop variables. All VariableSymbol creation sites now inherit the declared annotation, so references at their use sites report the correct source annotation.

Changed

  • SemanticsVersion.Major bumped from 1 to 2. Signals the new default nullability contract. Existing Z3 verification caches invalidate cleanly under the new version.
  • BCL call signatures carry Roslyn NullableAnnotation. The binder consults the metadata resolution to compute the source annotation for the three diagnostics above.
  • object.ToString() is narrowed to Annotated. Overrides can return null, so the previously NotAnnotated result was theoretically unsound.

Fixed

  • BoundVariableExpression inherits declared nullability from its VariableSymbol. Previously local references silently degraded to Oblivious.
  • NominalBoundType.Equals includes the NullableAnnotation field. An architecture test now guards against annotation-agnostic construction anti-patterns.

[0.13.2] - 2026-08-12

Hardening release. The compiler's control-flow and dataflow foundation is rebuilt on explicit semantics, and the build and release chain becomes hermetic and supply-chain verified.

0.13.1 was tagged but never published. Four publish attempts failed and nothing was ever pushed to nuget.org under that version. Rather than publish it from a main that had since gained a compiler-internals rewrite it did not describe, 0.13.1 is abandoned and its contents ship here, described accurately.

Changed

  • Z3 translation aligned with executable C# semantics. Integral literal width and signedness survive translation, and C# numeric promotions apply across arithmetic, comparison, equality, shifts and overflow. Operations with no executable C# semantics stay fail-closed rather than approximated. Translator semantics are versioned in proof results and cache validity, so existing verification caches are invalidated by this release and proofs are re-established under the corrected semantics.
  • CFG and dataflow rebuilt around explicit semantics. Control flow is constructed from explicit terminators and typed edges rather than positional inference; loop, exception, catch, finally, using, return, throw, break and continue route structurally. Dataflow fails explicitly rather than silently on non-convergence, and initialization, liveness and reaching-definitions analyses are symbol-keyed and semantically ordered.
  • Builds and releases are hermetic and supply-chain verified. Z3 is restored by an explicit bootstrap instead of downloaded mid-build, and every supported binary is verified against committed SHA-256 and byte-size pins before compilation. NuGet versions are centralized with committed lock files and locked restores, alongside offline build/test/pack, corrupt-asset, lock-mismatch and runtime-load gates, SPDX SBOMs and SLSA-style provenance.
  • Round-trip verification is failure-safe. Conversion runs in lossless mode and validates generated C# before publication, failing closed rather than reporting success on build, process or coverage failures. Excluded and failed items stay visible in reports.

Fixed

  • Z3 downloads survive a real outage. Every fetch path used --retry 5 --retry-delay 3 — and per curl(1), setting --retry-delay disables exponential backoff, so the flag that reads as hardening pinned retries to a flat 3s and gave up in about fifteen seconds, shorter than a routine CDN blip. Backoff is restored across every fetch site, including the publish path, and the Windows path — which had no retry at all — now backs off too.
  • MCP tools no longer refuse work because their host is using memory. The heavy-tool admission gate measured whole-process memory and applied everywhere, so a handler embedded in someone else's process charged that process's memory to the next tool call and refused it. The gate now applies only to the stdio server, which owns its process.

Removed

  • VS Code extension support is withdrawn. The extension tree, VSIX release assets, Marketplace publishing, and the single-file publish guard are removed. The language server is unaffected: calor lsp speaks standard LSP over stdio and works with any LSP-capable editor.

[0.13.1] - 2026-08-12

Completes the 0.13 project-model program: the persistent index and calor query ship — the item deferred three times in 0.13 — and four release gates close.

Not 0.14. The "Null-Safe .NET" content (metadata-backed .NET binding, typed semantics, non-nullable reference types, the 2.0.0 self-migration) remains unbuilt and will ship as §SEMVER{2.0.0} when it does.

Added

  • Persistent project index and calor query. calor index build writes a versioned index under obj/calor/; calor index status says whether it can still be trusted. calor query answers six facets: symbol, callers, callees, impact, contracts, assumptions. (A seventh, effects, was added in 0.15 — see calor query for the current list.) Two rules are enforced rather than documented — a stale index never answers, and every answer carries its residual, naming what could not be resolved rather than quietly omitting it.
  • calor rename — rename addressed by symbol identity rather than text, refusing rather than guessing on ambiguity, stale sources, collisions, and declarations split across files.
  • Dogfood utility — a real in-repo tool whose only source is Calor, built and run by CI.
  • Four release gates closed: full-vs-incremental diagnostic identity, index/query correctness, rename (with an apply-recompile-and-run behaviour oracle), and the performance envelope (index build 0.40s against a 30s budget; warm queries under a millisecond against 500ms).
  • Exact-span LSP refactoring — definition, references, and rename distinguish overloads, shadowed locals and fields, and same-spelled symbols across files.

Fixed

  • Module-qualified cross-module calls resolve. Writing §C{Module.Function} previously gave a worse result than the bare form — it was treated as an unknown external call and forbidden inside a pure function even when the callee was pure.
  • Generated C# is guaranteed Roslyn-valid; formatting is lossless and atomic.

[0.13.0] - 2026-08-11

The "Trustworthy Project Model" release. Headline: the structural-binding rebuild is complete — all 60 accepted expression classes bind structurally (the analysis-incomplete instrument is retired because nothing is incomplete), with stable symbol identities, full-signature overload resolution, exhaustive checker traversals, a resolved call graph, and LSP rename/references/cross-file resolution on top.

Release-gate scorecard, stated plainly

  • Green, measured: binding totality (zero incomplete on both measurement corpora, ratcheted in CI); the verifier-vs-runtime differential gate (65/65 modeled forms, 1,170/1,170 cases, zero mismatches, CI-blocking); the migration fixture registry (enforcing).
  • Not met, disclosed: the full-vs-incremental identity harness and the rename harness were registered but never built — they ship unmet and are the top of 0.13.x.
  • Deferred, named: the persistent index / calor query did not ship — the fourth deferral of an item the plan pre-committed not to defer again. First index work item in 0.13.x.

Added

  • Verifier-vs-runtime differential gate: 1,170 deterministic cases across all 65 frozen modeled forms × contract positions × nesting depths × polarity, with a compile-and-execute oracle. CI blocks on zero mismatches.
  • Structural binding for every expression class, with an authoritative dispatch table, reflection completeness tests, and a two-leg corpus ratchet that makes coverage regression a CI failure.
  • Stable SymbolIds and exact identifier spans consumed by the LSP (rename, references, cross-file resolution).
  • Top-level overload sets: duplicate signatures, ambiguity, and no-match are explicit diagnostics — a second same-name declaration no longer silently vanishes.

Changed

  • Proof-based guard elision is now opt-in (--elide-proven-guards): verification verdicts are diagnostic by default; a Proven result keeps its runtime guard unless you opt in. The differential gate that would justify flipping the default is now built and green; the flip is a deliberate next-cycle decision.
  • Verification cache keys are exhaustive and semantics-versioned: content-hashed over the modeled surface, namespaced by compiler-semantics version and solver budget, with collision-unsafe keys refused outright.
  • MSBuild incremental builds fingerprint every diagnostics-affecting input, including referenced assemblies and IL-analysis inputs — flipping any of them against a warm cache forces a recompile.

Removed

  • The osx-x64 Z3 native is no longer shipped. Upstream ships an arm64 binary under the x64 label, so Intel Macs installed successfully and silently lost verification — no honest asset exists to ship. Intel macOS is unsupported for verification (compilation is unaffected; Z3-dependent features report "Z3 unavailable" loudly). A native arch-vs-RID assertion in the packaging workflow prevents any recurrence.

Known channel state

  • The VS Code Marketplace remains at v0.3.8 (expired publish token, maintainer-only); this release does not change that.

[0.12.1] - 2026-08-07

A packaging release. v0.12.0 was tagged but never installable — both publish workflows failed, so nuget.org continued to serve 0.10.0 and the VS Code marketplace continued to serve 0.3.8. Nothing in the language, compiler, or verifier changed here.

Fixed

  • The VS Code extension builds again — though it still did not publish. The build failure below is fixed and all six platform packages are produced, but the publish step then failed on an expired marketplace token, so the marketplace remains at 0.3.8. That token, not the build, is why the extension has failed to publish on every release since v0.4.0 — the last successful publish was v0.3.8. The language server is packed as a single file, which promotes Assembly.Location to a build error. Fixed at both sites — one justified suppression where the empty location is a supported and already-guarded input, and one switch to the single-file-safe RuntimeEnvironment.GetRuntimeDirectory(), since the old expression would have silently emptied the framework probe root used by IL effect analysis.
  • The NuGet publish no longer fails its own checksum gate. The gate was right — the pinned Z3 binaries release had been republished underneath its manifest. Two underlying defects are fixed rather than papered over by rehashing the live assets: the managed Microsoft.Z3.dll was selected with find | head -1 across all build artifacts, and linux-arm64 was built from source despite upstream publishing a binary for it. Both now come from checksum-verified upstream archives, so the published assets are reproducible.
  • The shipped Z3 managed wrapper is now upstream's Release build. Because of that selection race, released packages had been carrying a Debug assembly compiled on a CI runner, with the runner's absolute build path embedded in it.
  • Z3 now loads on ARM64. Upstream's Windows x64 archive ships an AMD64-marked wrapper that cannot load in an arm64 process, and both download scripts had designated it. This affected linux-arm64 and win-arm64; it stayed hidden because x64 CI cannot observe it.
  • Republishing the shared Z3 binaries is now deliberate. The build workflow no longer runs on push — which is what invalidated the pin originally, since the commit that introduced the manifest re-triggered the workflow and republished every asset a day later.

Added

  • CI now exercises both release-only paths. Neither failure was catchable before release day, because the single-file publish and the binary pins were only ever exercised while publishing. The test suite now runs the same single-file publish the extension build uses, and a scheduled check verifies the binary pins daily against both the project's release and the upstream archives.

[0.12.0] - 2026-08-06

The soundness release. This entry covers the v0.11 range as well — there is no v0.11.0 tag; the maintainer folded v0.11 forward, so everything below ships together.

Benchmark and release gates

  • 217-program micro-benchmark: overall 1.32x Calor/C#, with Calor ahead in seven of eight metrics. Comprehension is 1.84x, error detection 1.49x, token economics 1.42x, and information density 0.98x (C# wins).
  • Statistics caveat: all 30 runs are identical deterministic static analyses. Zero-width confidence intervals and the reported p-values describe no run-to-run variance; they do not measure uncertainty over the corpus.
  • These numbers are not the release gates. PP-A1 passed all nine adoption items. PP-W5 recorded a 1.0016 point estimate as "no large tax detected", explicitly not a proof of parity.

Six distinct false-proof vectors were closed — cases where a proof deleted a runtime check that would have failed. Enumerated so the count is checkable: (1) D4 non-ordinal string comparison modes; (2) D4 bare StartsWith/EndsWith/IndexOf, which use the current culture in .NET and so diverge with no mode argument present to signal it; (3) D3, Z3 strings are null-free; (4) D12, Z3 models strings as UTF-8 bytes where .NET counts UTF-16 units; (5) D14, array and user-type sorts are total and non-null; (6) the third $length mint site, which the D14 fix's own first cut missed.

Items (1), (2), (3), (4) and the array half of (5) carry a recorded calor run versus calor run --verify reproduction. Only the user-type half of (5) rests on inspection of the encoding; item (6) is not reachable from --verify at all — it served a false proof to agents through the MCP refine tool, which is how it surfaced. That is six closures, not six review rounds, and the number of vectors still unfound is not knowable and is not claimed to be zero.

What closed the class was a change of mechanism, not a better enumeration. Hand-enumeration was tried at three levels — divergence rows, then Z3 sorts, then mint sites — and missed something at every one. Proofs carried by a sort Z3 models as total are now demoted to Assumed, which never elides, so the class closes by construction.

Added

  • calor import <package> generates effect manifests from real assemblies in three tiers — IL-derived, curated, and unresolved, which is surfaced loudly and excluded rather than defaulted to pure. Nothing is ever emitted as verified. Validated on Serilog and MediatR.
  • calor review-packet leads with the unproven remainder: seven-status counts, assumption lists, vacuity flags, counterexamples, and per-module interop/waiver fractions with waiver disclosure on the first line.
  • Adoption playbook and a tested eject path — an 11-test suite that compiles and executes ejected C# to pin exactly what survives leaving Calor.
  • Effect soundness (WS-W2). Invoking a delegate held in a value is now an error rather than an assumed-pure no-op; effect variance is checked on both the override and interface legs; interop is Assumed and propagates transitively; the known-pure list no longer contains mutators; and both silent => Empty catch-alls are replaced by exhaustive switches. --enforce-effects defaults on.
  • A self-contained Calor.Sdk package. The SDK carries its build task, compiler/runtime dependency closure, and per-RID Z3 natives. CI installs the packed SDK into a project with no project references and requires a real Z3 counterexample from the MSBuild task.
  • Telemetry is opt-in and metadata-only. Default invocations send nothing. Enabled telemetry excludes source, paths, diagnostic/exception text, stack traces, hostnames, help queries, and content hashes.

Changed

  • Postcondition elision is withdrawn for any signature naming an array or a non-primitive type. Z3's array and user-type sorts are total and non-null; .NET's are nullable references. Coarse on purpose — being wrong in the narrow direction deletes a runtime check, being wrong in the broad direction costs an optimization. Contract proving and reporting are unchanged.
  • The type checker is on by default, which required first fixing the checker: turning it on produced 92 test failures, every one of them a working program the checker refused. They trace to roughly eight defect classes — char alone accounts for 37 of the 92 — covering object, decimal, every sized integer, arrays, string concatenation and static member access. --no-type-check opts out on build; CALOR_NO_TYPE_CHECK=1 opts out everywhere, including run, test, watch and the MSBuild task inside the SDK.
  • Conversion no longer substitutes silently. Unknown operators, patterns, and compound assignments escalate to loud interop blocks instead of quietly becoming something else.
  • Conversion success is loss-aware. Generated output is compile-validated; the success line appears only for zero-loss output, while text and JSON name every semantic loss and preserved interop member.

Fixed

  • A quantified-contract lowering bug that discarded implication antecedents was fixed before release. It demonstrated that encoding soundness and runtime lowering soundness are separate surfaces.
  • Direct self-recursion no longer fails as an unknown effect target.
  • calor_refine no longer reports success after a failed compile.
  • Z3-backed CI tests can no longer silently turn green by skipping when the native solver is unavailable.
  • The VS Code package now includes vscode-languageclient, fixing its long-standing activation defect. The Marketplace publish itself did not succeed; the listing remains at v0.3.8.

Known limitations shipping with this release

All pre-existing rather than regressions: §MT contract violations report the wrong function id; a contract call carrying a keyword argument crashes the translator instead of diagnosing it; str is not yet non-nullable at the binder, which is what would let the string demotion be lifted; and IL-analysis inputs are not all included in the MSBuild task's warm-cache options hash.

[0.10.0] - 2026-07-30

The Guarantees release: every verification verdict an agent sees is now sound, honestly statused, and provenance-aware — and the upgrade is measured. On a paired A/B probe epoch, all three seeded contract defects moved from runtime exceptions to build-time refutations under the new verify gate, with zero catch regression and zero contract weakening across 30 eligible runs.

Added

  • Sound result binding (#807 closed). Result-referencing postconditions are checked against the encoded function body (SSA-style substitution over immutable §B chains, guard-clause branching, and value returns); bodies outside the encodable surface report honest unsupported — never a refutation against an unconstrained result.
  • Seven-status verdict vocabulary. assumed (holds conditionally on a named assumption set — division side conditions are the first producer; never elides runtime checks) and unavailable (no solver present) join the five existing statuses; vacuous proofs carry a vacuous flag and keep their runtime checks. Envelope schema 2.0.
  • Positive modeled-forms whitelist. unsupported is decided by an enumerated, CI-byte-checked whitelist rather than by accident; out-of-surface contracts name the offending construct.
  • Cross-module linking (#809 closed). Bare-name calls across modules emit qualified C# and link under MSBuild/csc, covering both the CLI multi-input driver and the MSBuild task.
  • MSBuild verify gate. CalorVerify=true runs Z3 contract verification inside builds; refutation warnings surface as MSBuild warnings with declaration attribution.
  • Mechanical weakening check. calor verify --weakening-check decides whether a declaration's contract was weakened between two versions (postcondition relaxed or precondition strengthened), with an intact-or-strengthened verdict for CI gating.

Fixed

  • Precondition guards are never elided on satisfiability results (#755). Only genuine ∀-proofs elide checks.
  • Contract-level vacuity is loud. Jointly-unsatisfiable precondition sets no longer mint silent Proven verdicts.
  • % semantics corrected (bvsrem, C# remainder), out-of-32-bit literals refused instead of silently wrapped, and the flagship proven-contracts sample repaired to prove 14/14 soundly.

[0.9.0] - 2026-07-29

The Loop release: the agent edit–feedback cycle is now a first-class, measured product surface — one machine-readable envelope across every CLI command and MCP tool, a transactional MCP write path with warm project sessions, and millisecond-scale edit feedback. Measured on paired A/B epochs: the new loop tooling cuts agent tokens-to-green ~35%, and Calor's enforcement caught 9/9 seeded defects where the C# toolkit caught 4/9.

Added

  • Transactional MCP write path with project sessions. calor_session_open/calor_session_close hold parsed project state; calor_file_write runs heal → check → atomic-apply-or-reject with the full diagnostic envelope on reject, project-wide reference checking, and write confinement under a pinned root (calor mcp --root).
  • Fault-tolerant parse mode. Broken files produce partial ASTs plus diagnostics (never a compile pass), so sessions keep working context across syntax errors; parser depth guards prevent runaway nesting from killing the process.
  • Warm feedback. Sessions reuse parse/bind state and a lazy call-target index; MCP edit→envelope latency measures P50 2 ms / P99 9 ms on a pinned 10k-line fixture, with calor watch incremental rebuilds at P50 27 ms. Telemetry streams (mcp-write/2, watch-rebuild/1) journal per-edit latency breakdowns.
  • Verification-outcome honesty. All verification outcomes route through a single five-status choke point (proven | refuted | unknown | timeout | unsupported); refuted contracts carry concrete counterexample models; a committed fixture corpus pins every status in CI.

Changed

  • One envelope everywhere. Every diagnostic-producing CLI command and MCP tool emits the shared envelope schema v1.1 (diagnostic code, span, enclosing declaration ID, severity, fix hint, verification payload) — 100% coverage over the enumerated surface, enforced by CI conformance tests. See CLI envelope schema.
  • calor verify exits 1 on refuted contracts (missing files and compile errors too); proven/inconclusive outcomes keep exit 0.

Fixed

  • CLI exit codes propagate on error paths across verify, convert, coverage, benchmark, migrate, and the effects subcommands (previously stomped to 0 by the invocation pipeline).
  • JSON mode always emits exactly one envelope document, including on error paths, with CLI-band diagnostics (Calor1310–1312).
  • Counterexample rendering filters internal solver variables across all producers.

[0.8.0] - 2026-07-23

The correctness-hardening release: every exit-0-then-broken-dotnet build hole in the binding/rebind/shadowing family is now caught at calor -i with a clear diagnostic, plus a systemic pass to keep every diagnostic agent-friendly.

Breaking

  • New hard-error diagnostics (Calor0254–0258) reject programs that previously compiled. Programs that used to exit 0 and then fail a downstream C# build now fail earlier at calor -i: array→concrete-collection (Calor0254), enclosing-scope shadowing (Calor0255), type-changing mutable rebind (Calor0256), foreach iteration-variable write (Calor0257), and same-scope duplicate §B (Calor0258).
  • Contract-verification result codes renumbered Calor0700–0705 → Calor0710–0715. Tooling filtering verification output on Calor0700–Calor0705 must switch to Calor0710–Calor0715; Calor0700/Calor0701 now unambiguously mean the semantics-version diagnostics.

Added

  • Array-to-collection type error (Calor0254). calor -i rejects binding, returning, reassigning, or passing an array where a concrete generic collection (List<T>, HashSet<T>, …) is expected — e.g. §B{lines:List<str>} §C{File.ReadAllLines} — instead of emitting C# that fails with CS0029. Collection interfaces (IList<T>, IEnumerable<T>) are still accepted.
  • Local-shadowing error (Calor0255). Rejects a §B that shadows a local, parameter, or loop variable already in an enclosing scope (CS0136), in both directions (an inner binding, or a loop variable reusing an enclosing name). Mutable rebinds (the accumulator idiom) and legal field-shadowing are unaffected.
  • Type-changing / mismatched mutable-rebind error (Calor0256). Rejects a mutable §B reassignment whose value is a different, non-convertible type — from an explicit annotation, a literal, or (new) an inferred reference/call return type. Uses primitive-category comparison, so implicit numeric widening is never falsely flagged.
  • Foreach iteration-variable rebind error (Calor0257). Rejects writing to a read-only §EACH/§EACHKV iteration variable (CS1656); §L for-loop variables and §EACH index counters stay reassignable.
  • Same-scope duplicate-binding error (Calor0258). Rejects two §B reusing a name in one scope (CS0128); the C#→Calor converter now emits array/collection reassignments as reassignments rather than duplicate creation blocks.
  • calor convert --passthrough. The CLI C#→Calor converter can now preserve unconvertible members verbatim as §CSHARP interop blocks so the output always parses, reporting how many members were preserved.
  • Surface-spelled diagnostics. Every diagnostic that echoes a type now prints the compact surface spelling (i32, str, bool, Option<str>) instead of the internal form, guarded by a test that scans the whole corpus for leaks.
  • Exemplar compile-checking. calor self-check docs now compiles every program in the agent syntax exemplar all the way through Roslyn's semantic model, catching type errors the Calor pipeline emits without complaint.
  • Converter §CSHARP fallback. When a C#-preserving mode is active and the emitted Calor for a member would not parse, that member is re-emitted as a §CSHARP interop block so the output is always valid Calor.

Fixed

  • Scope-aware mutable-rebind codegen. The emitter now tracks declared variables in a scope stack, so a mutable §B in a closed sibling block re-declares (valid) rather than assigning to an out-of-scope variable (CS0103); loop variables, catch/using bindings, and parameters are modeled too.

Changed

  • CI hardening. The Calor-first guard exempts the tests/ tree (xUnit tests are C# by nature), and the Z3 download retries transient network failures instead of breaking the build.

[0.7.0] - 2026-07-16

The agent dev-loop release: Phase 1 of the agent-native strategy complete — six items, each hardened by adversarial review.

Added

  • Source maps. The compiler emits #line directives mapping generated C# back to .calr source: downstream compiler errors, runtime stack traces, and debugger sessions report your .calr file and line instead of a generated .g.cs location.
  • calor run and calor test. One command to execute or test any .calr file or directory — no MSBuild wiring. Effects enforcement on by default (--permissive to relax, violations now visible as warnings), --verify/--contract-mode pass-through, process timeouts, real exit codes.
  • Structured diagnostics. --format text|json|sarif on compile and lint: stdout is always one machine-parseable document (status goes to stderr), early-exit errors included, honest exit codes. Schema documented in structured output.
  • calor format --heal. Best-effort source-level repair for files too broken to parse — indentation re-derivation, closer stripping, chain-clause re-alignment — with every ambiguous decision reported per file:line. New fixable indentation diagnostics (Calor0008/Calor0009/Calor0117) carry one-pass machine-applicable edits, and the MCP calor_check tool auto-heals in agent loops.
  • calor self-check docs. Machine-verifies the docs against the implementation — §-keywords vs the lexer, diagnostic codes, effect codes, and every fenced Calor example parsed with the real parser — and runs in CI, so documentation drift is now a build failure.
  • calor watch. Debounced incremental recompiles with NDJSON structured output; incremental caching shares the MSBuild engine with hardened trust boundaries (content hashed from compiled bytes, effect summaries required for cache hits, outputs verified by hash). Plain-compile caching is opt-in via --cache.

Fixed

  • Obligation fact scoping: guard facts were collected function-wide, so contradictory sibling branch guards could vacuously discharge every proof obligation in a function; facts are now scoped to the range they dominate with an UNSAT pre-check.
  • NullDereferenceChecker order-dependent classification of unwrap_or/unwrap_or_default.
  • Option/Result combinators resolve through effect manifests as pure-modulo-arguments; Calor surface types (?T, T!E) map to runtime manifest keys.
  • Agent-facing docs corrected (keyword accuracy, effect-code completeness) and now drift-guarded in CI.

Changed

  • New diagnostic bands: 1300–1399 (CLI findings and command-level errors, incl. doc-drift codes).
  • Calor0008/Calor0009 warnings fire on legacy tab/4-space-indented files (with attached one-pass fixes).

[0.6.8] - 2026-07-01

Added

  • calor fix --heal-closers — a source-level CLI that finishes the Calor0830 auto-heal story. Closer-form syntax (§/F, §/M, §/L, …) hard-errors at parse time, so calor format / calor lint --fix — which parse the file first — can never read, let alone heal, it. The new calor fix --heal-closers <root> [--log <file>] [--revert] [--dry-run] deletes legacy structural closers at the source level, rewriting a file into canonical indent-only form, and --revert --log restores it byte-exactly. It is lexer-backed, so a §/F embedded in a string literal or a // comment is left untouched, and removals are recorded as UTF-8 byte ranges so revert is exact even across non-ASCII text and CRLF endings. This delivers the CLI heal command deferred in v0.6.6.

Changed

  • Single-sourced return-value classification in a shared Analysis/ReturnShape. The "does this owner return a value" classification — for void/async-void functions and methods, iterators, constructors, setters, and event accessors — was duplicated between the Calor0205 pass and the contract verifier's result check. Both now defer to one classifier, which distinguishes the runtime shape (folding in async/iterator lowering) from the narrow header predicate (which keeps result referenceable in an iterator's postcondition, since an iterator still declares IEnumerable<T>). The refactor is behavior-preserving and leaves codegen untouched, retiring the shared-ReturnShape follow-up noted in v0.6.7.

[0.6.7] - 2026-07-01

Added

  • Calor0116 — malformed four-field §F/§AF function headers are now a parse error. A header like §F{f1:Add:i32:pub} looks reasonable but is silently wrong: function headers take at most {id:name:visibility}, and the return type belongs in the signature ((...) -> type). Left unflagged, the parser read the extra field's type as the visibility and discarded the real visibility, emitting a void method (e.g. void Add() { return 0; }, then a CS0127 in the generated C#). The parser now reports it up front. Only §F/§AF are affected — §MT/§AMT legitimately take a fourth modifier field.
  • Calor0205 — a value returned from a no-value owner is now a hard error. An always-on pass flags a value-returning §R expr in a void/async-void function or method, an iterator, a constructor, a property/indexer set/init accessor, or an event add/remove accessor — cases that previously produced non-compiling C# (CS0127 / CS1622), the classic one being a correct void header followed by §R INT:0. Being always-on, it is deliberately conservative for zero false positives: it flags only expressions that are definitely a value and never a valid void statement-expression, leaving calls, new, await, and ++/-- untouched. Together with Calor0116, this closes the deferred "value returned from void function" gap from v0.6.6.

Documentation

  • Every agent-readable surface was swept to indent-only syntax. The MCP primer, the editor/agent instruction templates, README.nuget, the evaluation skills doc, and the correct-Calor fields of the JSON resources were corrected so no agent-facing material still shows removed closer-form tags, four-field headers, or other syntax the compiler rejects — with a new compile-time guard that fails if any surface drifts back to teaching non-compiling forms.

[0.6.6] - 2026-07-01

Fixed

  • The calor://primer MCP resource now compiles. The agent primer served at calor://primer taught syntax the compiler rejects today — closer-form tags (§/F, §/M, §/I, §/L), ULID IDs, §RESULT, and empty §R — so an agent onboarded from it at session start wrote non-compiling Calor. It was rewritten to be fully indent-only and empirically compilable, with 3-field §F headers, arrow signatures, BCL-only effectful calls, a "Common mistakes" section, and a quick reference.
  • Calor0830 (legacy closer form) is now auto-healable. Its remediation previously told users to run calor format, which parses the file first and aborts on the very error it was meant to fix — a dead end. The diagnostic now attaches a quick-fix that deletes the closer line, surfaced through the LSP quick-fix and the calor_check apply MCP tool, and the message explains the block ends at its body's dedent.

Documentation

  • Teaching and reference docs no longer show removed closer-form syntax. The Markdown docs still claimed closers were "still accepted" and showed closer-form / stale pseudo-syntax, so an agent following Calor's own docs wrote non-compiling Calor. The false claims were corrected and the stale if / loop / match / class / try-catch examples were modernized to current indent-only syntax — each verified to compile.

Tests

  • Compile-time guards keep the primer honest in both directions. PrimerCompilesTests proves every correct module the primer teaches compiles under the same options calor_compile uses by default; PrimerMistakesRejectedTests proves every example in the primer's "these do NOT compile" section genuinely fails to compile — so Calor's own onboarding materials can't drift into teaching non-compiling code. Drift guards keep the curated mistake set in sync with the primer.

[0.6.5] - 2026-06-30

Fixed

  • The TokenEconomics benchmark metric now reports the composite it computes (the discarded-composite bug, #668). The calculator computed a composite advantage — the geometric mean of the token, character, and line ratios — and then discarded it, reporting the raw token-count ratio only despite the metric being named CompositeTokenEconomics. The category now reports the composite. The metric is deterministic, so its 95% CI equals its point estimate. This lifts the headline numbers — TokenEconomics from 1.11× (token-only) to 1.42× (composite) and overall from 1.28× to 1.32× — purely as a measurement correction, not a Calor improvement. The honest caveat stands: Calor still uses more raw tokens than C# on small programs (the §-sigil premium), but is more compact once character and line counts are included.

Changed

  • The deferred v0.7 TokenEconomics gate was recalibrated against the corrected metric. The old token-only target (lower-95%-CI > 1.122) is superseded by a composite gate of ≥ 1.40×, anchored to the measured 1.42× baseline. The public Token Economics metric page was corrected — it had previously and incorrectly reported the category as "C# wins".

[0.6.4] - 2026-06-16

Fixed

  • Parser: an elided §C call no longer steals the parent block's terminating Dedent. An elided §C{X} / §C{X} §A arg call that was the last statement of a function body followed by a sibling declaration (e.g. another §F) previously failed with Calor0100: Expected statement but found Func. Discovered while modernizing the TypeSystem sample for this release.

Added

  • 7 new TokenEconomics benchmark fixtures broadening corpus coverage of v0.6.3 expression-context call elision and v0.6 bind-inference (parser, formatter, delegation, and aggregation shapes), plus two neutral controls. Honest note: the TokenEconomics category scores raw token count only, and Calor's §-sigils cost more tokens than the equivalent C# on small programs, so these representative fixtures left the category at 1.11x — the v0.7 TokenEconomics gate remains open.

Internal

  • The TypeSystem sample and the 04_option_result E2E scenario were modernized to canonical v0.6.3 syntax — a legacy triply-nested §OK{§ARR…} array form (an artifact of mass C# → Calor conversion) was replaced with the canonical §OK value / §ERR "msg" short form, fixing the generated C# from Result.Ok<object, string>(new object[]{…}) to Result.Ok<int, string>(100).

Documentation

  • v0.6 bind-inference RFC §7 — the Calor0250 open question is resolved. The diagnostic has always shipped as a hard error; the "promote warning→error in v0.7?" bullet was a stale carry-over and is now marked resolved, backed by a permanent corpus-clean test pinning zero bind-inference firings across the sample and benchmark corpus.

[0.6.3] - 2026-06-13

Added

  • calor fix --elide-call-closers bulk migrator. New calor fix subcommand that rewrites existing .calr source trees to the v0.6.x call-closer-elided form: zero-arg §C{X} §/C → §C{X} and same-line one-arg §C{X} §A arg §/C → §C{X} arg. Multi-line forms, named-arg (§A[name] x), multi-arg, and ref/out/in arg modifiers are left untouched. Includes a canonical-emit safety net (re-parse the migrated source, re-emit both ASTs, drop the file's edits on any divergence) and --revert --log <file> for byte-for-byte round-trip. Mutually exclusive with --drop-structural-ids and --compact-ids; supports --dry-run.
  • LSP quick-fixes for strict bind-inference diagnostics Calor0251/Calor0252/Calor0253. Each diagnostic now ships a SuggestedFix that inserts the recommended :type annotation right before the closing } of the bind's attribute block. Templates: :Option<object> (for §NN), :object? (for null), :Vec<object> / :Map<object, object> / etc. arity-aware per the matched generic factory, and :f64 (for ambiguous numeric).
  • Calor.LanguageServer.DocumentState.Reanalyze now runs BindValidationPass so strict-bind diagnostics (and their quick-fixes) surface in editors; previously the LSP only ran the lexer/parser/binder and these diagnostics were CLI-only.

Changed

  • Expression-context §C calls now elide §/C by default for one-argument forms. CalorEmitter.Visit(CallExpressionNode) extends the v0.6.1 zero-arg elision and the v0.6.2 stmt-context one-arg elision to expression context: §C{target} arg (no §A, no §/C) when the argument is unnamed, the rendered first token is in the StartsWithExpressionStarter whitelist, and we are not inside an inline-sibling context.
  • Strict bind-inference diagnostics Calor0251/Calor0252/Calor0253 are now default-on. These flag bindings that cannot infer a concrete type without an explicit :type annotation: untyped §NN/null, well-known generic factory calls (Vec.empty, List.empty, etc.), and binary ops mixing integer and floating-point literals. Opt out for one release with --no-strict-bind-inference (CLI) or CompilationOptions.StrictBindInference = false (SDK).

Fixed

  • Parser: Calor0150 no longer fires across sibling-statement boundaries. When the next expression-start token after a one-arg elided call is on a different line, it is a sibling statement, not an ambiguous second positional arg.
  • Emitter: §LAM body, §WITH target, and §LIST/§HSET element emit sites now use AcceptInInlineSibling. These same-line sibling positions previously used raw node.X.Accept(this), which could silently corrupt the AST after the one-arg expression-context elision landed.

[0.6.2] - 2026-06-10

Added

  • Elision-aware TokenEconomics benchmark fixtures (VoidSequence, LogPipeline, PairLogger) exercise the new statement-context §/C elision path. Corpus is now 210 programs.

Changed

  • Statement-context §C calls now elide §/C by default (when safe). CalorEmitter.Visit(CallStatementNode) rewrites zero-argument calls as §C{target} and one-argument unnamed calls (with safe-prefix arguments) as §C{target} arg. Mirrors the v0.6.1 behavior for expression-context calls. Elision is gated by UseImplicitCallCloser and suppressed in inline-sibling contexts.

Removed

  • calor diagnose CLI command removed. Deprecated in v0.5.x with a stated removal target of v0.6.0; v0.6.2 completes that deprecation. For machine-readable diagnostics use the calor_check MCP tool with action: "diagnose" (or calor_compile with automatic fix application).

Fixed

  • Contract verifier: class methods, user-defined types, and visibility preservation. ContractSimplificationPass now preserves Visibility so the verifier reaches §MT members. ContractVerificationPass walks class-method bodies. Z3 translator gained support for user-defined types and dot-path field access (a.b.c).

[0.6.1] - 2026-06-09

Changed

  • ConversionContext.UseImplicitCallCloser now defaults to true (was false in v0.6.0). The C# → Calor converter now elides §/C for zero-argument calls by default, producing more idiomatic Calor output. The opt-out (UseImplicitCallCloser = false) is preserved.

Compatibility

  • Calor source emitted by v0.6.1 may not parse on v0.6.0 or earlier calor toolchains. The new default emits more zero-arg §C calls without explicit §/C. Sources that exercise the newly-fixed parser layouts (zero-arg §C immediately before Dedent, or followed by a same-column sibling opener) will mis-parse on v0.6.0. To produce v0.6.0-compatible output from v0.6.1, use any of:

    • CLI single-file: calor convert --explicit-call-closers <input.cs>
    • CLI project migration: calor migrate --explicit-call-closers <path>
    • MCP calor_convert / calor_migrate: "explicitCallClosers": true
    • SDK: new ConversionOptions { UseImplicitCallCloser = false }

    Round-trip (C# → Calor → C#) remains semantic/structural; the intermediate .calr is intentionally not byte-identical to v0.6.0 converter output unless the opt-out is used.

Fixed

  • Parser: §C standard form no longer swallows trailing Dedent. Previously, a zero-arg §C immediately followed by the end of its enclosing block could corrupt the structural parse of the surrounding method or §IF body. Now correctly distinguished from indent-aware block terminators.
  • Parser: §C no longer absorbs a same-column sibling structural opener on the next line. A sibling §IF / §MATCH / §NEW on the next line at the same indent is no longer silently absorbed as the call's inline argument. Both ParseCallExpression and ParseCallStatement now gate the inline-arg form on the candidate argument starting on the same source line as §C{target}. See Calls reference rule 6.
  • Emitter: zero-arg §C inside an inline-sibling context now keeps explicit §/C. With the new default, naively eliding §/C from a zero-arg call emitted inside another call's §A chain or inside an array/§ARR/§ROW/§SALLOC/§IDX2D initializer caused silent AST corruption (e.g. M(A(), 2) round-tripping as M(A(2))). The emitter now tracks an inline-sibling-context counter; zero-arg §/C elision is suppressed whenever a call is emitted inside such a context. Top-level / leaf-position calls (binding initializers, return values) still elide as before.
  • Parser: §C statement form now supports zero-arg implicit close before sibling statements. Previously a zero-arg §C{target} followed by a sibling statement on the next line at the same indent reported Calor0100. The statement-form parser now recognizes the zero-arg implicit close when the current token is not §A, §/C, Dedent, or Eof.

[0.6.0] - 2026-06-04

Added

  • §C call-closer elision (RFC v0.6-call-closer-elision). Expression-context §C{target} calls may now omit the trailing §/C in two cases: (1) zero arguments — §B{n} §C{items.Count} is equivalent to §B{n} §C{items.Count} §/C; (2) exactly one inline argument (no §A) — §B{y} §C{Math.Abs} x is equivalent to §B{y} §C{Math.Abs} §A x §/C. The parser disambiguates nested elided calls (e.g., §C{Foo.bar} §C{Baz.qux} y ≡ Foo.bar(Baz.qux(y))) by counting consecutive §/C closers relative to enclosing §A depth. Trailing member access on inline arguments binds to the argument (§C{Identity} obj?.Length ≡ Identity(obj?.Length)); trailing member access on zero-arg calls binds to the call result (§C{Maybe}?.Length ≡ Maybe()?.Length). The explicit §A / §/C form continues to parse unchanged. See Calls reference.
  • Calor0150 AmbiguousCallContinuation — New diagnostic in the reserved Calor0150-0159 range. Fires when an elided §C already consumed one inline argument and is followed by either a second expression-start token or a §A token. The fix message recommends the explicit form.
  • ConversionContext.UseImplicitCallCloser emitter flag. Opt-in property on Migration/ConversionContext. When true, CalorEmitter elides §/C for zero-argument calls. Default false for backward compatibility. One-argument elision is intentionally deferred to v0.6.1 pending context-aware tracking inside Lisp argument lists.
  • §B bind-inference formalization (RFC v0.6-bind-inference-formalization). The four supported §B forms — §B{name} (requires initializer), §B{name} initializer (inferred), §B{name:type} (explicit, no initializer), §B{name:type} initializer (explicit wins) — and the binder's shallow inference rule are now documented at Bindings.
  • Calor0250 BindRequiresTypeOrInitializer — §B{name} with no :type and no initializer is now a hard error reported through BindValidationPass. Replaces the pre-v0.6 silent fallback that bound x as INT and produced wrong-typed C# with no diagnostic.
  • Calor0251 / Calor0252 / Calor0253 strict-mode bind-inference diagnostics (opt-in via --strict-bind-inference). Three new diagnostics in the Calor0250-0259 range, each silenced by an explicit :type annotation, scheduled to become default-on in v0.7:
    • Calor0251 BindCannotInferNullLiteral — fires on §B{x} §NN or §B{x} null.
    • Calor0252 BindCannotInferGenericReturn — fires on §B{x} §C{Vec.empty} §/C and other well-known generic factory targets.
    • Calor0253 BindAmbiguousNumeric — fires on §B{x} (+ INT:0 FLOAT:0.0) — a binary op mixing integer and floating-point literal operands.
  • Two new syntax-reference pages — Bindings and Calls, with full disambiguation tables, examples, and diagnostic catalogues.
  • v6 compact stable identifiers (default). IdGenerator.Generate(IdKind) now mints 12-char Crockford-lowercase compact IDs (f_7k9m2npqrstv). The legacy 26-char Crockford-uppercase ULID form (f_01J5X7K9M2NPQRSTABWXYZ12) remains accepted by the parser, validator, and migration tooling. Saves ~9.7 tokens per ID in agent-facing serialisations.
  • calor fix --compact-ids <root> — bulk repo-wide migrator from legacy ULID payloads to v6 compact payloads. Two-pass design with deterministic compact derivation, within-file and cross-file collision detection, and byte-exact revert via --revert --log <file>. Idempotent on already-migrated source.
  • IdValidator accepts both compact and legacy ULID forms. New predicates IsCompactId, IsLegacyUlidId, and IsCanonicalId. New Calor0821 LegacyUlidPayload diagnostic code reserved for the opt-in lint that flags ULID payloads.

Changed

  • Migration/CalorEmitter.Visit(CallExpressionNode) — Zero-argument calls in expression context now conditionally elide §/C when ConversionContext.UseImplicitCallCloser is true. Multi-argument and one-argument paths are unchanged in v0.6.0 (zero-arg-only elision) pending the v0.6.1 context-aware enablement.

Fixed

  • Binder no longer silently defaults §B{x} to INT. A §B{name} with neither a :type annotation nor an initializer was silently treated as INT by the pre-v0.6 binder, producing wrong-typed C# with no diagnostic. v0.6 surfaces this as Calor0250.

[0.5.1] - 2026-06-03

Added

  • Indent-only Calor source — Calor source is now indent-delimited end to end. The parser, migration emitter (Migration/CalorEmitter), and calor format all produce/accept indent form as canonical. Closer-form structural tags (§/F{id}, §/M{id}, §/I{id}, §/L{id}, …) have been removed from the emitter and rejected by the strict CLI compile path. The 09_codegen_bugfixes self-test fixture was migrated alongside; round-trip emission stays byte-identical.
  • calor --input … --output … --allow-legacy-closers — Escape hatch on the CLI compile path for users mid-migration. Default is strict (legacy closer form produces Calor0830); calor format rewrites a file in canonical indent form.
  • CompilationOptions.RejectLegacyClosers — Opt-in compilation flag; the CLI sets it to true by default, other API surfaces (MSBuild <CompileCalor>, MCP tools, LSP) keep the lax default for now.
  • calor fix --drop-structural-ids <root> — Bulk, mechanical, byte-reversible source rewriter that strips {id} from structural closing tags. Records every removal in a migration.log.json and supports --revert --log <file> to restore the original bytes exactly. See docs/cli/fix.md.
  • Calor0820 LegacyStructuralId / Calor0830 LegacyCloserForm — Opt-in lints flagging legacy IDs on closing tags and legacy closer-form structural tags, with fix patches pointing at calor fix and calor format respectively.
  • Optional closing-tag IDs everywhere — Structural closing tags (§/M, §/F, §/AF, §/L, §/I, §/TR, §/CL, §/IN, §/PR, §/MT) may omit the trailing {id} block. Both forms continue to parse; the parser pairs closers with their nearest matching opener by structural nesting.
  • Phase 5 — Product docs migrated to indent-only syntax. README, docs/, and website/content/ now teach indent-form Calor as the canonical surface; closer-form is mentioned only in legacy callouts that point at calor fix for migration.

Changed

  • Lint no longer flags leading indentation or blank lines. With indent form now canonical, the two formatting lint rules introduced for the closer-form "agent-optimized" surface (leading whitespace, blank lines) have been removed from both calor lint and the check MCP tool. Indentation and blank lines are first-class.
  • Benchmark metric calculators score indent form, not closer tags. The four heuristic calculators under tests/Calor.Evaluation/Metrics/ used to award credit for the presence of paired structural closing tags as a proxy for "scope boundaries are explicit". With indent form canonical, the dedent IS the scope-boundary signal, so those bonuses now reward indented body lines per structural opener. Net score magnitude is preserved.

Fixed

  • Migration/CalorEmitter.Visit(CatchClauseNode) — Catch filters now emit §WHEN (matching the token form the parser produces and the §WHEN already emitted by match-arm guards). Previously emitted a bare lowercase WHEN that did not round-trip cleanly.
  • MatchExpressionNode as a §B binding initializer — match expressions used as binding initializers now emit indented §K arms relative to the enclosing block. Previously hardcoded 2/4-space indents could trigger Calor0099 dedent errors when the binding lived inside a deeper function body.

[0.5.0] - 2026-04-22

Added

  • Roslyn 5.3.0 upgrade — Migration pipeline now uses Roslyn 5.3.0 (C# 14 support), enabling conversion of modern C# files using lambda parameter modifiers, out in lambda parameters, and other C# 13/14 features.
  • LanguageVersion.Preview parse option — The C# parser now accepts the broadest possible C# syntax.

Changed

  • Non-exhaustive match on Option<T> / Result<T,E> is now an error — exhaustive match on known sum types is mandatory syntax (Calor0500).
  • Microsoft.CodeAnalysis.CSharp upgraded from 4.8.0 to 5.3.0 across all projects.

[0.4.9] - 2026-04-21

Added

  • Cross-assembly IL analysis — Opt-in compile-time analysis that traces method calls through referenced .NET assemblies to discover effects not covered by manifests. Enabled via <CalorEnableILAnalysis>true</CalorEnableILAnalysis>. Handles async state machines, iterator methods, delegate creation, and virtual dispatch. See Cross-Assembly IL Analysis guide.
  • Cross-module effect propagation — Multi-file Calor projects now verify effect contracts at file boundaries. A caller that invokes a public function declared in another module must declare that callee's effects. Violations produce Calor0410; public functions without §E produce the new Calor0417 warning.
  • Multi-file CLI — calor --input a.calr --input b.calr compiles multiple files and runs the cross-module effect pass across them. Single-file invocations are unchanged.
  • MSBuild cross-module enforcement — The CompileCalor task runs the cross-module pass over every .calr in the project. Works correctly on warm builds via persistent per-module effect summaries in the build cache.
  • Effect summary cache (schema v2.0) — Each module's public function declarations, internal names, and call-site listings are cached alongside the content hash, so incremental builds retain complete cross-module coverage without re-parsing skipped files.
  • Cross-Module Effect Propagation guide — Contract model, bare-name vs. qualified calls, incremental build semantics, CLI + MSBuild integration.

Changed

  • --input option in the calor CLI now accepts multiple values.
  • Build state cache format bumped from 1.0 to 2.0 — existing caches auto-invalidate on first build after upgrade.
  • Options hash includes EffectKind enum shape — future enum changes automatically invalidate caches, preventing stale summaries from silently dropping effects.

[0.4.8] - 2026-04-20

Added

  • Incremental compilation — CompileCalor MSBuild task now owns all incremental logic with a two-level cache gate: (mtime, size) stat check then SHA256 content hash. Global invalidation on compiler DLL, options, effect manifest, or output directory changes.
  • calor effects suggest CLI command — Analyzes Calor source files and generates a .calor-effects.suggested.json manifest template for unresolved external calls. Supports --json for agent consumption, --merge for additive updates to existing manifests.
  • Shared ExternalCallCollector — Extracted from InteropEffectCoverageCalculator, extended to walk class methods and constructors. Resolves variable types via §NEW initializer scanning.
  • Incremental build benchmark — Measures cold, warm (no changes), and warm (1 file changed) build times
  • Effect manifests .NET ecosystem guide — ~170 covered types, resolution mechanics, custom manifest authoring, CLI tools

[0.4.7] - 2026-04-20

Added

  • Static analysis for class members — The --analyze flag now examines methods, constructors, property accessors, operators, indexers, and event accessors (previously only top-level functions were analyzed)
  • Verification-gated reporting — --analyze only reports proven findings by default (Z3-confirmed or constant analysis); use --all-findings for lower-confidence results
  • Taint hop-count tracking — Taint analysis tracks propagation steps; single-hop parameter-to-sink flows filtered by default to reduce false positives
  • Bug pattern detection in class members — Division by zero, null dereference, integer overflow, off-by-one, path traversal, command injection, and SQL injection detection now covers all class member bodies
  • Arity-aware overload resolution — Correct overload resolved by argument count, preventing wrong return types from flowing into Z3
  • Constructor initializer binding — : base()/: this() arguments visible to bug pattern checkers
  • 33 new unit tests for class member binding, scope, overloads, dataflow, and end-to-end analysis
  • New --all-findings CLI flag for showing all analysis findings including inconclusive results
  • New static analysis documentation page

Fixed

  • False positive elimination — Unhandled expression types return opaque expressions instead of literal zero, eliminating false division-by-zero reports
  • this.field shadowing — this.field resolves from class scope, not method scope
  • Throw-to-catch CFG edges — Throw statements inside try blocks now flow to catch blocks
  • Assignment dataflow — x = 1 no longer reports x as used before write

Validated

  • 47 open-source projects scanned — 23 verified findings, ~90% true positive rate
  • Real findings: ILSpy null dereferences, FluentFTP path traversal, ASP.NET Core path traversal

[0.4.6] - 2026-04-18

Added

  • Effect system: .NET framework manifests — Tier B effect manifests for 30+ common .NET framework interfaces (ILogger, DbContext, HttpClient, ControllerBase, etc.)
  • Effect system: ecosystem library manifests — Manifests for Serilog, Newtonsoft.Json, Dapper, MediatR, AutoMapper, FluentValidation, Polly
  • Effect system: BCL manifest expansion — New manifests for System.Text.Json, Regex, Concurrent collections, Crypto types
  • Effect system: variable type resolution — Enforcement pass resolves instance method calls via initializer tracking
  • 95 new enforcement tests (210 total)

Fixed

  • Effect system: unified resolver — Consolidated three parallel effect systems into a single manifest-based resolver
  • Parser: compound effect codes — Fixed silent mis-parsing of chained compound codes

[0.4.5] - 2026-04-16

Fixed

  • Restrict expression-only match case to lambda-only patterns
  • PLIST nested call depth, call-as-pattern, OPTION compaction

Earlier Releases

See the full changelog on GitHub for versions prior to 0.4.5.