Guarantees and verification
What is actually proven, by which gate, and what is explicitly not claimed.
A Contract is this project's noun for the validated type graph projected out of a Candid .did interface — the type nodes, the declarations, the optional actor — together with a byte-exact canonical encoding of that graph and the SHA-256 identities computed over those bytes. Claims about a structure like that are worth exactly what the evidence behind them is worth, so this page keeps two things apart that are easy to blur: a gate that runs, and a property that has been shown to hold.
Everything named here is a file you can open. Where a claim is supported only by an implementation and not yet by independent evidence, this page says so rather than rounding up.
The battery every change runs#
No change reaches the default branch without passing all of this. It is written down in CLAUDE.md and split across two workflows, each of which runs on every pull request and every push to main. Verify carries the formatting, lint and test gates and the feature-graph verifier; Release candidate carries the packaging and semver verifiers, because those two ask their question of the published archive rather than of the repository tree.
cargo fmt --check && git diff --check
cargo clippy --all-targets --all-features --locked -- -D warnings
cargo test --all-targets --locked
cargo test --all-targets --locked --no-default-features
python3 tests/fixtures/packaging/verify_package_manifest.py --locked
python3 tests/fixtures/packaging/verify_feature_graph.py
python3 tests/fixtures/packaging/verify_semver.pyFormatting, lints and the feature matrix#
Clippy runs over all targets — library, binary, tests, benchmarks, examples — with -D warnings, so a warning is a failure rather than a note somebody scrolls past. Tests run twice in the battery above, but the Verify feature matrix runs all six supported feature combinations — no features, host-value, compiler, compiler,host-value, filesystem-compiler, and every feature — because a change that only compiles with default features is a change that breaks a real consumer. Clippy is run again over four of those six, and rustdoc over the reduced and full surfaces, so an intra-doc link to a feature-gated item cannot break the reduced build.
Suites with meaningful pure-model coverage build under --no-default-features and gate individual cases, so the conformance vectors and the canonicalization property tests run with no features at all — the configuration that links no Candid engine. Suites that need a feature throughout declare required-features and are skipped rather than silently emptied.
The packaging verifier and the positive allowlist#
The published archive is defined by a positive include list in Cargo.toml, not by an exclude list. The difference is which way a mistake falls: with exclude, a file added tomorrow ships until somebody remembers to keep it out; with include, a new path is outside the archive until it is named on purpose.
include = [
"/src/**/*.rs",
"/examples/**/*.rs",
"/docs/**/*.md",
"/README.md",
"/CHANGELOG.md",
"/LICENSE",
]verify_package_manifest.py asserts that policy in both directions, and each direction has its own failure mode. Nothing internal ships: tests/, benches/, fuzz/, .github/, agent assets and any scratch directory in a contributor's working tree must all be absent. Nothing required is missing: every src/**.rs, examples/**.rs and docs/**.md on disk must be present in the archive, with per-directory floors — at least 20 paths under src/, 6 under examples/, 10 under docs/ — so an include glob that silently matches nothing cannot pass. The same check reports strays: a file under a published directory that the globs do not match, such as a src/table.json behind an include_str! or a docs/diagram.svg an ADR links to, compiles and renders here and is absent for a consumer, so the script names it rather than letting the first installer find it. And in the other direction no unexplained path is admitted: every packaged path must be accounted for by a rule in the script, and the manifest metadata itself is compared against recorded values, so relaxing the allowlist is a deliberate edit to the script in the same change.
The archive holds 57 paths today. The script prints that count; what it asserts is the set membership in both directions that produces it, plus the floors. The number is a recorded expectation in Cargo.toml and the contributor guide, not a hard-coded assertion in the verifier — worth knowing if you are relying on it.
The feature-graph verifier, and why cargo metadata is the wrong tool#
The feature split is a claim about dependency graphs: a consumer who only reads Contracts must not be made to build a Candid source engine, and a consumer who never touches a filesystem must not be made to build a filesystem capability crate. verify_feature_graph.py checks that claim by resolving cargo tree --edges no-dev per feature set and per target triple, and asserting that the base graph excludes candid, candid_parser, cap-std and ic_principal; that host-value adds only ic_principal; that compiler adds the parser stack but no cap-std; and that cap-std appears only for filesystem-compiler on targets that have a filesystem.
The script used to walk normal and build edges over a cargo metadata resolve. That was sound while the package was alone in its workspace and stopped being sound the moment it gained a sibling: the generator crate's golden tests dev-depend on candid-core with compiler on, and a whole-workspace resolve unifies that feature into candid-core before any edge filtering can help. --edges no-dev excludes dev-dependencies from resolution itself, not from a walk over an already-resolved graph.
The two verifiers answer different questions: feature selection bounds what a consumer must build, while the include allowlist bounds what a consumer must download. So default-features = false does not shrink the archive, and it cannot subtract features from a build where something else already asked for more.
The semver gate — and what it actually proves#
verify_semver.py runs cargo semver-checks against the published baseline on both published surfaces. Read its claim precisely: it is evidence that no breaking change reached the published surfaces unnoticed — not that none occurred. The crate is pre-1.0 and reserves the right to change the public API, the serialized shapes, the canonical bytes and every identity computed over them, so a gate that hard-failed on a break would fight intended work, and a warning nobody must act on would be ignored. A reported break must instead be acknowledged in the ## Unreleased section of CHANGELOG.md by a list item beginning with a bolded BREAKING marker.
Two details in that gate are load-bearing. The marker has to start a list item, because the changelog entry documenting the gate necessarily contains the word while explaining it — a substring match would be satisfied by its own documentation and would then pass every break forever. And the script forces --release-type patch: without it, a branch whose version has not been bumped compares X -> X, which cargo semver-checks reads as "assume major" and answers by running zero checks. A gate that passes because it asked nothing is worse than no gate.
Cross-implementation evidence#
This is the part that is actually interesting. A test written against the same implementation it tests can only confirm that the implementation agrees with itself. Several claims here are instead checked by a second implementation that shares no code with the first.
Two compilation backends, compared against each other#
There are two ways to get from sources to a checked result: an in-memory backend that merges the bundle into one virtual program through the candid_parser merged-program APIs, and a native backend that materialises files for the upstream file checker used by compile_did_file. Rather than trusting that they agree, src/compile/differential.rs runs the same bundles through both and requires byte-identical canonical Contracts, identities and provenance for valid input, and identical stable diagnostic codes, phases and resource triples for invalid input — across type imports, service imports, diamonds, recursion, actorless and class actors, and merge, type-check, duplicate-binding and parse failures.
Two independent Python references#
The canonical byte format is specified in docs/canonicalization-v1.md; the Rust code is a reference implementation, not the specification. A standard-library Python canonicaliser recomputes every conformance vector's canonical graph, payload bytes, domain preimage and identities from the raw, non-canonical inputs — without calling Rust — and CI runs it as its own job.
python3 tests/fixtures/conformance/verify_vectors.py
python3 tests/fixtures/artifact-identity/verify_artifact_ids.pyThe second reference is deliberately separate, so the closed semantic conformance set keeps meaning exactly what it did. It recomputes every vector's artifact_id from the exact bytes on disk, pins the whole domain-framing preimage as hex for the empty vector, and asserts that the nine documents embedding one shared contract_id have nine distinct artifact identities — which is the whole reason a detached exact-octet identity exists. Those fixtures are marked text eol=lf in .gitattributes, because an exact-octet identity cannot survive a checkout that rewrites line endings.
The TypeScript codec, in both directions, against the reference Candid crate#
The upstream candid Rust crate is the reference wire-format implementation, and crates/candid-core-ts/tests/wire_vectors.rs holds both halves of the differential. Emit: every vector is encoded with candid into a golden of hex bytes that the TypeScript decoder must accept and interpret to the same domain value. Verify back: the TypeScript encoder's own bytes are checked in and decoded here with candid, and must equal the case's value — so the reference validates this project's encoder, not only the reverse. The type environment comes from candid_parser over the fixture .did source directly, deliberately not through candid-core's own compiler, so the reference path shares no code with the model under test.
The wasm CLI against the native binary, and the browser on a pinned Chrome#
The WebAssembly build wraps the same Rust code the native binary uses, and the repository asserts that rather than assuming it: the generated module for every fixture must be byte-identical to the reviewed golden, and the envelope output must be byte-identical to a committed fixture the native candid-core compile --envelope binary actually produced — the entire document, producer included. The same comparisons run again through the real compiled WebAssembly under Node, and two clean builds must produce identical bytes.
The browser suite is a runtime claim, not a build claim. It compiles a self-contained source and a four-source imported bundle inside headless Chrome, pins the resulting contract_id, interface_id, source_bundle_id, exact logical sources and exact import edges, and asserts that an unresolvable import, an exhausted limit, cancellation and an explicit deadline all fail with stable codes on a target with no filesystem and no clock. Three things are pinned exactly — wasm-pack at 0.14.0, the browser at Chrome for Testing 150.0.7871.124, and the ChromeDriver taken from that same build — so no rolling channel is involved, and the job prints both versions before running. Every case body is shared with the native jobs, which is what stops the pinned identities drifting between the two.
The tsc equality gate over the generated goldens#
The generator emits, for each declaration, a line of the form const $X: $.Schema<$X> = $.c.rec(() => …), exported under the Candid name X. Because Schema<in out T> is invariant, that line compiles only if the type the builder infers is equal to the reviewed alias in both directions. Running tsc --noEmit over the checked-in goldens with the exact-pinned TypeScript is therefore a proof about the mapping, not a lint. Alongside it, the schema runtime's own suite includes a cross-check that schemas built dynamically from the golden Contract JSON documents validate exactly the values the generated builders describe — same verdicts, same issue codes, same paths, same messages, asserted by deep equality on the whole result.
Fuzzing: candid-core-fuzz#
The fuzz harness is publish = false and declares its own [workspace] table with its own tracked fuzz/Cargo.lock, so libfuzzer-sys never enters the published crate's resolved graph or Cargo's feature unification. It mirrors the library's feature surface, and each target declares the feature that owns the API it drives. There are seven targets:
| Target | Entry point it drives | Feature |
|---|---|---|
source_parsing | compile_did | compiler |
contract_json | Contract::from_json_with_limits | base |
canonicalization | RawContract decode then ContractDraft build | base |
resolver_ids | SourceId::parse | compiler |
provenance | Compilation::from_slice_with_limits | compiler |
host_value | HostValue::from_json_with_limits | host-value |
envelope_json | ContractEnvelope::from_slice_with_limits | base |
Pull requests compile every target and replay its tracked seed and regression corpora with -runs=0, so a target that stops compiling, or a previously fixed crash that returns, fails on the pull request rather than on the next weekly run. The replay performs no mutation and is therefore deterministic. Both fuzz jobs first assert that fuzz/Cargo.lock is current, because cargo fuzz accepts no --locked flag of its own. A weekly job does the actual mutation and uploads crash artifacts.
Benchmarks measure; they never gate#
Benchmarks exist and are reproducible, and the comparison tool refuses rather than guesses: it records both halves of a run's identity — corpus fingerprints, feature set and metric units from the Rust side; toolchain, target, host, codegen flags and the bench binary's resolved dependency graph from its own side — and exits non-zero with "no comparison was made" when they disagree. Every report leads with control benchmarks that execute only upstream parser code, because a shift the controls share with the rest of the table measures the machinery rather than the change.
No timing or allocation measurement ever fails a workflow in this repository. That is a recorded decision, not an unfinished feature: hosted runners are shared and thermally variable, and a threshold calibrated from their noise produces a red check maintainers learn to ignore. Detection without enforcement — the numbers are surfaced with absolute values, intervals and raw data, and a human decides.
That policy is also why no page on this site states a performance number. A measurement no workflow defends is a measurement that can drift without anyone noticing, and publishing one as a property would be exactly the blurring this page exists to avoid. Treat any single benchmark result as a reason to investigate, never as proof of a regression.
The honest ledger#
Seven architecture decision records define the protocol boundaries that must hold before the Contract format could be called stable. Readiness is tracked per decision, and the two states mean different things. Implemented, verification pending means a working reference implementation exists and its required-verification list is written down, while at least one gate on that list has no recorded evidence. Verified means those specific gates completed and the evidence — pull request, commit, CI run — is recorded.
As the files stand today:
- ADR 0002 (version schema, semantics and canonical bytes independently) — Verified. The independent Python reference reproduced every canonical graph, payload byte, domain preimage, Contract ID and interface ID across the 11 required scenarios in PR #73, and the dedicated
conformance-referencejob passed in a named CI run recorded indocs/verification.md. - ADRs 0001 and 0003–0007 — Implemented, verification pending, all six.
ADR 0002's result promotes nothing else. ADR 0007 in particular has an independent Python reference and a dedicated CI job, and its status is still pending, because the repository records no run of that job. Wiring a CI job is not evidence. A workflow file that exists, a script that is committed, and a gate that would fire are three descriptions of intent; the evidence is a recorded run against a named commit, and until there is one, the honest status is pending.
The Contract format is not a stable v1, and the beta does not promote it to one. Any pre-1.0 release may change the public API, the serialized shapes, the canonical bytes, and therefore every identity computed over them — so pin an exact version, and treat a stored contract_id as stable within a version rather than across versions. If a property matters enough to depend on, the vectors, the references and the fixtures are all in the repository under an allowlist that keeps them out of the published archive: clone it and run them yourself.
Next: Design decisions covers what each ADR actually decided and the consequence you can see in the API; Status, versions and releases covers what is published, what changed in each version, and the pre-1.0 policy in full. Content-addressed identities explains the four identities the evidence above is mostly about.