Canonical bytes

Why a byte-exact encoding exists, and what it guarantees across languages, machines, and releases.

The identities on the previous page are SHA-256 digests. A digest is only as portable as the bytes it is taken over, and "the bytes our Rust serializer happens to emit" is not portable at all. JSON lets you order object keys freely, add whitespace anywhere, and spell the same number several ways; two programs that agree completely about meaning will still disagree about the digest.

So candid-core does not derive its identities from a serializer. It derives them from a written specification — docs/canonicalization-v1.md, the candid-core-canon-1 profile — which fixes every byte of every identity preimage in language-independent terms. The Rust crate is a reference implementation of that document, not the document itself. This page explains the guarantee and what it costs you to rely on; the specification has the byte-level detail.

Four version markers, four kinds of change#

Schema shape, Candid semantics, and byte-level normalisation change for different reasons and on different timelines. Coupling them to one version number forces every consumer to treat every change as the same kind of break, so every persisted Contract declares four markers instead of one.

json
{
  "format": "candid-core",
  "format_version": 1,
  "semantics_profile": "candid-1",
  "canonicalization_profile": "candid-core-canon-1"
}
MarkerWhat it pinsRejected with
format That this document is a candid-core Contract at all. unsupported_contract_format at $.format
format_version The JSON field set and tagged unions — the document's shape. unsupported_format_version at $.format_version
semantics_profile Which Candid type system is being interpreted. unsupported_semantics_profile at $.semantics_profile
canonicalization_profile Graph minimization, collection ordering, re-indexing, the canonical bytes, and the identity domains. unsupported_canonicalization_profile at $.canonicalization_profile

Unknown values fail closed everywhere: in the Rust validator, and in the TypeScript Contract loader, which reuses the same four codes. Two of the four — semantics_profile and canonicalization_profile — are also inside the identity preimages, so a profile change is visible in the digest rather than silently reinterpreting an old one.

candid-core-canon-1 is frozen. Any observable change to the bytes it defines requires a new profile name, never an edit to this one. That is the whole point of naming it: producer can record an implementation fix without touching a digest, while a real normalisation change has to announce itself.

Normalise the graph, then write the bytes#

Canonicalization runs in two stages, both before anything is hashed.

Minimize#

Candid type definitions are equi-recursive: an alias, or a second definition with identical structure, does not create a new wire type. So the arena is partitioned into classes of nodes that behave identically — including across cycles — and one node is kept per class. Two records spelled out separately in the source collapse into one node. The partition is computed by repeatedly refining classes from exact signature bytes, so its numbering never depends on the input's node ordering.

This is pinned by a vector rather than merely asserted. The duplicate_semantic_nodes conformance vector's raw input has four type nodes — two identical record nodes each pointing at its own nat node — and the pinned canonical output has two, with both declarations targeting node 0.

Re-index#

The surviving nodes are renumbered by a depth-first walk that starts at the actor and then visits declaration roots. Record and variant fields sort by label ID ascending; service methods by ID, then name, then function reference; declarations by name. All string comparisons are unsigned UTF-8 byte order, which for well-formed strings is Unicode scalar order — deliberately not UTF-16 code-unit order, so U+FF61 sorts before U+10000 here.

Two ordering facts are easy to get wrong when reimplementing. func argument and result lists and class init lists are positional: their order is semantic and is never sorted. And the traversal that assigns node indices visits declaration targets in ascending class-ID order, even though the output declarations array is sorted by name — the declaration_root_order vector exists to catch an implementation that used name order for traversal.

Because the walk starts at the actor, every actor-reachable node lands in a contiguous block at the front of the arena. That prefix is exactly the interface_id payload: never a node short, never a declaration-only node extra.

The constrained JSON writer#

The normalised payload is then rendered by a writer with no room for choice.

No whitespace
None anywhere. The separators are exactly {, }, [, ], : and ,.
Object keys sorted, compared as UTF-8 bytes
Every key in the payload vocabulary is one of 22 fixed ASCII schema names (actor, args, canonicalization_profile, class, declarations, fields, format, format_version, function, id, init, inner, kind, methods, mode, name, primitive, results, semantics_profile, service, type, types). Arrays are never sorted by the writer: array order is already semantic, decided by the re-indexing stage.
Numbers in shortest decimal form
The vocabulary contains only unsigned integers that fit in a u32. They are written with no sign, no leading zeros, no exponent and no fraction. Booleans, null, floats, negatives and integers at or above 2³² do not occur, and a conforming implementation must not emit them.
A fixed string escape set
\" and \\; the five short escapes \b \t \n \f \r; a six-character \u00-plus-two-lowercase-hex-digits form for the remaining control characters U+0000–U+001F; and literal UTF-8 bytes for everything else, including U+007F and every non-ASCII scalar. Printable characters are never \uXXXX-escaped.
Unicode hygiene, with no repairs
Input strings must already be well-formed scalar sequences. A writer that meets an unpaired surrogate must fail — it must not normalize, substitute U+FFFD, or drop anything. There is no Unicode normalization at any stage, so NFC é and NFD e + U+0301 stay distinct byte sequences and produce different digests.
This is a constrained JCS profile, not general RFC 8785

Within this vocabulary the output is byte-identical to RFC 8785, but two JCS behaviours are deliberately out of scope. RFC 8785 sorts property names by UTF-16 code units; this writer sorts by UTF-8 bytes — the orders differ only for a key mixing a supplementary-plane scalar with one in U+E000–U+FFFF, which cannot happen when every key is a fixed ASCII name. And RFC 8785 serializes numbers through IEEE-754 ES6 rules, which never comes up when the vocabulary is u32 only. A general JCS library that is byte-correct on this vocabulary may be substituted, but the conformance vectors are the arbiter, not the library's own claims.

The preimage and the digest#

Hashing is domain-separated. For a payload P with domain D:

text
preimage = utf8(D) ++ 00 ++ canonical-json-bytes(P)
digest   = SHA-256(preimage)
identity = D ++ ":sha256:" ++ lowercase-hex(digest)

One 0x00 byte separates the domain from the payload. There is no length field and no second kind label, because the domain already names the kind and the payload is one contiguous run. Uppercase hex spellings are invalid.

Recompute one yourself#

tests/fixtures/conformance/actorless.identity.json pins the canonical payload text, the same bytes as hex, the full preimage as hex, and the resulting ID for a one-line source file: type LibraryValue = record { value: nat; note: text };. Here is the canonical payload in full — one line, no whitespace, keys sorted:

jsontests/fixtures/conformance/actorless.identity.json
{"canonicalization_profile":"candid-core-canon-1","declarations":[{"name":"LibraryValue","type":0}],"format":"candid-core","format_version":1,"semantics_profile":"candid-1","types":[{"fields":[{"id":834174833,"type":1},{"id":1225398258,"type":2}],"kind":"record"},{"kind":"primitive","primitive":"nat"},{"kind":"primitive","primitive":"text"}]}

Read what the rules did. Keys are sorted, so canonicalization_profile comes first and types last, and inside each node fields precedes kind. There is no actor key at all — an actorless Contract omits the member entirely, and "actor": null never appears in a payload. The field names are gone, replaced by their Candid label IDs: 834174833 is value and 1225398258 is note, sorted ascending. And the two profile markers are inside the bytes being hashed.

This rebuilds the preimage and the digest from those bytes with nothing but the Python standard library — no Rust, no candid-core. Deriving the bytes themselves from a raw graph is the rest of the job, and the reference implementation described below does that for every scenario:

bash
python3 - <<'EOF'
import json, hashlib
pins = json.load(open("tests/fixtures/conformance/actorless.identity.json"))
preimage = pins["domain"].encode() + b"\x00" + pins["jcs"].encode()
assert preimage.hex() == pins["preimage_hex"]
recomputed = pins["domain"] + ":sha256:" + hashlib.sha256(preimage).hexdigest()
print(recomputed)
print("matches the pin:", recomputed == pins["contract_id"])
EOF
text
candid-core:contract:v1:sha256:d43274872cdb6c503456065d12c26b512ba9e3eac5b0a9533c8f9716293c6e18
matches the pin: True

The saved document is not the canonical bytes#

This trips people up, so it is worth stating flatly. The canonical bytes exist only as a hash preimage. They are never what you write to disk.

Contract::to_json_pretty() revalidates, recanonicalizes, and then pretty-prints the canonical Contract with ordinary indentation and the struct's own field order. That document is presentation. It embeds the identities, and those identities were computed over the constrained bytes, which are not in the file at all.

The consequence is the one you want: the same Contract can be handed to you minified, pretty-printed, or with its keys in any order, and it still decodes to the same contract_id — because decoding recanonicalizes and recomputes rather than reading the stored value as fact. The repository pins both halves. jcs_identity_is_independent_of_input_object_key_order moves the actor key to the end of a document and asserts the decoded contract_id is unchanged; the contract_envelope and contract_envelope_compact artifact vectors are the same envelope pretty-printed and minified, with one shared contract_id and two different artifact IDs.

Two implementations, one manifest#

A specification nobody has independently implemented is a hope, not a guarantee. So the format is checked by two implementations that share nothing but the fixtures.

tests/fixtures/conformance/manifest.json pins 11 required scenarios: actorless, empty actor, class actor, basic service, direct recursion, mutual recursion, an idl_hash collision, Unicode, duplicate semantic nodes, arena permutations, and declaration-root traversal order. Between them they carry 24 raw input graphs — deliberately noncanonical arrangements mixed with already-canonical ones, which is what pins idempotence — and, per scenario, the expected canonical graph, the canonical JSON as text and as hex, the domain preimage, and the contract identity, plus the interface identity for every scenario that has an actor. Scenarios with several raw inputs must all converge on one identical canonical result.

  • tests/conformance_vectors.rs drives the manifest from Rust, with supporting exact-fixture coverage in tests/adr_conformance.rs and ordering, normalization and permutation properties in tests/canonical_properties.rs.
  • tests/fixtures/conformance/verify_vectors.py is an independent canonicalizer written against the specification using the Python standard library only. It recomputes every canonical graph, payload byte, preimage and ID from the raw vectors without touching the Rust implementation. CI runs it as its own job, Independent conformance reference.

Detached artifact identity has a parallel setup, kept deliberately separate so the closed semantic conformance set keeps meaning exactly what it did: tests/fixtures/artifact-identity/manifest.json pins 10 vectors plus 6 cross-kind entries, verify_artifact_ids.py recomputes them all in Python, and CI runs that as Independent artifact identity reference. A third job hashes documents inside a real headless browser and checks the per-kind framing anchors there against the same literals the native tests pin, which is what would catch a digest that differed between native and WebAssembly builds.

An implementation conforms to candid-core-canon-1 if and only if it reproduces every pinned value of every manifest vector. That is the whole definition, and it is what makes the same digest reachable from a language nobody here has written yet. The Python reference is the existence proof: it was written against the specification, and it agrees with the Rust crate down to the byte.

Resource policy is not identity

Every entry point is bounded — see Limits, budgets and diagnostics — but limit values, work-accounting formulas, and which limit fires first are explicitly non-normative and must not influence a single canonical byte. Two conforming implementations with different budgets differ only in which inputs they refuse, never in the bytes or IDs of an input both accept.

What the guarantee is, and is not#

What it is: identity stability across machines, toolchains and languages is a property you can test, and the project tests it. Given a Contract graph, the digest does not depend on your platform, your JSON library, your locale, or whether you got there from a .did file or by building the graph directly. ADR 0002 records this decision and is the only one of the seven foundation ADRs currently marked Verified; the evidence it cites is the Python reference reproducing every canonical graph, payload byte, domain preimage and ID across the 11 required scenarios.

What it is not: a promise that today's digests are permanent. The crate is 0.1.0-beta.3, and until 1.0 any release may change the public API, the serialized shapes, the canonical bytes, and every identity computed over them. The profile mechanism is what makes such a change announceable and detectable rather than silent — it does not prevent it. Pin an exact version and recompute after an upgrade.

Next#