Content-addressed identities

Three identities travel inside a Contract and one is computed on demand. What each one means, and what moves it.

A .did file has many spellings for one meaning. You can rename a type, reorder record fields, reorder methods, reindent everything and add a page of comments without changing a single byte that a caller puts on the wire. Text comparison reports all of those as differences. It also misses real ones: two files that look almost identical can describe incompatible interfaces.

candid-core answers the question with hashes instead. It normalises a Contract — the validated, arena-based type graph it compiles a .did file into — and then hashes fixed projections of it. Three identities are stored in the documents it emits: two inside the Contract itself, and one inside the SourceInfo provenance sidecar that travels beside it. A fourth is computed on demand over a serialized document's exact octets. They deliberately answer different questions, and the useful skill is knowing which one to compare.

The shape of an identity string#

Every identity in this project is one string with the same three parts: a domain naming what kind of thing was hashed, the literal :sha256:, and 64 lowercase hexadecimal digits.

text
candid-core:contract:v1:sha256:d43274872cdb6c503456065d12c26b512ba9e3eac5b0a9533c8f9716293c6e18
candid-core:interface:v1:sha256:99f400820c0ca2fb7c32cc3ab1df6d6c000bc614cffe0e350b6a9a576da18d48
candid-core:source-bundle:v1:sha256:52cb9ba7ed7105a36de5a5f2665ef82261080406133498ed4aa8b3cac6f9bcca
candid-core:artifact:contract-json:v1:sha256:66c1371d29c896c2b292edc5dc1d344bf39103c5a1011141ed6883ace3e95401

The domain is not decoration. It is also the first thing in the hash preimage, which is utf8(domain), then a single 0x00 byte, then the payload bytes. That is what makes the identity spaces non-confusable: hashing one byte sequence under two domains gives two unrelated digests, so a source-bundle digest can never be mistaken for a contract digest even by accident. Validation enforces the spelling too — a Contract whose identities.contract is not exactly candid-core:contract:v1:sha256: plus 64 characters drawn from 0-9a-f is rejected with invalid_contract_id_format before any digest is compared, and the interface half has its own invalid_interface_id_format.

interface_id — the callable wire interface#

Contract::interface_id() returns Option<&str>. It hashes exactly four things: the semantics_profile, the canonicalization_profile, the part of the canonical type arena reachable from the actor, and the actor node itself.

What it leaves out is the point. format and format_version are excluded, so a document-shape revision does not by itself invalidate a compatibility cache. Declaration names are excluded, and so is every type node the actor cannot reach. Equality means one thing: the same callable actor interface under the same profiles. That is the right key for "can my client still call this canister?".

A Contract with no actor — a .did that declares only types — has no interface identity at all. The interface key is omitted from the document rather than written as null, and validation rejects an actorless Contract that claims one anyway. An empty actor, service : {}, is a real actor and does have one.

contract_id — the whole semantic Contract#

Contract::contract_id() returns &str; every valid Contract has one. It hashes format, format_version, both profiles, the complete canonical type arena, the declarations array including the names, and the actor when there is one.

Two things it excludes are worth stating out loud, because people assume otherwise. The identities block itself is excluded, since that is what is being derived. And producer — the tool name, tool version, and the candid/candid_parser versions recorded in the document — is excluded, deliberately. Producer metadata is unverified provenance; binding it would move every existing contract_id whenever an unrelated tool re-emitted the same Contract.

Also outside it: envelope extensions, the SourceInfo provenance sidecar, source text, comments, formatting, and packaging. Two files with different producers, different extensions, or different sidecars share one contract_id.

contract_id is the wrong thing to sign for a file

Because it excludes producer, extensions, the sidecar, and the encoding, a registry entry or signature keyed on contract_id commits to strictly less than the file the publisher believes they shipped. The repository's own vectors show it: nine checked-in documents spanning all three artifact kinds embed one identical contract_id and have nine distinct artifact IDs. When the octets are the claim, commit to an artifact ID.

source_bundle_id — the raw source files#

SourceInfo::source_bundle_id() is not a semantic identity, and it is the one people misread most often. It hashes a two-member payload: the canonical list of raw sources as {name, source} pairs, and their import edges as {from, import, to, kind} records. name is a normalised logical source URI such as memory:/root.did; source is the literal file text.

So reindenting a file or editing a comment does move it. That is the job. It answers "did the inputs change?", which is exactly the question a build cache asks. It is available only when the sidecar is (the compiler feature), since it describes what was compiled.

What it excludes is everything derived from those sources — the rederived documentation strings, the declaration and method provenance, the field-label table. Those live in the same SourceInfo but not in its bundle identity, so it identifies the input bundle rather than the whole sidecar. And because the logical source name is inside the payload, the same text compiled inline (compile_did names its source memory:/inline.did) and read from a workspace file are different bundles.

You can recompute one with nothing but a JSON canonicaliser and SHA-256. Against a checked-in fixture, this prints True:

bash
python3 - <<'EOF'
import json, hashlib
doc = json.load(open("tests/fixtures/artifact-identity/artifacts/compilation.json"))
info = doc["source_info"]
payload = {"sources": info["sources"], "imports": info["imports"]}
jcs = json.dumps(payload, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
digest = hashlib.sha256(b"candid-core:source-bundle:v1" + b"\x00" + jcs).hexdigest()
print("candid-core:source-bundle:v1:sha256:" + digest == info["source_bundle_id"])
EOF

artifact_id — the detached, exact-octet identity#

None of the three identifies a complete serialized document. That gap is what artifact_id_with_limits and artifact_id_with_context fill. They hash the exact octet sequence you hand them under a kind-specific domain, and hand the digest back.

Detached is the operative word. Nothing is stored. No serialized artifact gains an id field for it. No decode computes it implicitly, so no existing validation order or error precedence changes because it exists. If you want one, you ask for one.

artifact_id_with_limits
rust
pub fn artifact_id_with_limits(
    kind: ArtifactKind,
    bytes: &[u8],
    limits: &Limits,
) -> Result<String, ContractValidationError>

Hashes bytes under kind's domain. artifact_id_with_context is the same computation under a full RuntimeContext, so a caller's cancellation token and deadline are observed before the first byte and again at every 64 KiB chunk boundary — failing closed rather than returning a partial digest. There is deliberately no shorter convenience form.

ArtifactKind is #[non_exhaustive] and has three frozen variants. Selecting a domain is all it does: no variant parses or validates anything.

ArtifactKindDomainWhat a valid document of that kind carries
ContractJsonV1 candid-core:artifact:contract-json:v1 The Contract alone, producer included. No extensions, no sidecar.
ContractEnvelopeJsonV1 candid-core:artifact:contract-envelope-json:v1 A Contract plus its namespaced extension map. No sidecar.
CompilationJsonV1 candid-core:artifact:compilation-json:v1 A Contract plus its optional SourceInfo sidecar. No extensions.

This example is the rustdoc on the function, which runs as a doctest. It demonstrates the three properties at once: the domain is the literal prefix of the ID, the kind separates the digest space, and one space is a different artifact.

rustsrc/artifact_id.rs
use candid_core::{artifact_id_with_limits, ArtifactKind, Limits};

let document = br#"{"contract":{}}"#;
let id = artifact_id_with_limits(
    ArtifactKind::ContractEnvelopeJsonV1,
    document,
    &Limits::default(),
)?;
assert!(id.starts_with("candid-core:artifact:contract-envelope-json:v1:sha256:"));

// The same bytes under any other kind are a different identity.
for other in [ArtifactKind::ContractJsonV1, ArtifactKind::CompilationJsonV1] {
    let rehashed = artifact_id_with_limits(other, document, &Limits::default())?;
    assert_ne!(id, rehashed);
}

// One byte of whitespace is a different artifact.
let reformatted = artifact_id_with_limits(
    ArtifactKind::ContractEnvelopeJsonV1,
    br#"{"contract": {}}"#,
    &Limits::default(),
)?;
assert_ne!(id, reformatted);

Note what those bytes are: {"contract":{}} is not a valid Contract document at all. An artifact ID exists for it anyway, because computing one is not a validity claim. Equal IDs mean equal kind and equal octets, under the SHA-256 collision assumption — nothing more. It is a content address, not a credential: it establishes neither semantic equality, nor structural validity, nor authenticity, producer truth, or signature trust. This crate defines no signer model, key format, signature algorithm, trust policy, or registry protocol. Validate the artifact separately through the bounded parse entry point for its kind, then use the ID as the value a signature commits to.

Both entry points are bounded the same way. max_input_bytes is enforced against the slice before any hashing, reported as resource input_bytes; max_artifact_identity_work is then charged one unit per artifact byte plus the fixed domain framing, reported as artifact_identity_work. No other counter is consumed, so content-addressing a document can neither starve nor be starved by canonicalization on a shared budget. If you raise max_input_bytes above its 4 MiB default, raise the work limit too, or oversized artifacts start failing on the wrong counter.

What moves which id?#

This is the table to keep. Moves means a different digest; unchanged means byte-identical. Every cell was checked against src/canonical.rs, the two normative identity specifications, and the tests — the notes under the table say where.

Change interface_id contract_id source_bundle_id artifact ID
Add or edit a comment in the source a Unchanged Unchanged Moves Moves — compilation
contract, envelope: unchanged
Reformat whitespace in the source b Unchanged Unchanged Moves Moves — compilation
contract, envelope: unchanged
Rename a type declaration c Unchanged Moves Moves Moves — all three kinds
Reorder two top-level declarations d Unchanged Unchanged Moves Moves — compilation
contract, envelope: unchanged
Reorder fields inside a record e Unchanged Unchanged Moves Moves — compilation
contract, envelope: unchanged
Rename a record field f Moves if the field is actor-reachable Moves Moves Moves — all three kinds
Add a method to the service g Moves Moves Moves Moves — all three kinds
Change producer metadata h Unchanged Unchanged Unchanged Moves — all three kinds
Re-serialize the same Contract with different key order i Unchanged Unchanged Unchanged Moves — all three kinds

The artifact column assumes a valid serialized document of the named kind, because that is the only case where "does this change reach the bytes?" has a general answer. ArtifactKind itself never parses anything, so the honest statement for arbitrary input is shorter: if the octets differ, the ID differs.

Where each row is checked#

a Comment
identities_make_distinct_equality_claims in tests/adr_conformance.rs compiles // first and // second in front of the same service and asserts equal contract_id, different source_bundle_id. The compilation and compilation_source_doc vectors differ only in one doc comment and have different artifact IDs and bundle IDs with an identical contract_id.
b Whitespace
The bundle payload stores source as the literal file text (SourceFileInfo in src/model/source_info.rs, populated in src/compile/lower.rs), and no source text enters either semantic payload. The interface half is pinned directly by interface_ids_are_deterministic_and_ignore_provenance_but_track_wire_semantics in tests/contract_foundation.rs, whose two sources differ in formatting, field order, method order, and type alias.
c Type rename
contract_id_changes_when_declaration_names_change in tests/adr_conformance.rs. Declaration names are in the contract payload and not in the interface payload, and canonical node numbering traverses the actor first and then declaration roots in quotient class-ID order — never name order — so a rename cannot shift the actor-reachable prefix.
d Declaration order
No test literally swaps two type declarations in .did source text. What is pinned: the arena_permutation conformance case ships three raw inputs that arrange one graph's arena three different ways — listing its two declarations in both orders across them — and asserts all three converge on one pinned contract_id and interface_id, and the proptest arbitrary_arena_permutations_of_a_nasty_graph_converge in tests/canonical_properties.rs shuffles declaration listing order and asserts the canonical Contract and both IDs are unchanged. Read those two if you want to check this row yourself.
e Field order
equivalent_source_ordering_preserves_semantic_identity in tests/canonical_properties.rs compiles record { right: text; left: nat } and record { left: nat; right: text } behind the same service and asserts the whole compiled Contracts are equal, which covers both semantic IDs. Canonicalization sorts record and variant fields by label ID ascending.
f Field rename
A Field in the graph is { id: u32, ty: TypeRef } — the spelling is gone, replaced by the Candid label hash. So this row is precisely "moves whenever the new spelling hashes to a different label ID". note is 1225398258 and memo is 1213809850, so that rename moves them. The nearest direct test changes a numeric label from 1 to 2 and asserts a different interface_id (tests/contract_foundation.rs).
g New method
The interface payload is the actor plus its reachable arena — §9.2 of docs/canonicalization-v1.md — and a service node's canonical signature writes each method's ID and name, which is §5 of the same document. Both are implemented in src/canonical.rs. A related assertion in tests/contract_foundation.rs pins that changing one method's mode from query to update moves interface_id.
h Producer
rewritten_producer_bytes_move_the_artifact_id_and_not_the_semantic_ids in tests/artifact_identity.rs rebrands a producer to candid-core-fork/9.9.9 and asserts both semantic IDs are unchanged while the envelope artifact ID moves.
i Re-serialization
jcs_identity_is_independent_of_input_object_key_order in tests/adr_conformance.rs moves the actor key to the end of the document and asserts the decoded contract_id is unchanged. The contract_envelope and contract_envelope_compact vectors are the same envelope pretty-printed and minified: identical embedded contract_id, different artifact IDs.
Field names collide, and then nothing moves

The Candid label hash is 32-bit and collides. idl_hash("jhwlzguu") and idl_hash("jsyrjsvk") are both 2823597088. The collision itself is pinned as a unit test in src/name_hash.rs, which asserts the two hash equal; the literal 2823597088 is pinned by the hash_collision conformance vector. Renaming a field between two colliding spellings produces a byte-identical Contract, so neither semantic ID moves at all. Renaming a method between them still moves both, because a ServiceMethod keeps its name alongside its id — the text is what you need to make a call. Method names are semantic here; field names are not.

A worked example#

examples/semantic_equivalence.rs is the whole argument in one program. Two sources with a different type name, a different comment, a different field order and a different method order compile to the same wire interface and to different source bytes.

rustexamples/semantic_equivalence.rs
use candid_core::compile_did;
use std::error::Error;

fn main() -> Result<(), Box<dyn Error>> {
    let first = compile_did(
        r#"
        type Payload = record { owner: principal; amount: nat };
        service : {
          z: (Payload) -> () query;
          a: (Payload) -> () query;
        };
        "#,
    )?;
    let second = compile_did(
        r#"
        // Different name, documentation, field order, and method order.
        type Transfer = record { amount: nat; owner: principal };
        service : {
          a: (Transfer) -> () query;
          z: (Transfer) -> () query;
        };
        "#,
    )?;

    assert_eq!(
        first.contract().interface_id(),
        second.contract().interface_id()
    );
    assert_ne!(
        first.source_info().unwrap().source_bundle_id(),
        second.source_info().unwrap().source_bundle_id()
    );
    Ok(())
}
bash
cargo run --example semantic_equivalence

Read it against the table. Rows a, e and one method reordering all fire, so source_bundle_id differs and interface_id does not. Row c also fires — Payload versus Transfer — so these two Contracts have different contract_ids, which is why the example compares interface identities rather than contract identities.

The canonical document is where this becomes concrete. tests/fixtures/conformance/basic.did declares type Payload = record { owner: principal; amount: nat }; service : { transfer: (Payload) -> () };. Compiling it gives an arena whose node 0 is the actor's service — canonical numbering traverses the actor first — and whose record fields are ordered by label ID ascending rather than by source position. Here the two happen to coincide; a source that wrote amount first would produce this same document:

jsontests/fixtures/conformance/basic.contract.json
{
  "format": "candid-core",
  "format_version": 1,
  "semantics_profile": "candid-1",
  "canonicalization_profile": "candid-core-canon-1",
  "identities": {
    "contract": "candid-core:contract:v1:sha256:0b553cb65eb436a7ac5d35869d0016d043867bab0a97e844350ae72a5b4aea7d",
    "interface": "candid-core:interface:v1:sha256:99f400820c0ca2fb7c32cc3ab1df6d6c000bc614cffe0e350b6a9a576da18d48"
  },
  "types": [
    { "kind": "service", "methods": [
      { "name": "transfer", "id": 3664621355, "function": 1 }
    ] },
    { "kind": "func", "args": [2], "results": [], "mode": "update" },
    { "kind": "record", "fields": [
      { "id": 947296307, "type": 3 },
      { "id": 3573748184, "type": 4 }
    ] },
    { "kind": "primitive", "primitive": "principal" },
    { "kind": "primitive", "primitive": "nat" }
  ],
  "declarations": [
    { "name": "Payload", "type": 2 }
  ],
  "actor": { "kind": "service", "service": 0 }
}

The producer block is elided above. 3664621355 is the Candid hash of transfer, 947296307 of owner, 3573748184 of amount; all three reproduce under h = h * 223 + byte (mod 2³²). The field names are nowhere in the document.

Embedded identities are recomputed, not trusted#

The two semantic identities are written into the document, which raises the obvious question: what stops someone editing them? Nothing does — and nothing needs to, because decoding never reads them as fact. Contract::from_json recanonicalizes the incoming graph and compares. A mismatch fails with contract_id_mismatch at path $.identities.contract or interface_id_mismatch at $.identities.interface, and the message names both the expected and the found value.

rustexamples/trust_boundary.rs
let compilation = compile_did("service : { ping: () -> (nat) query };")?;
let canonical_json = compilation.contract().to_json_pretty()?;
let accepted = Contract::from_json(&canonical_json)?;
println!("validated {} type nodes", accepted.types().len());

let mut tampered: serde_json::Value = serde_json::from_str(&canonical_json)?;
tampered["identities"]["contract"] = serde_json::json!(
    "candid-core:contract:v1:sha256:0000000000000000000000000000000000000000000000000000000000000000"
);
let rejected = Contract::from_json(&serde_json::to_string(&tampered)?).unwrap_err();
println!("tampered semantic identity rejected: {rejected}");

A presented SourceInfo sidecar gets the same treatment, and more of it: its bundle must already be canonically sorted, its source_bundle_id must match a recomputation (source_bundle_id_mismatch), and the whole bundle is recompiled through the same pipeline so every derived provenance field has to match. See The trust boundary for why raw and validated types are separate types at all.

Pre-1.0: the digests themselves are not stable yet

The crate is 0.1.0-beta.3. Until 1.0, any release may change the public Rust API, the serialized Contract, compilation and envelope shapes, the canonical bytes, and therefore every identity computed over them. The versioning mechanism described on Canonical bytes is what makes such a change announceable; it is not a promise that today's digests survive. Pin an exact version — candid-core = "=0.1.0-beta.3" — and recompute identities after any upgrade rather than assuming stored ones still match.

Next#