Content-addressed identities
Three identities travel inside a Contract and one is computed on demand. What each one means, and what moves it.
A .did file has many spellings for one meaning. You can rename a
type, reorder record fields, reorder methods, reindent everything and add a page
of comments without changing a single byte that a caller puts on the wire. Text
comparison reports all of those as differences. It also misses real ones: two
files that look almost identical can describe incompatible interfaces.
candid-core answers the question with hashes instead. It normalises a
Contract — the validated, arena-based type
graph it compiles a .did file into — and then hashes fixed
projections of it. Three identities are stored in the documents it emits: two
inside the Contract itself, and one inside the SourceInfo
provenance sidecar that travels beside it. A fourth is computed on demand over a
serialized document's exact octets. They deliberately answer different
questions, and the useful skill is knowing which one to compare.
The shape of an identity string#
Every identity in this project is one string with the same three parts: a
domain naming what kind of thing was hashed, the literal
:sha256:, and 64 lowercase hexadecimal digits.
candid-core:contract:v1:sha256:d43274872cdb6c503456065d12c26b512ba9e3eac5b0a9533c8f9716293c6e18
candid-core:interface:v1:sha256:99f400820c0ca2fb7c32cc3ab1df6d6c000bc614cffe0e350b6a9a576da18d48
candid-core:source-bundle:v1:sha256:52cb9ba7ed7105a36de5a5f2665ef82261080406133498ed4aa8b3cac6f9bcca
candid-core:artifact:contract-json:v1:sha256:66c1371d29c896c2b292edc5dc1d344bf39103c5a1011141ed6883ace3e95401The domain is not decoration. It is also the first thing in the hash
preimage, which is utf8(domain), then a single 0x00
byte, then the payload bytes. That is what makes the identity spaces
non-confusable: hashing one byte sequence under two domains gives two unrelated
digests, so a source-bundle digest can never be mistaken for a contract digest
even by accident. Validation enforces the spelling too — a Contract whose
identities.contract is not exactly
candid-core:contract:v1:sha256: plus 64 characters drawn from
0-9a-f is rejected with invalid_contract_id_format
before any digest is compared, and the interface half has its own
invalid_interface_id_format.
interface_id — the callable wire interface#
Contract::interface_id() returns
Option<&str>. It hashes exactly four things: the
semantics_profile, the canonicalization_profile, the
part of the canonical type arena reachable from the actor, and the actor node
itself.
What it leaves out is the point. format and
format_version are excluded, so a document-shape revision does not
by itself invalidate a compatibility cache. Declaration names are excluded, and
so is every type node the actor cannot reach. Equality means one thing: the same
callable actor interface under the same profiles. That is the right key for
"can my client still call this canister?".
A Contract with no actor — a .did that declares only types — has
no interface identity at all. The interface key is omitted from the
document rather than written as null, and validation rejects an
actorless Contract that claims one anyway. An empty actor,
service : {}, is a real actor and does have one.
contract_id — the whole semantic Contract#
Contract::contract_id() returns &str; every
valid Contract has one. It hashes format,
format_version, both profiles, the complete canonical type arena,
the declarations array including the names, and the actor when there is
one.
Two things it excludes are worth stating out loud, because people assume
otherwise. The identities block itself is excluded, since that is
what is being derived. And producer — the tool name, tool version,
and the candid/candid_parser versions recorded in the
document — is excluded, deliberately. Producer metadata is unverified
provenance; binding it would move every existing contract_id
whenever an unrelated tool re-emitted the same Contract.
Also outside it: envelope extensions, the SourceInfo provenance
sidecar, source text, comments, formatting, and packaging. Two files with
different producers, different extensions, or different sidecars share one
contract_id.
Because it excludes producer, extensions, the sidecar, and the encoding, a
registry entry or signature keyed on contract_id commits to
strictly less than the file the publisher believes they shipped. The
repository's own vectors show it: nine checked-in documents spanning all three
artifact kinds embed one identical contract_id and have nine
distinct artifact IDs. When the octets are the claim, commit to an artifact
ID.
source_bundle_id — the raw source files#
SourceInfo::source_bundle_id() is not a semantic identity, and
it is the one people misread most often. It hashes a two-member payload: the
canonical list of raw sources as {name, source} pairs, and their
import edges as {from, import, to, kind} records. name
is a normalised logical source URI such as memory:/root.did;
source is the literal file text.
So reindenting a file or editing a comment does move it. That is the job. It answers "did the inputs change?", which is exactly the question a build cache asks. It is available only when the sidecar is (the compiler feature), since it describes what was compiled.
What it excludes is everything derived from those sources — the
rederived documentation strings, the declaration and method provenance, the
field-label table. Those live in the same SourceInfo but not in its
bundle identity, so it identifies the input bundle rather than the whole
sidecar. And because the logical source name is inside the payload, the same
text compiled inline (compile_did names its source
memory:/inline.did) and read from a workspace file are different
bundles.
You can recompute one with nothing but a JSON canonicaliser and SHA-256.
Against a checked-in fixture, this prints True:
python3 - <<'EOF'
import json, hashlib
doc = json.load(open("tests/fixtures/artifact-identity/artifacts/compilation.json"))
info = doc["source_info"]
payload = {"sources": info["sources"], "imports": info["imports"]}
jcs = json.dumps(payload, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
digest = hashlib.sha256(b"candid-core:source-bundle:v1" + b"\x00" + jcs).hexdigest()
print("candid-core:source-bundle:v1:sha256:" + digest == info["source_bundle_id"])
EOFartifact_id — the detached, exact-octet identity#
None of the three identifies a complete serialized document. That gap is what
artifact_id_with_limits and artifact_id_with_context
fill. They hash the exact octet sequence you hand them under a kind-specific
domain, and hand the digest back.
Detached is the operative word. Nothing is stored. No
serialized artifact gains an id field for it. No decode computes it
implicitly, so no existing validation order or error precedence changes because
it exists. If you want one, you ask for one.
pub fn artifact_id_with_limits(
kind: ArtifactKind,
bytes: &[u8],
limits: &Limits,
) -> Result<String, ContractValidationError>Hashes bytes under kind's domain.
artifact_id_with_context is the same computation under a full
RuntimeContext, so a caller's cancellation token and deadline
are observed before the first byte and again at every 64 KiB chunk
boundary — failing closed rather than returning a partial digest. There is
deliberately no shorter convenience form.
ArtifactKind is #[non_exhaustive] and has three
frozen variants. Selecting a domain is all it does: no variant parses
or validates anything.
ArtifactKind | Domain | What a valid document of that kind carries |
|---|---|---|
ContractJsonV1 |
candid-core:artifact:contract-json:v1 |
The Contract alone, producer included. No extensions, no sidecar. |
ContractEnvelopeJsonV1 |
candid-core:artifact:contract-envelope-json:v1 |
A Contract plus its namespaced extension map. No sidecar. |
CompilationJsonV1 |
candid-core:artifact:compilation-json:v1 |
A Contract plus its optional SourceInfo sidecar. No extensions. |
This example is the rustdoc on the function, which runs as a doctest. It demonstrates the three properties at once: the domain is the literal prefix of the ID, the kind separates the digest space, and one space is a different artifact.
use candid_core::{artifact_id_with_limits, ArtifactKind, Limits};
let document = br#"{"contract":{}}"#;
let id = artifact_id_with_limits(
ArtifactKind::ContractEnvelopeJsonV1,
document,
&Limits::default(),
)?;
assert!(id.starts_with("candid-core:artifact:contract-envelope-json:v1:sha256:"));
// The same bytes under any other kind are a different identity.
for other in [ArtifactKind::ContractJsonV1, ArtifactKind::CompilationJsonV1] {
let rehashed = artifact_id_with_limits(other, document, &Limits::default())?;
assert_ne!(id, rehashed);
}
// One byte of whitespace is a different artifact.
let reformatted = artifact_id_with_limits(
ArtifactKind::ContractEnvelopeJsonV1,
br#"{"contract": {}}"#,
&Limits::default(),
)?;
assert_ne!(id, reformatted);Note what those bytes are: {"contract":{}} is not a valid
Contract document at all. An artifact ID exists for it anyway, because computing
one is not a validity claim. Equal IDs mean equal kind and equal octets, under
the SHA-256 collision assumption — nothing more. It is a content address, not a
credential: it establishes neither semantic equality, nor structural validity,
nor authenticity, producer truth, or signature trust. This crate defines no
signer model, key format, signature algorithm, trust policy, or registry
protocol. Validate the artifact separately through the bounded parse entry point
for its kind, then use the ID as the value a signature commits to.
Both entry points are bounded the same way. max_input_bytes is
enforced against the slice before any hashing, reported as resource
input_bytes; max_artifact_identity_work is then charged
one unit per artifact byte plus the fixed domain framing, reported as
artifact_identity_work. No other counter is consumed, so
content-addressing a document can neither starve nor be starved by
canonicalization on a shared budget. If you raise max_input_bytes
above its 4 MiB default, raise the work limit too, or oversized artifacts start
failing on the wrong counter.
What moves which id?#
This is the table to keep. Moves means a different digest;
unchanged means byte-identical. Every cell was checked against
src/canonical.rs, the two normative identity specifications, and
the tests — the notes under the table say where.
| Change | interface_id |
contract_id |
source_bundle_id |
artifact ID |
|---|---|---|---|---|
| Add or edit a comment in the source a | Unchanged | Unchanged | Moves | Moves — compilation contract, envelope: unchanged |
| Reformat whitespace in the source b | Unchanged | Unchanged | Moves | Moves — compilation contract, envelope: unchanged |
| Rename a type declaration c | Unchanged | Moves | Moves | Moves — all three kinds |
| Reorder two top-level declarations d | Unchanged | Unchanged | Moves | Moves — compilation contract, envelope: unchanged |
| Reorder fields inside a record e | Unchanged | Unchanged | Moves | Moves — compilation contract, envelope: unchanged |
| Rename a record field f | Moves if the field is actor-reachable | Moves | Moves | Moves — all three kinds |
| Add a method to the service g | Moves | Moves | Moves | Moves — all three kinds |
Change producer metadata h |
Unchanged | Unchanged | Unchanged | Moves — all three kinds |
| Re-serialize the same Contract with different key order i | Unchanged | Unchanged | Unchanged | Moves — all three kinds |
The artifact column assumes a valid serialized document of the named kind,
because that is the only case where "does this change reach the bytes?" has a
general answer. ArtifactKind itself never parses anything, so the
honest statement for arbitrary input is shorter: if the octets differ, the ID
differs.
Where each row is checked#
- a Comment
identities_make_distinct_equality_claimsintests/adr_conformance.rscompiles// firstand// secondin front of the same service and asserts equalcontract_id, differentsource_bundle_id. Thecompilationandcompilation_source_docvectors differ only in one doc comment and have different artifact IDs and bundle IDs with an identicalcontract_id.- b Whitespace
- The bundle payload stores
sourceas the literal file text (SourceFileInfoinsrc/model/source_info.rs, populated insrc/compile/lower.rs), and no source text enters either semantic payload. The interface half is pinned directly byinterface_ids_are_deterministic_and_ignore_provenance_but_track_wire_semanticsintests/contract_foundation.rs, whose two sources differ in formatting, field order, method order, and type alias. - c Type rename
contract_id_changes_when_declaration_names_changeintests/adr_conformance.rs. Declaration names are in the contract payload and not in the interface payload, and canonical node numbering traverses the actor first and then declaration roots in quotient class-ID order — never name order — so a rename cannot shift the actor-reachable prefix.- d Declaration order
- No test literally swaps two
typedeclarations in.didsource text. What is pinned: thearena_permutationconformance case ships three raw inputs that arrange one graph's arena three different ways — listing its two declarations in both orders across them — and asserts all three converge on one pinnedcontract_idandinterface_id, and the proptestarbitrary_arena_permutations_of_a_nasty_graph_convergeintests/canonical_properties.rsshuffles declaration listing order and asserts the canonical Contract and both IDs are unchanged. Read those two if you want to check this row yourself. - e Field order
equivalent_source_ordering_preserves_semantic_identityintests/canonical_properties.rscompilesrecord { right: text; left: nat }andrecord { left: nat; right: text }behind the same service and asserts the whole compiled Contracts are equal, which covers both semantic IDs. Canonicalization sorts record and variant fields by label ID ascending.- f Field rename
- A
Fieldin the graph is{ id: u32, ty: TypeRef }— the spelling is gone, replaced by the Candid label hash. So this row is precisely "moves whenever the new spelling hashes to a different label ID".noteis 1225398258 andmemois 1213809850, so that rename moves them. The nearest direct test changes a numeric label from1to2and asserts a differentinterface_id(tests/contract_foundation.rs). - g New method
- The interface payload is the actor plus its reachable arena — §9.2 of
docs/canonicalization-v1.md
— and a
servicenode's canonical signature writes each method's ID and name, which is §5 of the same document. Both are implemented insrc/canonical.rs. A related assertion intests/contract_foundation.rspins that changing one method's mode fromqueryto update movesinterface_id. - h Producer
rewritten_producer_bytes_move_the_artifact_id_and_not_the_semantic_idsintests/artifact_identity.rsrebrands a producer tocandid-core-fork/9.9.9and asserts both semantic IDs are unchanged while the envelope artifact ID moves.- i Re-serialization
jcs_identity_is_independent_of_input_object_key_orderintests/adr_conformance.rsmoves theactorkey to the end of the document and asserts the decodedcontract_idis unchanged. Thecontract_envelopeandcontract_envelope_compactvectors are the same envelope pretty-printed and minified: identical embeddedcontract_id, different artifact IDs.
The Candid label hash is 32-bit and collides.
idl_hash("jhwlzguu") and idl_hash("jsyrjsvk") are
both 2823597088. The collision itself is pinned as a unit test in
src/name_hash.rs, which asserts the two hash equal; the literal
2823597088 is pinned by the hash_collision conformance vector.
Renaming a field between two colliding spellings produces a
byte-identical Contract, so neither semantic ID moves at all. Renaming a
method between them still moves both, because a
ServiceMethod keeps its name alongside its
id — the text is what you need to make a call. Method names are
semantic here; field names are not.
A worked example#
examples/semantic_equivalence.rs is the whole argument in one
program. Two sources with a different type name, a different comment, a
different field order and a different method order compile to the same wire
interface and to different source bytes.
use candid_core::compile_did;
use std::error::Error;
fn main() -> Result<(), Box<dyn Error>> {
let first = compile_did(
r#"
type Payload = record { owner: principal; amount: nat };
service : {
z: (Payload) -> () query;
a: (Payload) -> () query;
};
"#,
)?;
let second = compile_did(
r#"
// Different name, documentation, field order, and method order.
type Transfer = record { amount: nat; owner: principal };
service : {
a: (Transfer) -> () query;
z: (Transfer) -> () query;
};
"#,
)?;
assert_eq!(
first.contract().interface_id(),
second.contract().interface_id()
);
assert_ne!(
first.source_info().unwrap().source_bundle_id(),
second.source_info().unwrap().source_bundle_id()
);
Ok(())
}cargo run --example semantic_equivalenceRead it against the table. Rows a, e and one method
reordering all fire, so source_bundle_id differs and
interface_id does not. Row c also fires —
Payload versus Transfer — so these two Contracts have
different contract_ids, which is why the example compares interface
identities rather than contract identities.
The canonical document is where this becomes concrete.
tests/fixtures/conformance/basic.did declares
type Payload = record { owner: principal; amount: nat }; service : {
transfer: (Payload) -> () };. Compiling it gives an arena whose node 0
is the actor's service — canonical numbering traverses the actor first — and
whose record fields are ordered by label ID ascending rather than by source
position. Here the two happen to coincide; a source that wrote
amount first would produce this same document:
{
"format": "candid-core",
"format_version": 1,
"semantics_profile": "candid-1",
"canonicalization_profile": "candid-core-canon-1",
"identities": {
"contract": "candid-core:contract:v1:sha256:0b553cb65eb436a7ac5d35869d0016d043867bab0a97e844350ae72a5b4aea7d",
"interface": "candid-core:interface:v1:sha256:99f400820c0ca2fb7c32cc3ab1df6d6c000bc614cffe0e350b6a9a576da18d48"
},
"types": [
{ "kind": "service", "methods": [
{ "name": "transfer", "id": 3664621355, "function": 1 }
] },
{ "kind": "func", "args": [2], "results": [], "mode": "update" },
{ "kind": "record", "fields": [
{ "id": 947296307, "type": 3 },
{ "id": 3573748184, "type": 4 }
] },
{ "kind": "primitive", "primitive": "principal" },
{ "kind": "primitive", "primitive": "nat" }
],
"declarations": [
{ "name": "Payload", "type": 2 }
],
"actor": { "kind": "service", "service": 0 }
}The producer block is elided above. 3664621355 is the Candid
hash of transfer, 947296307 of owner, 3573748184 of
amount; all three reproduce under
h = h * 223 + byte (mod 2³²). The field names are nowhere
in the document.
Embedded identities are recomputed, not trusted#
The two semantic identities are written into the document, which raises the
obvious question: what stops someone editing them? Nothing does — and nothing
needs to, because decoding never reads them as fact. Contract::from_json
recanonicalizes the incoming graph and compares. A mismatch fails with
contract_id_mismatch at path $.identities.contract or
interface_id_mismatch at $.identities.interface, and
the message names both the expected and the found value.
let compilation = compile_did("service : { ping: () -> (nat) query };")?;
let canonical_json = compilation.contract().to_json_pretty()?;
let accepted = Contract::from_json(&canonical_json)?;
println!("validated {} type nodes", accepted.types().len());
let mut tampered: serde_json::Value = serde_json::from_str(&canonical_json)?;
tampered["identities"]["contract"] = serde_json::json!(
"candid-core:contract:v1:sha256:0000000000000000000000000000000000000000000000000000000000000000"
);
let rejected = Contract::from_json(&serde_json::to_string(&tampered)?).unwrap_err();
println!("tampered semantic identity rejected: {rejected}");A presented SourceInfo sidecar gets the same treatment, and more
of it: its bundle must already be canonically sorted, its
source_bundle_id must match a recomputation
(source_bundle_id_mismatch), and the whole bundle is recompiled
through the same pipeline so every derived provenance field has to match. See
The trust boundary for why raw and validated
types are separate types at all.
The crate is 0.1.0-beta.3. Until 1.0, any release may change
the public Rust API, the serialized Contract, compilation and envelope shapes,
the canonical bytes, and therefore every identity computed over them. The
versioning mechanism described on Canonical
bytes is what makes such a change announceable; it is not a
promise that today's digests survive. Pin an exact version —
candid-core = "=0.1.0-beta.3" — and recompute identities after
any upgrade rather than assuming stored ones still match.
Next#
The written-down byte encoding these digests are taken over, and why an independent implementation can reproduce them.
Concept Sources, imports and provenanceHow a multi-file bundle is resolved, and what the SourceInfo sidecar carries besides its bundle identity.
The arena, declarations and actor that canonicalization normalises before any of this is hashed.