Sources, imports and provenance
How multi-file DID bundles are resolved hermetically, and how the result is re-derivable rather than trusted.
Real Candid interfaces are rarely one file. A .did can import
another, that one can import a third, and two of them can import the same fourth. Something
has to decide what each import string points at and produce the bytes. In most tools that
something is ambient: the compiler opens whatever the path resolves to on the machine it
happens to be running on.
candid-core does not do that. Import resolution is a capability you hand in: an object with two methods that turns an import spelling into a canonical identifier and returns bytes for it. The compiler never reaches for a file or a socket on its own, so whatever authority the resolver has is the only authority in play. The result is that the same bundle compiles identically in a CLI, in CI, and inside a browser — and that the exact sources it used can be recorded, shipped, and later re-derived rather than believed.
Logical source IDs#
Every source in a compile is named by a SourceId: a normalised logical URI of
the form <scheme>:/<path>, such as memory:/api/types.did
or workspace:/root.did. It is deliberately not a host path. Two hosts with
different directory layouts that supply the same logical bundle get the same identities.
pub struct SourceId(String);
impl SourceId {
pub fn parse(value: impl AsRef<str>) -> Result<Self, ResolveError>;
pub fn as_str(&self) -> &str;
pub fn scheme(&self) -> &str;
pub fn path(&self) -> &str;
}
The path grammar is UTF-8 with / separators, and it is strict. Empty segments, a
leading /, backslashes, control characters and colons inside a segment are all
rejected; . is dropped and .. removes one preceding segment but may
not escape the root. A scheme is at least two characters, starts with a lowercase ASCII
letter, and otherwise contains lowercase letters, digits or -. That two-character
minimum exists for one reason: it stops a Windows drive prefix like C: from
being read as a scheme.
SourceId::parse with no scheme defaults to memory, which is why the
bare name "api/root.did" becomes memory:/api/root.did. The type also
has FromStr and TryFrom<&str>, and its
Deserialize routes through parse — so a decoded ID is always already
canonical.
The SourceResolver capability#
This is the whole interface between candid-core and wherever your sources live. Two required methods; the rest have defaults.
pub trait SourceResolver {
fn identify(&self, from: Option<&SourceId>, import: &str) -> Result<SourceId, ResolveError>;
fn load(&self, id: &SourceId, limits: &Limits) -> Result<ResolvedSource, ResolveError>;
/// Loads with cooperative runtime cancellation and deadline checks.
/// Implementations that do long-running work may override this method to
/// checkpoint internally.
fn load_with_context(
&self,
id: &SourceId,
context: &crate::RuntimeContext,
) -> Result<ResolvedSource, ResolveError> { /* default */ }
fn resolve(
&self,
from: Option<&SourceId>,
import: &str,
limits: &Limits,
) -> Result<ResolvedSource, ResolveError> { /* default */ }
fn resolve_with_context(
&self,
from: Option<&SourceId>,
import: &str,
context: &crate::RuntimeContext,
) -> Result<ResolvedSource, ResolveError> { /* default */ }
}
identify answers "given the import string import, written inside
source from (or None for the entry), what is its canonical ID?"
load answers "what are the bytes for that ID?" Nothing else is delegated: depth
limits, byte limits, cycle detection, digest checking and provenance all stay on the
compiler's side.
The supported boundary is synchronous, immutable source data supplied by the
host. There is no async variant and no future to poll, because the
compiler consumes one bundle that does not change while it is being compiled. If your sources
need fetching, fetch them first and hand over the result.
What a resolver returns, and how it is checked#
pub struct ResolvedSource {
pub id: SourceId,
pub source: String,
pub digest: String,
}
impl ResolvedSource {
pub fn verify(&self) -> Result<(), ResolveError>;
}
digest is sha256:<hex> over the text. A custom resolver has
three ways to be rejected before its bytes are used at all:
| Failure | Code |
|---|---|
identify returned an ID that is not already canonical | did_invalid_source_id |
load returned a source whose id differs from the one requested | did_resolver_identity_mismatch |
The declared digest does not match the returned bytes | did_source_digest_mismatch |
Failures come back as ResolveError { code, message, resource_limit }. The
code is the stable, machine-matchable part; message text is not an interface. All
three fields are public, so a resolver you write constructs one directly.
The two resolvers that ship#
MemoryResolver#
A map the host fills in. IDs must use the memory scheme; a bare relative name is
normalised into it. This is the resolver for tests, editors, agents, and any host — a browser
included — that already holds its sources.
impl MemoryResolver {
pub fn new() -> Self;
pub fn insert(
&mut self,
id: impl AsRef<str>,
source: impl Into<String>,
) -> Result<(), ResolveError>;
pub fn with_source(
mut self,
id: impl AsRef<str>,
source: impl Into<String>,
) -> Result<Self, ResolveError>;
}
A non-memory entry ID, or an attempt to resolve from a non-memory
source, fails with did_source_scheme_mismatch. with_source is the
chaining form of insert.
WorkspaceResolver and its sandbox#
WorkspaceResolver is the one resolver that maps logical
workspace:/… segments onto native paths, which is why it — and the
cap-std capability crate it uses — lives behind
filesystem-compiler rather than
compiler.
impl WorkspaceResolver {
pub fn new(root: impl AsRef<Path>) -> Result<Self, ResolveError>;
pub fn root(&self) -> &Path;
}
Construction opens a directory capability for the authorised root and holds it. Every
subsequent path is opened relative to that same open handle, never re-derived from a string
and re-opened from the top. Two properties follow. An absolute symlink out of the root, or a
.. that pops past it, fails with
did_import_outside_workspace, while an ordinary permission denial keeps the
distinct code did_file_read_error — you can tell "this tried to escape" from
"this file is unreadable". And because authorisation and reading use one handle, an attacker
who replaces a path between the two cannot swap in a file from outside; a test races symlink
replacement against two thousand loads and asserts no read ever escapes.
Reads are bounded as they happen: a source is read through a capped reader and exceeding
max_source_bytes produces a structured source_bytes resource
failure rather than a large allocation.
How imports become a bundle#
A bundle is the entry source plus every source reachable through its imports,
each loaded exactly once and treated as immutable for the rest of the compile. The bundle is
what gets type-checked, and it is what source_bundle_id hashes.
Candid has two import forms, and the loader records which one produced each edge:
import "types.did";-
Recorded as
SourceImportKind::Type. Brings in that file's type declarations. import service "registry.did";-
Recorded as
SourceImportKind::Service. Additionally brings in that file's main service, so the entry's actor can be resolved from it. A source reached only by such an edge that declares no main service is an error, reported against that source.
Diamonds load once#
The loader keys already-loaded sources by canonical SourceId. When the same
target is reached again it does not re-load; it only records whether any edge to it
was a service import. This conformance test wraps a MemoryResolver in a counter
and pins the behaviour — four sources, four load calls, even though
common.did is reached through both a.did and b.did:
let mut inner = MemoryResolver::new();
inner
.insert(
"root.did",
r#"import "a.did"; import "b.did"; service : { get: () -> (Common) };"#,
)
.unwrap();
inner
.insert("a.did", r#"import "common.did"; type A = Common;"#)
.unwrap();
inner
.insert("b.did", r#"import "common.did"; type B = Common;"#)
.unwrap();
inner.insert("common.did", "type Common = nat;").unwrap();
let loads = Arc::new(AtomicUsize::new(0));
let resolver = CountingResolver {
inner,
loads: loads.clone(),
};
let compilation = compile_with_resolver(
"root.did",
&resolver,
CompileOptions::default(),
&RuntimeContext::default(),
)
.unwrap();
assert_eq!(compilation.source_info().unwrap().sources().len(), 4);
assert_eq!(loads.load(Ordering::SeqCst), 4);
Wrapping load like this is also the practical way to see exactly which sources a
compile pulled in — the resolver is the only thing that touches them.
What bounds the traversal#
The import graph is walked with an explicit work list, not recursion, and every step is
charged against the caller's budget. The limits that apply are all defaults of the
interactive_v1 profile:
| Limit | Default | What it bounds | Resource name on failure |
|---|---|---|---|
max_import_depth | 64 | How far a chain of imports may nest below the entry | import_depth |
max_import_edges | 1024 | Total import statements across the whole bundle | import_edges |
max_sources | 256 | Distinct sources loaded | sources |
max_source_bytes | 1 MiB | Any single source | source_bytes |
max_bundle_bytes | 8 MiB | All sources together | bundle_bytes |
max_source_id_bytes | 1024 | A logical ID, and separately each import spelling | source_id_bytes |
Exhaustion produces a resource_limit_exceeded diagnostic in the
Load phase carrying a { resource, limit, observed } triple — not a
panic and not a truncated result. A cycle is caught separately and reported as
did_import_cycle. The import-spelling check is deliberately run last, after the
whole graph has loaded, so a resolver that aliases a very long spelling onto a short canonical
target cannot slip past the check the provenance sidecar performs later.
A worked example: a bundle with no filesystem#
This is a runnable program in the repository — cargo run --example
hermetic_bundle. It needs only the compiler feature,
which is why the identical code runs in a browser.
//! Imported-bundle compilation with no filesystem.
//!
//! The host holds every source itself and hands the compiler one immutable
//! logical bundle. Nothing is read from or written to disk, so this is
//! `compiler` surface — the same code a browser-WASM host runs.
use candid_core::{compile_with_resolver, CompileOptions, MemoryResolver, RuntimeContext};
use std::error::Error;
fn main() -> Result<(), Box<dyn Error>> {
let mut resolver = MemoryResolver::new();
resolver.insert(
"api/root.did",
r#"import "types.did"; service : { read: () -> (Item) query };"#,
)?;
resolver.insert(
"api/types.did",
"type Item = record { id: nat64; label: text };",
)?;
let compilation = compile_with_resolver(
"api/root.did",
&resolver,
CompileOptions::default(),
&RuntimeContext::default(),
)?;
let source_info = compilation
.source_info()
.ok_or("source provenance was not requested")?;
println!("contract: {}", compilation.contract().contract_id());
println!("source bundle: {}", source_info.source_bundle_id());
for source in source_info.sources() {
println!("- {} ({} bytes)", source.name, source.source.len());
}
Ok(())
}
The two bare names normalise to memory:/api/root.did and
memory:/api/types.did. The import "types.did" inside
root.did resolves relative to the importing source's directory, which is why the
second source has to live under api/ too. CompileOptions::default()
sets include_source_info: true, which is why source_info() is
Some; pass CompileOptions { include_source_info: false } when you
want only the Contract. The Contract and its identities are byte-identical either way.
The provenance sidecar#
The canonical Contract deliberately throws away everything that is about the source text
rather than the wire semantics. SourceInfo is where that goes: a separate,
separately versioned document produced alongside the Contract, bound to it by
contract_id and part of no identity.
| Field | What it records |
|---|---|
source_info_version | Currently 1. Any other value is rejected with unsupported_source_info_version. |
contract_id | The Contract this sidecar belongs to; a mismatch is source_contract_id_mismatch. |
source_bundle_id | A content identity over the raw sources and import edges alone. |
sources | { name, source } for the entry and every resolved import — the exact text, comments included. |
imports | One entry per edge: { from, import, to, kind }. import is the caller's original spelling, verbatim; to is the canonical target the resolver chose. |
declarations | Which source each type Foo = … came from, plus its doc comments. |
field_labels | For every source occurrence of a record or variant label: its origin, an AST-shaped occurrence path, the container node, the wire id, the SourceLabel, and doc comments. |
methods | Method occurrences with origin, path and doc comments. |
function_arguments | Argument and result names by position — read: (id: nat) -> … keeps id. |
actors | Which source declared the service block, plus its doc comments. |
SourceLabel is the piece that only the sidecar can supply. Candid hashes a named
label, an explicit numeric label and a positional tuple field to the same kind of 32-bit
number, so the Contract cannot tell them apart. The sidecar can:
/// Source spelling is intentionally separate from the semantic field ID.
/// `positional` differentiates tuple syntax from an explicitly numeric label.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(tag = "kind", rename_all = "snake_case", deny_unknown_fields)]
pub enum SourceLabel {
Named { name: String },
Numeric,
Positional,
}Here is a real sidecar, the provenance half of an artifact-identity fixture, for a two-file bundle very like the example above:
"source_info": {
"source_info_version": 1,
"contract_id": "candid-core:contract:v1:sha256:5b4d7090e72cefb298b0d5e2941dc2ed1f6ac44957163d75dd406ac5e30c930d",
"source_bundle_id": "candid-core:source-bundle:v1:sha256:52cb9ba7ed7105a36de5a5f2665ef82261080406133498ed4aa8b3cac6f9bcca",
"sources": [
{
"name": "memory:/root.did",
"source": "import \"types.did\";\n/// Ledger service.\nservice : { read: (id: nat) -> (Item) query };\n"
},
{
"name": "memory:/types.did",
"source": "/// An item.\ntype Item = record { id: nat; label: text };\n"
}
],
"imports": [
{
"from": "memory:/root.did",
"import": "types.did",
"to": "memory:/types.did",
"kind": "type"
}
],
… "declarations", "field_labels" and "methods" elided …
"function_arguments": [
{
"origin": {
"kind": "actor",
"source": "memory:/root.did"
},
"path": "actor.methods[0].function.args[0]",
"function": 1,
"direction": "argument",
"position": 0,
"name": "id"
}
],
"actors": [
{
"source": "memory:/root.did",
"docs": [
"/ Ledger service."
]
}
]
}Provenance is re-derived, not trusted#
This is the property that makes the sidecar worth shipping. When somebody hands you a
SourceInfo, candid-core does not check that it looks internally consistent. It
takes the source bundle embedded inside it, recompiles it through the same backend
that produced it, and accepts the sidecar only if the re-derived Contract identity and every
derived collection match exactly.
pub fn try_from_raw(
raw: RawSourceInfo,
contract: &Contract,
limits: &Limits,
) -> Result<Self, ContractValidationError>
Recompiles the embedded bundle and compares. A bundle that re-derives a different Contract
fails with source_contract_rederivation_mismatch; a presented collection that
differs from the re-derived one fails with
source_info_provenance_mismatch, and the violation's path names
which collection disagreed — $.sources, $.imports,
$.declarations, $.field_labels, $.methods,
$.function_arguments or $.actors.
try_from_raw_with_context takes a resource policy.
The rederivation runs through compile_resolved_bundle — the same internal
function compile_with_resolver calls — so public compilation and provenance
authentication cannot drift apart into two implementations that quietly disagree. A validated
SourceInfo therefore means "this matched a fresh recompilation on your budget",
not "this parsed".
One consequence worth internalising: source_bundle_id and
contract_id answer different questions. Editing a comment or reformatting a file
moves the bundle ID and leaves the Contract ID exactly where it was, because the first hashes
bytes and the second hashes meaning. Cache compiled output on the bundle ID; ask "is this the
same interface?" with the Contract or interface ID.
Content-addressed identities has the full table.
The hermetic guarantee, stated plainly#
candid-core never fetches anything, never resolves asynchronously, and never consults
a registry. Its entire production dependency set is serde,
serde_json, sha2, hex, and — behind features —
candid, candid_parser, ic_principal and
cap-std. None of those is an HTTP client. The only filesystem access on the whole
surface is WorkspaceResolver, which requires a feature you enable and a root you
name.
A resolver you write can front whatever you like — a package registry, an editor's open buffers, a browser cache — and the point is that this authority is then explicitly yours. A host can put a permission prompt in front of the resolver, because the resolver is the single place where anything is reached for.
Two target facts#
On wasm32-unknown-unknown, cap-std is not in the dependency graph
at all — it is declared under cfg(not(target_os = "unknown")). The
filesystem-compiler items still compile for that
target; they have no filesystem to reach. WorkspaceResolver::new
returns did_workspace_root_error, and load returns
did_file_read_error. Use MemoryResolver or your own resolver
there.
compile_did takes one self-contained string and does no import resolution. If
the parsed source declares any import or import service, it fails
in the Load phase with did_import_requires_file, one note per
import path, and a message that names both alternatives:
"DID source contains imports; supply the bundle through a SourceResolver and compile it with compile_with_resolver, or use compile_did_file to read it from a native workspace"
Do not work around this by concatenating your files. import service merge
semantics are not textual, and the concatenation would compile to something else.
The third entry point, compile_did_file
(filesystem-compiler), roots a
WorkspaceResolver at the file's parent directory and uses the file name
as the entry. It is not a wrapper around compile_with_resolver: it deliberately
keeps a second, native backend that writes the loaded bundle into a private temporary
directory and hands the entry to candid_parser::check_file, which is what keeps
the file checker's own diagnostic mapping. So
compile_did_file("api/root.did") authorises api/
and nothing above it, and an import of "../shared/types.did" pops past the root
and fails with did_import_outside_workspace. If your bundle spans directories,
root a WorkspaceResolver higher yourself and call
compile_with_resolver.
Next#
Compiling Candid sources covers the three entry points and their option and context variants in full. Limits, budgets and diagnostics explains the budget the loader charges against, including deadlines and cancellation, and the Contract graph describes the structure all of this produces.