Sources, imports and provenance

How multi-file DID bundles are resolved hermetically, and how the result is re-derivable rather than trusted.

Real Candid interfaces are rarely one file. A .did can import another, that one can import a third, and two of them can import the same fourth. Something has to decide what each import string points at and produce the bytes. In most tools that something is ambient: the compiler opens whatever the path resolves to on the machine it happens to be running on.

candid-core does not do that. Import resolution is a capability you hand in: an object with two methods that turns an import spelling into a canonical identifier and returns bytes for it. The compiler never reaches for a file or a socket on its own, so whatever authority the resolver has is the only authority in play. The result is that the same bundle compiles identically in a CLI, in CI, and inside a browser — and that the exact sources it used can be recorded, shipped, and later re-derived rather than believed.

Logical source IDs#

Every source in a compile is named by a SourceId: a normalised logical URI of the form <scheme>:/<path>, such as memory:/api/types.did or workspace:/root.did. It is deliberately not a host path. Two hosts with different directory layouts that supply the same logical bundle get the same identities.

rustsrc/resolver.rs
pub struct SourceId(String);

impl SourceId {
    pub fn parse(value: impl AsRef<str>) -> Result<Self, ResolveError>;
    pub fn as_str(&self) -> &str;
    pub fn scheme(&self) -> &str;
    pub fn path(&self) -> &str;
}

The path grammar is UTF-8 with / separators, and it is strict. Empty segments, a leading /, backslashes, control characters and colons inside a segment are all rejected; . is dropped and .. removes one preceding segment but may not escape the root. A scheme is at least two characters, starts with a lowercase ASCII letter, and otherwise contains lowercase letters, digits or -. That two-character minimum exists for one reason: it stops a Windows drive prefix like C: from being read as a scheme.

SourceId::parse with no scheme defaults to memory, which is why the bare name "api/root.did" becomes memory:/api/root.did. The type also has FromStr and TryFrom<&str>, and its Deserialize routes through parse — so a decoded ID is always already canonical.

The SourceResolver capability#

This is the whole interface between candid-core and wherever your sources live. Two required methods; the rest have defaults.

rustsrc/resolver.rs
pub trait SourceResolver {
    fn identify(&self, from: Option<&SourceId>, import: &str) -> Result<SourceId, ResolveError>;

    fn load(&self, id: &SourceId, limits: &Limits) -> Result<ResolvedSource, ResolveError>;

    /// Loads with cooperative runtime cancellation and deadline checks.
    /// Implementations that do long-running work may override this method to
    /// checkpoint internally.
    fn load_with_context(
        &self,
        id: &SourceId,
        context: &crate::RuntimeContext,
    ) -> Result<ResolvedSource, ResolveError> { /* default */ }

    fn resolve(
        &self,
        from: Option<&SourceId>,
        import: &str,
        limits: &Limits,
    ) -> Result<ResolvedSource, ResolveError> { /* default */ }

    fn resolve_with_context(
        &self,
        from: Option<&SourceId>,
        import: &str,
        context: &crate::RuntimeContext,
    ) -> Result<ResolvedSource, ResolveError> { /* default */ }
}

identify answers "given the import string import, written inside source from (or None for the entry), what is its canonical ID?" load answers "what are the bytes for that ID?" Nothing else is delegated: depth limits, byte limits, cycle detection, digest checking and provenance all stay on the compiler's side.

The supported boundary is synchronous, immutable source data supplied by the host. There is no async variant and no future to poll, because the compiler consumes one bundle that does not change while it is being compiled. If your sources need fetching, fetch them first and hand over the result.

What a resolver returns, and how it is checked#

rustsrc/resolver.rs
pub struct ResolvedSource {
    pub id: SourceId,
    pub source: String,
    pub digest: String,
}

impl ResolvedSource {
    pub fn verify(&self) -> Result<(), ResolveError>;
}

digest is sha256:<hex> over the text. A custom resolver has three ways to be rejected before its bytes are used at all:

FailureCode
identify returned an ID that is not already canonicaldid_invalid_source_id
load returned a source whose id differs from the one requesteddid_resolver_identity_mismatch
The declared digest does not match the returned bytesdid_source_digest_mismatch

Failures come back as ResolveError { code, message, resource_limit }. The code is the stable, machine-matchable part; message text is not an interface. All three fields are public, so a resolver you write constructs one directly.

The two resolvers that ship#

MemoryResolver#

A map the host fills in. IDs must use the memory scheme; a bare relative name is normalised into it. This is the resolver for tests, editors, agents, and any host — a browser included — that already holds its sources.

rustsrc/resolver.rs
impl MemoryResolver {
    pub fn new() -> Self;
    pub fn insert(
        &mut self,
        id: impl AsRef<str>,
        source: impl Into<String>,
    ) -> Result<(), ResolveError>;
    pub fn with_source(
        mut self,
        id: impl AsRef<str>,
        source: impl Into<String>,
    ) -> Result<Self, ResolveError>;
}

A non-memory entry ID, or an attempt to resolve from a non-memory source, fails with did_source_scheme_mismatch. with_source is the chaining form of insert.

WorkspaceResolver and its sandbox#

WorkspaceResolver is the one resolver that maps logical workspace:/… segments onto native paths, which is why it — and the cap-std capability crate it uses — lives behind filesystem-compiler rather than compiler.

rustsrc/resolver.rs
impl WorkspaceResolver {
    pub fn new(root: impl AsRef<Path>) -> Result<Self, ResolveError>;
    pub fn root(&self) -> &Path;
}

Construction opens a directory capability for the authorised root and holds it. Every subsequent path is opened relative to that same open handle, never re-derived from a string and re-opened from the top. Two properties follow. An absolute symlink out of the root, or a .. that pops past it, fails with did_import_outside_workspace, while an ordinary permission denial keeps the distinct code did_file_read_error — you can tell "this tried to escape" from "this file is unreadable". And because authorisation and reading use one handle, an attacker who replaces a path between the two cannot swap in a file from outside; a test races symlink replacement against two thousand loads and asserts no read ever escapes.

Reads are bounded as they happen: a source is read through a capped reader and exceeding max_source_bytes produces a structured source_bytes resource failure rather than a large allocation.

How imports become a bundle#

A bundle is the entry source plus every source reachable through its imports, each loaded exactly once and treated as immutable for the rest of the compile. The bundle is what gets type-checked, and it is what source_bundle_id hashes.

Candid has two import forms, and the loader records which one produced each edge:

import "types.did";
Recorded as SourceImportKind::Type. Brings in that file's type declarations.
import service "registry.did";
Recorded as SourceImportKind::Service. Additionally brings in that file's main service, so the entry's actor can be resolved from it. A source reached only by such an edge that declares no main service is an error, reported against that source.

Diamonds load once#

The loader keys already-loaded sources by canonical SourceId. When the same target is reached again it does not re-load; it only records whether any edge to it was a service import. This conformance test wraps a MemoryResolver in a counter and pins the behaviour — four sources, four load calls, even though common.did is reached through both a.did and b.did:

rusttests/adr_conformance.rs
let mut inner = MemoryResolver::new();
inner
    .insert(
        "root.did",
        r#"import "a.did"; import "b.did"; service : { get: () -> (Common) };"#,
    )
    .unwrap();
inner
    .insert("a.did", r#"import "common.did"; type A = Common;"#)
    .unwrap();
inner
    .insert("b.did", r#"import "common.did"; type B = Common;"#)
    .unwrap();
inner.insert("common.did", "type Common = nat;").unwrap();
let loads = Arc::new(AtomicUsize::new(0));
let resolver = CountingResolver {
    inner,
    loads: loads.clone(),
};
let compilation = compile_with_resolver(
    "root.did",
    &resolver,
    CompileOptions::default(),
    &RuntimeContext::default(),
)
.unwrap();
assert_eq!(compilation.source_info().unwrap().sources().len(), 4);
assert_eq!(loads.load(Ordering::SeqCst), 4);

Wrapping load like this is also the practical way to see exactly which sources a compile pulled in — the resolver is the only thing that touches them.

What bounds the traversal#

The import graph is walked with an explicit work list, not recursion, and every step is charged against the caller's budget. The limits that apply are all defaults of the interactive_v1 profile:

LimitDefaultWhat it boundsResource name on failure
max_import_depth64How far a chain of imports may nest below the entryimport_depth
max_import_edges1024Total import statements across the whole bundleimport_edges
max_sources256Distinct sources loadedsources
max_source_bytes1 MiBAny single sourcesource_bytes
max_bundle_bytes8 MiBAll sources togetherbundle_bytes
max_source_id_bytes1024A logical ID, and separately each import spellingsource_id_bytes

Exhaustion produces a resource_limit_exceeded diagnostic in the Load phase carrying a { resource, limit, observed } triple — not a panic and not a truncated result. A cycle is caught separately and reported as did_import_cycle. The import-spelling check is deliberately run last, after the whole graph has loaded, so a resolver that aliases a very long spelling onto a short canonical target cannot slip past the check the provenance sidecar performs later.

A worked example: a bundle with no filesystem#

This is a runnable program in the repository — cargo run --example hermetic_bundle. It needs only the compiler feature, which is why the identical code runs in a browser.

rustexamples/hermetic_bundle.rs
//! Imported-bundle compilation with no filesystem.
//!
//! The host holds every source itself and hands the compiler one immutable
//! logical bundle. Nothing is read from or written to disk, so this is
//! `compiler` surface — the same code a browser-WASM host runs.

use candid_core::{compile_with_resolver, CompileOptions, MemoryResolver, RuntimeContext};
use std::error::Error;

fn main() -> Result<(), Box<dyn Error>> {
    let mut resolver = MemoryResolver::new();
    resolver.insert(
        "api/root.did",
        r#"import "types.did"; service : { read: () -> (Item) query };"#,
    )?;
    resolver.insert(
        "api/types.did",
        "type Item = record { id: nat64; label: text };",
    )?;

    let compilation = compile_with_resolver(
        "api/root.did",
        &resolver,
        CompileOptions::default(),
        &RuntimeContext::default(),
    )?;
    let source_info = compilation
        .source_info()
        .ok_or("source provenance was not requested")?;
    println!("contract: {}", compilation.contract().contract_id());
    println!("source bundle: {}", source_info.source_bundle_id());
    for source in source_info.sources() {
        println!("- {} ({} bytes)", source.name, source.source.len());
    }
    Ok(())
}

The two bare names normalise to memory:/api/root.did and memory:/api/types.did. The import "types.did" inside root.did resolves relative to the importing source's directory, which is why the second source has to live under api/ too. CompileOptions::default() sets include_source_info: true, which is why source_info() is Some; pass CompileOptions { include_source_info: false } when you want only the Contract. The Contract and its identities are byte-identical either way.

The provenance sidecar#

The canonical Contract deliberately throws away everything that is about the source text rather than the wire semantics. SourceInfo is where that goes: a separate, separately versioned document produced alongside the Contract, bound to it by contract_id and part of no identity.

FieldWhat it records
source_info_versionCurrently 1. Any other value is rejected with unsupported_source_info_version.
contract_idThe Contract this sidecar belongs to; a mismatch is source_contract_id_mismatch.
source_bundle_idA content identity over the raw sources and import edges alone.
sources{ name, source } for the entry and every resolved import — the exact text, comments included.
importsOne entry per edge: { from, import, to, kind }. import is the caller's original spelling, verbatim; to is the canonical target the resolver chose.
declarationsWhich source each type Foo = … came from, plus its doc comments.
field_labelsFor every source occurrence of a record or variant label: its origin, an AST-shaped occurrence path, the container node, the wire id, the SourceLabel, and doc comments.
methodsMethod occurrences with origin, path and doc comments.
function_argumentsArgument and result names by position — read: (id: nat) -> … keeps id.
actorsWhich source declared the service block, plus its doc comments.

SourceLabel is the piece that only the sidecar can supply. Candid hashes a named label, an explicit numeric label and a positional tuple field to the same kind of 32-bit number, so the Contract cannot tell them apart. The sidecar can:

rustsrc/model/source_info.rs
/// Source spelling is intentionally separate from the semantic field ID.
/// `positional` differentiates tuple syntax from an explicitly numeric label.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(tag = "kind", rename_all = "snake_case", deny_unknown_fields)]
pub enum SourceLabel {
    Named { name: String },
    Numeric,
    Positional,
}

Here is a real sidecar, the provenance half of an artifact-identity fixture, for a two-file bundle very like the example above:

jsontests/fixtures/artifact-identity/artifacts/compilation.json
  "source_info": {
    "source_info_version": 1,
    "contract_id": "candid-core:contract:v1:sha256:5b4d7090e72cefb298b0d5e2941dc2ed1f6ac44957163d75dd406ac5e30c930d",
    "source_bundle_id": "candid-core:source-bundle:v1:sha256:52cb9ba7ed7105a36de5a5f2665ef82261080406133498ed4aa8b3cac6f9bcca",
    "sources": [
      {
        "name": "memory:/root.did",
        "source": "import \"types.did\";\n/// Ledger service.\nservice : { read: (id: nat) -> (Item) query };\n"
      },
      {
        "name": "memory:/types.did",
        "source": "/// An item.\ntype Item = record { id: nat; label: text };\n"
      }
    ],
    "imports": [
      {
        "from": "memory:/root.did",
        "import": "types.did",
        "to": "memory:/types.did",
        "kind": "type"
      }
    ],

    … "declarations", "field_labels" and "methods" elided …

    "function_arguments": [
      {
        "origin": {
          "kind": "actor",
          "source": "memory:/root.did"
        },
        "path": "actor.methods[0].function.args[0]",
        "function": 1,
        "direction": "argument",
        "position": 0,
        "name": "id"
      }
    ],
    "actors": [
      {
        "source": "memory:/root.did",
        "docs": [
          "/ Ledger service."
        ]
      }
    ]
  }

Provenance is re-derived, not trusted#

This is the property that makes the sidecar worth shipping. When somebody hands you a SourceInfo, candid-core does not check that it looks internally consistent. It takes the source bundle embedded inside it, recompiles it through the same backend that produced it, and accepts the sidecar only if the re-derived Contract identity and every derived collection match exactly.

SourceInfo::try_from_raw compiler
rust
pub fn try_from_raw(
    raw: RawSourceInfo,
    contract: &Contract,
    limits: &Limits,
) -> Result<Self, ContractValidationError>

Recompiles the embedded bundle and compares. A bundle that re-derives a different Contract fails with source_contract_rederivation_mismatch; a presented collection that differs from the re-derived one fails with source_info_provenance_mismatch, and the violation's path names which collection disagreed — $.sources, $.imports, $.declarations, $.field_labels, $.methods, $.function_arguments or $.actors. try_from_raw_with_context takes a resource policy.

The rederivation runs through compile_resolved_bundle — the same internal function compile_with_resolver calls — so public compilation and provenance authentication cannot drift apart into two implementations that quietly disagree. A validated SourceInfo therefore means "this matched a fresh recompilation on your budget", not "this parsed".

One consequence worth internalising: source_bundle_id and contract_id answer different questions. Editing a comment or reformatting a file moves the bundle ID and leaves the Contract ID exactly where it was, because the first hashes bytes and the second hashes meaning. Cache compiled output on the bundle ID; ask "is this the same interface?" with the Contract or interface ID. Content-addressed identities has the full table.

The hermetic guarantee, stated plainly#

candid-core never fetches anything, never resolves asynchronously, and never consults a registry. Its entire production dependency set is serde, serde_json, sha2, hex, and — behind features — candid, candid_parser, ic_principal and cap-std. None of those is an HTTP client. The only filesystem access on the whole surface is WorkspaceResolver, which requires a feature you enable and a root you name.

A resolver you write can front whatever you like — a package registry, an editor's open buffers, a browser cache — and the point is that this authority is then explicitly yours. A host can put a permission prompt in front of the resolver, because the resolver is the single place where anything is reached for.

Two target facts#

WorkspaceResolver has nothing to open on browser WASM

On wasm32-unknown-unknown, cap-std is not in the dependency graph at all — it is declared under cfg(not(target_os = "unknown")). The filesystem-compiler items still compile for that target; they have no filesystem to reach. WorkspaceResolver::new returns did_workspace_root_error, and load returns did_file_read_error. Use MemoryResolver or your own resolver there.

compile_did refuses a source that contains imports

compile_did takes one self-contained string and does no import resolution. If the parsed source declares any import or import service, it fails in the Load phase with did_import_requires_file, one note per import path, and a message that names both alternatives:

"DID source contains imports; supply the bundle through a SourceResolver and compile it with compile_with_resolver, or use compile_did_file to read it from a native workspace"

Do not work around this by concatenating your files. import service merge semantics are not textual, and the concatenation would compile to something else.

The third entry point, compile_did_file (filesystem-compiler), roots a WorkspaceResolver at the file's parent directory and uses the file name as the entry. It is not a wrapper around compile_with_resolver: it deliberately keeps a second, native backend that writes the loaded bundle into a private temporary directory and hands the entry to candid_parser::check_file, which is what keeps the file checker's own diagnostic mapping. So compile_did_file("api/root.did") authorises api/ and nothing above it, and an import of "../shared/types.did" pops past the root and fails with did_import_outside_workspace. If your bundle spans directories, root a WorkspaceResolver higher yourself and call compile_with_resolver.

Next#

Compiling Candid sources covers the three entry points and their option and context variants in full. Limits, budgets and diagnostics explains the budget the loader charges against, including deadlines and cancellation, and the Contract graph describes the structure all of this produces.