Limits, budgets and diagnostics

Every entry point is bounded, cancellable and structured. This is what makes untrusted .did input safe to accept.

A .did file is source text, and a tool that reads one is usually reading bytes it did not write: pasted into an editor, fetched from a canister, produced by an agent, or pulled from a remote registry. Valid Candid structure is not the same as safe cost. A source can be perfectly well formed and still be shaped so that reading it exhausts memory, spins the CPU, or overflows the call stack.

candid-core answers that in one place rather than leaving it to each caller. Every public entry point that touches untrusted bytes (parse, compile, validate, canonicalize, hash) runs under a resource policy — one you pass in, or Limits::default() for the no-argument conveniences — spends one budget while it works, and, when that budget runs out, returns a structured diagnostic naming the resource, the limit, and the value it observed. The rule is recorded in ADR 0005: bound all untrusted work, avoid recursive execution, fail closed.

Three failure modes, one answer#

Three distinct things can go wrong, and each is bounded separately.

Unbounded time
A graph algorithm whose cost is quadratic or worse in something the author of the input chooses freely. A DAG of shared type aliases re-expanded once per incoming edge costs O(2n) node visits from a 932-byte source. Bounded by counted work units and by an optional deadline.
Unbounded allocation
A document that is small on the wire but expands into a large in-memory structure: a million fields, a megabyte of names, a deduplication memo holding one entry per traversal state. Bounded by byte gates applied before decoding, and by counts on every collection the graph holds.
Unbounded recursion
A few kilobytes of opt opt opt … that drives a recursive parser off the end of the call stack. Graph and import algorithms use explicit work queues instead of call-stack recursion, and nesting is counted lexically in the raw text before any recursive third-party parser is invoked at all, so depth is a budget decision rather than a crash.

In all three cases the intended outcome is the same: a structured error value, returned normally rather than a panic or an unbounded run, and a failed operation never returns a partially validated Contract — the crate's validated in-memory model of an interface. That is the designed behaviour rather than a proven property: ADR 0005 is recorded as Implemented, verification pending, and the crate itself names one way a host can still lose the structured refusal — an allocator that gives out before a counter does aborts instead. See max_type_preflight_work on the limits reference.

Limits: the policy object#

Limits is a plain value carrying 27 numeric bounds plus one optional deadline. Its fields are private, so there is no struct literal: you start from a named profile and chain with_* builders, and you read values back through getters.

rustsrc/limits.rs
use candid_core::{Limits, LimitsProfile};

let limits = LimitsProfile::InteractiveV1
    .limits()
    .with_max_input_bytes(64 * 1024)
    .with_deadline_unix_ms(Some(2_000_000_000_000));
assert_eq!(limits.max_input_bytes(), 64 * 1024);

LimitsProfile is a named, versioned set of frozen default numbers. InteractiveV1 is the only released profile, and Limits::default() is exactly LimitsProfile::InteractiveV1.limits(). "Frozen" is a real promise: those numbers never change. A future retuning ships as a new enum variant with a new wire name rather than as an edit to an existing one, which is what lets a serialized policy name its profile instead of copying 27 platform-dependent values.

Two consequences follow. Limits never participate in Contract identity, so raising one for a trusted workload cannot change any hash the crate computes. And every limit accepts 0: a zero limit is a defined, fail-closed policy rather than a rejected configuration, refusing any input that consumes that resource at all. The complete list of defaults is on the limits reference.

RuntimeContext, cancellation and deadlines#

RuntimeContext is a Limits policy plus the host-local controls for one operation — today, a cancellation token. Anything taking a &RuntimeContext can be cancelled; anything taking a bare &Limits can only be bounded and deadlined. The limits field is public but the token is not, so a struct literal deliberately fails to compile and a future control can be added without breaking callers.

rust
use candid_core::{CancellationToken, Limits, RuntimeContext};

let token = CancellationToken::new();
let context = RuntimeContext::new(Limits::default()).with_cancellation(token.clone());

// … hand `context` to a long compilation on a worker thread, then:
token.cancel();

A CancellationToken is a cloneable flag backed by one shared atomic; cancelling any clone cancels them all. It is cooperative, not preemptive: nothing is killed mid-instruction, and the operation notices at its next checkpoint and returns the operation_cancelled code for its domain. The token is never serialized. A RuntimeContext serializes as {"limits": …} and nothing else, and a context decoded from JSON comes back with a fresh, uncancelled token.

A deadline is set as Limits::with_deadline_unix_ms(Some(unix_ms)). It is configured in wall-clock milliseconds but enforced against a monotonic clock: the budget converts the remaining duration into a monotonic instant once, when the operation begins, so a clock adjustment mid-flight cannot lengthen or shorten an operation already running. Any deadline at or before the current time, Some(0) included, makes every bounded operation fail closed with operation_deadline_exceeded before it does any work. If the platform cannot represent the remaining duration monotonically, the code treats the deadline as already reached, because an explicit deadline must never silently become unbounded.

Bare wasm32-unknown-unknown has no clock

On that target SystemTime::now and Instant::now panic rather than returning an error, so a deadline cannot be measured and a failed measurement cannot be reported. Reading the clock anyway would abort the module, the one outcome the fail-closed rule exists to prevent. The budget therefore records only whether a deadline was configured and reports any such deadline as already elapsed, reaching you through the ordinary operation_deadline_exceeded result. None stays unbounded exactly as it does natively, and cancellation and every quantitative limit are unaffected. A host that needs a real browser deadline should drive a CancellationToken from its own timer; the crate takes no web-time, js-sys or wasm-bindgen production dependency.

One shared budget per operation#

Each context-aware public operation creates exactly one internal budget, and every stage draws on that same instance: the byte gate, JSON decode, source loading, the depth preflight walks, lowering, validation, canonicalization, provenance checks. No stage gets a fresh allowance at its boundary. The budget holds the limits, the monotonic deadline snapshot, the cancellation token, and a map of per-resource running totals.

It meters two different kinds of thing, and the distinction is deliberate.

  • Charged resources accumulate. Work units are charged, so a second charge of 2 against a limit of 3 fails reporting an observed value of 4.
  • Observed resources record a high-water mark and take the maximum. Retained artifacts such as the input document or the type-node count are observed, so a later stage revalidating the same artifact does not count it twice.

"Shared" has a practical consequence: raising the structural limit that let you build something does not by itself make that thing renderable. The to_json_pretty_with_limits serializers validate before rendering and then charge the emitted byte length against max_canonicalization_work, on top of whatever construction already spent from the same counter. You can parse a document successfully and then fail to print it back out.

One budget per operation, not per context: RuntimeContext builds a fresh budget on every entry point call. Reusing a context reuses the policy and the cancellation token, never a running allowance.

Lexical nesting versus semantic depth#

Two pairs of limits look redundant and are not. For DID source, max_source_nesting scans the token stream of the raw text and counts open delimiters plus any run of opt/vec constructors, before the recursive upstream parser is ever invoked; max_type_depth counts the checked semantic type depth afterwards. For host values, the crate's self-describing JSON encoding of Candid values, max_value_nesting counts JSON { and [ containers in a constant-stack pre-scan before serde_json runs, while max_value_depth counts semantic value depth after decoding.

They are separate because the units do not convert. One vec level costs two JSON containers; one record level costs three. A single limit could not report an honest observed value for both, and reporting a dishonest one would send you to adjust the wrong knob. So a document rejected by the pre-scan always reports value_nesting, never value_depth. The lexical check is also what keeps the rejection a budget decision rather than a stack-exhaustion abort: rejecting costs constant stack, so no input can force a crash by being deeper than the limit.

What a failure looks like#

Every failure domain in the crate (compilation, Contract validation, provenance validation, host value validation) produces items of one serializable type, Diagnostic. The outer collections stay domain-specific for Rust ergonomics (CompileError.diagnostics, ContractValidationError.violations, HostValueValidationError.violations), but the items are identical.

rustsrc/diagnostics.rs
pub struct Diagnostic {
    pub code: String,
    pub phase: Option<DiagnosticPhase>,
    pub severity: Option<Severity>,
    pub path: Option<String>,
    pub message: String,
    pub span: Option<SourceSpan>,
    pub related: Vec<RelatedLocation>,
    pub notes: Vec<String>,
    pub resource_limit: Option<ResourceLimitInfo>,
}

pub struct ResourceLimitInfo {
    pub resource: String,
    pub limit: u64,
    pub observed: u64,
}

Domains differ only in which optional fields they fill in. Compile diagnostics always carry phase (parse, type_check, load or lower) and severity (error is the only variant today). Validation violations always carry path, the $.… semantic or value path, and never carry phase or severity. Absent fields are omitted from JSON entirely, so each domain's serialized shape stays minimal.

ResourceLimitInfo is the part that makes a resource failure actionable. Its limit and observed are fixed-width u64, not platform usize, so the same failure serializes to the same numeric text on a 32-bit browser target and a 64-bit server. Every conversion between domains preserves the triple verbatim; no path in the crate is allowed to reduce a resource failure to message text.

SourceSpan comes in exactly two forms. An exact span carries start_byte/end_byte offsets genuinely valid for the named source's original text. A source-scoped location names a logical source with no offsets at all, which is what a backend reports when it type-checked against rewritten text whose offsets would point at the wrong characters in your file. Half-spans and empty spans are rejected on decode.

The diagnostics cap and its sentinel#

Contract structure validation collects its violations up to max_diagnostics (default 100); compile diagnostics come from the upstream checker and are not capped by it. When more violations are observed than fit, the last retained slot is replaced by a resource_limit_exceeded violation at path $ whose triple names resource diagnostics, the cap as limit, and the true observed count as observed. Later observations update that same sentinel in place, so the collection never grows past the cap.

A cap of 0 pins the guarantee. It can hold no real violation, so the sentinel becomes the single retained item, which means an invalid input never yields an empty error collection and you can always rely on violations[0] existing.

json
[
  {
    "code": "resource_limit_exceeded",
    "path": "$",
    "message": "resource diagnostics exceeded limit 0; observed at least 1",
    "resource_limit": { "resource": "diagnostics", "limit": 0, "observed": 1 }
  }
]

Because the sentinel overwrites rather than appends, a cap of 1 with two violations leaves you the sentinel and nothing else. Set the cap high enough to keep the diagnostics you want to display.

Codes are the interface; message text is not

code, the structured path, and the resource_limit triple are the stable, machine-matchable surface. Human-readable message text is explicitly not a stable interface and may be reworded in any release. If you are writing code that reacts to a failure, match on code and read resource_limit. No consumer should regex a message. The same applies to the resource string itself: match it as data, because a resource name is not always the limit field name minus its max_ prefix.

A limit tripping, end to end#

examples/bounded_parsing.rs is a runnable program that trips two different limits on purpose. Run it with:

bash
cargo run --example bounded_parsing

It first compiles a one-method service, renders the Contract to JSON, then tries to read that JSON back under a policy that allows 64 bytes of input:

rustexamples/bounded_parsing.rs
// The byte gate runs before serde_json is invoked, so an oversized document
// is rejected without being decoded. This bounds peak allocation against
// the chosen ceiling; it does not reject element-by-element during decode.
let context = RuntimeContext::new(Limits::default().with_max_input_bytes(64));
let rejected =
    Contract::from_slice_with_context(contract_json.as_bytes(), &context).unwrap_err();
println!("oversized Contract rejected: {}", resource_limit(&rejected));

The failure names input_bytes, with limit 64 and observed set to the document's real byte length. The document is never decoded: the gate fires first, which is what bounds peak allocation rather than merely detecting the problem afterwards. The resource_limit helper in the same file does the only thing a caller should do, which is find the first violation carrying resource_limit and read the triple off it.

The program then raises max_input_bytes enough to parse both documents, and starves the render instead:

rustexamples/bounded_parsing.rs
let render = raised.clone().with_max_canonicalization_work(16);
let starved_render = contract.to_json_pretty_with_limits(&render).unwrap_err();
let metadata = starved_render
    .violations
    .iter()
    .find_map(|violation| violation.resource_limit.as_ref())
    .ok_or("expected resource metadata")?;
println!(
    "render starved: resource={} limit={} observed={}",
    metadata.resource, metadata.limit, metadata.observed
);

That is the shared budget made visible: the policy that successfully parsed the document still cannot print it, because rendering charges the emitted bytes against a counter the caller left low.

The serialized item is the same shape in the compile domain, with phase and severity added. Starving max_canonicalization_work down to 1 and compiling the empty service service : {}; produces exactly this, pinned by tests/diagnostics_contract.rs:

json
[
  {
    "code": "resource_limit_exceeded",
    "phase": "lower",
    "severity": "error",
    "path": "$",
    "message": "$: resource canonicalization_work exceeded limit 1; observed 2",
    "resource_limit": { "resource": "canonicalization_work", "limit": 1, "observed": 2 }
  }
]

The structured violation crossed a domain boundary, from Contract validation into a compile diagnostic, item by item — keeping its path and its triple rather than being flattened into a message string.

Where the bounds are enforced#

Bounds are not a wrapper you can forget to apply. Contract, ContractEnvelope, Compilation and HostValue deliberately do not implement serde's Deserialize at all, because a trait implementation has no argument position for a resource policy and could only ever decode under limits the library chose rather than the ones you chose. Untrusted bytes go through from_json_with_limits, from_slice_with_context and their siblings; the no-argument conveniences such as Contract::from_json run the same bounded path under Limits::default(), as does the candid-core binary, which has no flags for changing it.