Limits, budgets and diagnostics
Every entry point is bounded, cancellable and structured. This is what makes untrusted .did input safe to accept.
A .did file is source text, and a tool that reads one is usually reading bytes it
did not write: pasted into an editor, fetched from a canister, produced by an agent, or pulled
from a remote registry. Valid Candid structure is not the same as safe cost. A source can be
perfectly well formed and still be shaped so that reading it exhausts memory, spins the CPU, or
overflows the call stack.
candid-core answers that in one place rather than leaving it to each caller. Every public entry
point that touches untrusted bytes (parse, compile, validate, canonicalize, hash) runs under a
resource policy — one you pass in, or Limits::default() for the no-argument
conveniences — spends one budget while it works, and, when that budget runs out, returns a
structured diagnostic naming the resource, the limit, and the value it observed. The rule is recorded in
ADR
0005: bound all untrusted work, avoid recursive execution, fail closed.
Three failure modes, one answer#
Three distinct things can go wrong, and each is bounded separately.
- Unbounded time
- A graph algorithm whose cost is quadratic or worse in something the author of the input chooses freely. A DAG of shared type aliases re-expanded once per incoming edge costs O(2n) node visits from a 932-byte source. Bounded by counted work units and by an optional deadline.
- Unbounded allocation
- A document that is small on the wire but expands into a large in-memory structure: a million fields, a megabyte of names, a deduplication memo holding one entry per traversal state. Bounded by byte gates applied before decoding, and by counts on every collection the graph holds.
- Unbounded recursion
-
A few kilobytes of
opt opt opt …that drives a recursive parser off the end of the call stack. Graph and import algorithms use explicit work queues instead of call-stack recursion, and nesting is counted lexically in the raw text before any recursive third-party parser is invoked at all, so depth is a budget decision rather than a crash.
In all three cases the intended outcome is the same: a structured error value, returned normally
rather than a panic or an unbounded run, and a failed operation never returns a partially
validated Contract — the crate's validated in-memory model of an interface. That is the designed
behaviour rather than a proven property: ADR 0005 is recorded as
Implemented, verification pending, and the crate itself names one way a host can still
lose the structured refusal — an allocator that gives out before a counter does aborts instead.
See max_type_preflight_work on the limits
reference.
Limits: the policy object#
Limits is a plain value carrying 27 numeric bounds plus one optional deadline. Its
fields are private, so there is no struct literal: you start from a named profile and chain
with_* builders, and you read values back through getters.
use candid_core::{Limits, LimitsProfile};
let limits = LimitsProfile::InteractiveV1
.limits()
.with_max_input_bytes(64 * 1024)
.with_deadline_unix_ms(Some(2_000_000_000_000));
assert_eq!(limits.max_input_bytes(), 64 * 1024);
LimitsProfile is a named, versioned set of frozen default numbers.
InteractiveV1 is the only released profile, and Limits::default() is
exactly LimitsProfile::InteractiveV1.limits(). "Frozen" is a real promise: those
numbers never change. A future retuning ships as a new enum variant with a new wire name rather
than as an edit to an existing one, which is what lets a serialized policy name its profile
instead of copying 27 platform-dependent values.
Two consequences follow. Limits never participate in Contract identity, so
raising one for a trusted workload cannot change any hash the crate computes. And every
limit accepts 0: a zero limit is a defined, fail-closed policy rather than a
rejected configuration, refusing any input that consumes that resource at all. The complete list
of defaults is on the limits reference.
RuntimeContext, cancellation and deadlines#
RuntimeContext is a Limits policy plus the host-local controls for one
operation — today, a cancellation token. Anything taking a &RuntimeContext can be
cancelled; anything taking a bare &Limits can only be bounded and deadlined. The
limits field is public but the token is not, so a struct literal deliberately fails
to compile and a future control can be added without breaking callers.
use candid_core::{CancellationToken, Limits, RuntimeContext};
let token = CancellationToken::new();
let context = RuntimeContext::new(Limits::default()).with_cancellation(token.clone());
// … hand `context` to a long compilation on a worker thread, then:
token.cancel();
A CancellationToken is a cloneable flag backed by one shared atomic; cancelling any
clone cancels them all. It is cooperative, not preemptive: nothing is killed mid-instruction, and
the operation notices at its next checkpoint and returns the operation_cancelled code
for its domain. The token is never serialized. A RuntimeContext serializes as
{"limits": …} and nothing else, and a context decoded from JSON comes back with a
fresh, uncancelled token.
A deadline is set as Limits::with_deadline_unix_ms(Some(unix_ms)). It is
configured in wall-clock milliseconds but enforced against a monotonic clock: the
budget converts the remaining duration into a monotonic instant once, when the operation begins,
so a clock adjustment mid-flight cannot lengthen or shorten an operation already running. Any
deadline at or before the current time, Some(0) included, makes every bounded
operation fail closed with operation_deadline_exceeded before it does any work. If the
platform cannot represent the remaining duration monotonically, the code treats the deadline as
already reached, because an explicit deadline must never silently become unbounded.
On that target SystemTime::now and Instant::now panic rather
than returning an error, so a deadline cannot be measured and a failed measurement cannot be
reported. Reading the clock anyway would abort the module, the one outcome the fail-closed rule
exists to prevent. The budget therefore records only whether a deadline was configured
and reports any such deadline as already elapsed, reaching you through the ordinary
operation_deadline_exceeded result. None stays unbounded exactly as it
does natively, and cancellation and every quantitative limit are unaffected. A host that needs a
real browser deadline should drive a CancellationToken from its own timer; the crate
takes no web-time, js-sys or wasm-bindgen production
dependency.
One shared budget per operation#
Each context-aware public operation creates exactly one internal budget, and every stage draws on that same instance: the byte gate, JSON decode, source loading, the depth preflight walks, lowering, validation, canonicalization, provenance checks. No stage gets a fresh allowance at its boundary. The budget holds the limits, the monotonic deadline snapshot, the cancellation token, and a map of per-resource running totals.
It meters two different kinds of thing, and the distinction is deliberate.
- Charged resources accumulate. Work units are charged, so a second charge of 2 against a limit of 3 fails reporting an observed value of 4.
- Observed resources record a high-water mark and take the maximum. Retained artifacts such as the input document or the type-node count are observed, so a later stage revalidating the same artifact does not count it twice.
"Shared" has a practical consequence: raising the structural limit that let you build
something does not by itself make that thing renderable. The
to_json_pretty_with_limits serializers validate before rendering and then charge the
emitted byte length against max_canonicalization_work, on top of whatever
construction already spent from the same counter. You can parse a document successfully and then
fail to print it back out.
One budget per operation, not per context: RuntimeContext builds a fresh
budget on every entry point call. Reusing a context reuses the policy and the cancellation token,
never a running allowance.
Lexical nesting versus semantic depth#
Two pairs of limits look redundant and are not. For DID source,
max_source_nesting scans the token stream of the raw text and counts open delimiters
plus any run of opt/vec constructors, before the recursive upstream
parser is ever invoked; max_type_depth counts the checked semantic type depth
afterwards. For host values, the crate's self-describing JSON encoding of Candid values,
max_value_nesting counts JSON { and [ containers in a
constant-stack pre-scan before serde_json runs, while max_value_depth
counts semantic value depth after decoding.
They are separate because the units do not convert. One vec level costs two JSON
containers; one record level costs three. A single limit could not report an honest
observed value for both, and reporting a dishonest one would send you to adjust the
wrong knob. So a document rejected by the pre-scan always reports value_nesting,
never value_depth. The lexical check is also what keeps the rejection a budget
decision rather than a stack-exhaustion abort: rejecting costs constant stack, so no input can
force a crash by being deeper than the limit.
What a failure looks like#
Every failure domain in the crate (compilation, Contract validation, provenance validation, host
value validation) produces items of one serializable type, Diagnostic. The outer
collections stay domain-specific for Rust ergonomics (CompileError.diagnostics,
ContractValidationError.violations,
HostValueValidationError.violations), but the items are identical.
pub struct Diagnostic {
pub code: String,
pub phase: Option<DiagnosticPhase>,
pub severity: Option<Severity>,
pub path: Option<String>,
pub message: String,
pub span: Option<SourceSpan>,
pub related: Vec<RelatedLocation>,
pub notes: Vec<String>,
pub resource_limit: Option<ResourceLimitInfo>,
}
pub struct ResourceLimitInfo {
pub resource: String,
pub limit: u64,
pub observed: u64,
}
Domains differ only in which optional fields they fill in. Compile diagnostics always carry
phase (parse, type_check, load or
lower) and severity (error is the only variant today).
Validation violations always carry path, the $.… semantic or value path,
and never carry phase or severity. Absent fields are omitted from JSON
entirely, so each domain's serialized shape stays minimal.
ResourceLimitInfo is the part that makes a resource failure actionable. Its
limit and observed are fixed-width u64, not platform
usize, so the same failure serializes to the same numeric text on a 32-bit browser
target and a 64-bit server. Every conversion between domains preserves the triple verbatim; no
path in the crate is allowed to reduce a resource failure to message text.
SourceSpan comes in exactly two forms. An exact span carries
start_byte/end_byte offsets genuinely valid for the named source's
original text. A source-scoped location names a logical source with no offsets at
all, which is what a backend reports when it type-checked against rewritten text whose offsets
would point at the wrong characters in your file. Half-spans and empty spans are rejected on
decode.
The diagnostics cap and its sentinel#
Contract structure validation collects its violations up to max_diagnostics
(default 100); compile diagnostics come from the upstream checker and are not capped by it. When
more violations are observed than fit, the last retained slot is replaced by a resource_limit_exceeded
violation at path $ whose triple names resource diagnostics, the cap as
limit, and the true observed count as observed. Later observations update
that same sentinel in place, so the collection never grows past the cap.
A cap of 0 pins the guarantee. It can hold no real violation, so the sentinel becomes
the single retained item, which means an invalid input never yields an empty error
collection and you can always rely on violations[0] existing.
[
{
"code": "resource_limit_exceeded",
"path": "$",
"message": "resource diagnostics exceeded limit 0; observed at least 1",
"resource_limit": { "resource": "diagnostics", "limit": 0, "observed": 1 }
}
]Because the sentinel overwrites rather than appends, a cap of 1 with two violations leaves you the sentinel and nothing else. Set the cap high enough to keep the diagnostics you want to display.
code, the structured path, and the resource_limit triple
are the stable, machine-matchable surface. Human-readable message text is
explicitly not a stable interface and may be reworded in any release. If you are
writing code that reacts to a failure, match on code and read
resource_limit. No consumer should regex a message. The same applies to the
resource string itself: match it as data, because a resource name is not always the limit field
name minus its max_ prefix.
A limit tripping, end to end#
examples/bounded_parsing.rs is a runnable program that trips two different limits on
purpose. Run it with:
cargo run --example bounded_parsingIt first compiles a one-method service, renders the Contract to JSON, then tries to read that JSON back under a policy that allows 64 bytes of input:
// The byte gate runs before serde_json is invoked, so an oversized document
// is rejected without being decoded. This bounds peak allocation against
// the chosen ceiling; it does not reject element-by-element during decode.
let context = RuntimeContext::new(Limits::default().with_max_input_bytes(64));
let rejected =
Contract::from_slice_with_context(contract_json.as_bytes(), &context).unwrap_err();
println!("oversized Contract rejected: {}", resource_limit(&rejected));
The failure names input_bytes, with limit 64 and observed
set to the document's real byte length. The document is never decoded: the gate fires first, which
is what bounds peak allocation rather than merely detecting the problem afterwards. The
resource_limit helper in the same file does the only thing a caller should do, which
is find the first violation carrying resource_limit and read the triple off it.
The program then raises max_input_bytes enough to parse both documents, and starves
the render instead:
let render = raised.clone().with_max_canonicalization_work(16);
let starved_render = contract.to_json_pretty_with_limits(&render).unwrap_err();
let metadata = starved_render
.violations
.iter()
.find_map(|violation| violation.resource_limit.as_ref())
.ok_or("expected resource metadata")?;
println!(
"render starved: resource={} limit={} observed={}",
metadata.resource, metadata.limit, metadata.observed
);That is the shared budget made visible: the policy that successfully parsed the document still cannot print it, because rendering charges the emitted bytes against a counter the caller left low.
The serialized item is the same shape in the compile domain, with phase and
severity added. Starving max_canonicalization_work down to 1 and
compiling the empty service service : {}; produces exactly this, pinned by
tests/diagnostics_contract.rs:
[
{
"code": "resource_limit_exceeded",
"phase": "lower",
"severity": "error",
"path": "$",
"message": "$: resource canonicalization_work exceeded limit 1; observed 2",
"resource_limit": { "resource": "canonicalization_work", "limit": 1, "observed": 2 }
}
]
The structured violation crossed a domain boundary, from Contract validation into a compile
diagnostic, item by item — keeping its path and its triple rather than being
flattened into a message string.
Where the bounds are enforced#
Bounds are not a wrapper you can forget to apply. Contract,
ContractEnvelope, Compilation and HostValue deliberately do
not implement serde's Deserialize at all, because a trait implementation has no
argument position for a resource policy and could only ever decode under limits the library chose
rather than the ones you chose. Untrusted bytes go through from_json_with_limits,
from_slice_with_context and their siblings; the no-argument conveniences such as
Contract::from_json run the same bounded path under Limits::default(),
as does the candid-core binary, which has no flags for changing it.
All 27 limits, their default values, and the portable serialized configuration.
Rust crate Working with limitsProfiles, builders, deadlines and cancellation from Rust.
Reference Diagnostics referenceThe structured error shape shared by the Rust crate and the TypeScript runtime.