The binary codec
Schema-directed Candid encoding and decoding in TypeScript, fail-closed on every input.
This page describes the surface the repository builds, published as
@candid-core/schema 0.3.0-beta.1 under the npm beta dist-tag. latest is still
0.2.0, which has the old one: a decoded
principal is an object with toText() and encoding accepts one, the encoded type table
depends on how the schema was built, and a misspelled limit is ignored.
Migrating from 0.2.0 lists every difference with before and
after code, and each is recorded in the package changelog under 0.3.0-beta.1.
@candid-core/schema/codec turns JavaScript values into Candid wire bytes and back.
It is written in TypeScript, has no runtime dependencies, and exports exactly six functions:
encode, encodeArgs, decode, decodeArgs,
principalTextFromBytes and principalBytesFromText. Everything it needs
to know about your data comes from a schema — the plain object a
c.* builder returns, the module the
generator emits, or a schema
built at runtime from a Contract document.
Nothing here throws for a value or a byte string. Every call returns a discriminated result you
branch on, and every failure carries a machine-readable code and a $-rooted path into
the offending value. That holds for hostile bytes and hostile values alike: a getter that throws
mid-encode becomes an unreadable_value issue, not an escaping exception. The one
thing that does throw is a malformed options object — see
the budgets below.
The four entry points#
A Candid message is an argument sequence, not a single value, so
encodeArgs and decodeArgs are the primary pair.
encode and decode wrap the one-argument case and re-root the issue
paths so they read exactly like validate's.
export function encodeArgs(
schemas: readonly AnyFieldSchema[],
values: readonly unknown[],
options: CodecOptions = {},
): EncodeResultexport type EncodeResult =
| { readonly ok: true; readonly bytes: Uint8Array }
| { readonly ok: false; readonly issues: readonly CodecIssue[] };
One schema and one value per argument. Mismatched lengths fail with
invalid_length before anything is written.
export function decodeArgs(
schemas: readonly AnyFieldSchema[],
bytes: Uint8Array,
options: CodecOptions = {},
): DecodeResultexport type DecodeResult =
| { readonly ok: true; readonly values: readonly unknown[] }
| { readonly ok: false; readonly issues: readonly CodecIssue[] };One decoded value per expected argument. Extra trailing wire arguments are skipped; a trailing argument the wire does not carry follows the record-field rule below.
export function encode<T>(
schema: Schema<T>,
value: unknown,
options: CodecOptions = {},
): EncodeResult
export function decode<T>(
schema: Schema<T>,
bytes: Uint8Array,
options: CodecOptions = {},
): { ok: true; value: unknown } | { ok: false; issues: readonly CodecIssue[] }The one-argument convenience wrappers.
The type parameter on decode<T> constrains the schema, not the
payload: a successful result is { ok: true; value: unknown }, and
decodeArgs returns readonly unknown[]. The bytes came from outside
your program, so the runtime does not hand you a static promise about them. Narrow the value
yourself, or pass it through validate — the reference-vector suite pins that each
decoded value validates ok against the schema it decoded under.
Here is the whole surface in one round trip:
import { c, principal, type Infer } from "@candid-core/schema";
import { validate } from "@candid-core/schema/validate";
import { encode, decode } from "@candid-core/schema/codec";
const Account = c.record({ owner: c.principal, balance: c.nat });
type Account = Infer<typeof Account>; // { owner: Principal; balance: bigint }
const value: Account = { owner: principal("aaaaa-aa"), balance: 5n };
validate(Account, value); // { ok: true } | { ok: false, issues }
const encoded = encode(Account, value); // { ok: true, bytes } | { ok: false, issues }
if (encoded.ok) {
decode(Account, encoded.bytes);
}
A decoded principal is its canonical text, a plain string typed as the branded
Principal — the codec never constructs an SDK class — so a decoded value
serializes with JSON.stringify, survives structuredClone, and compares
with ===. In the other direction encoding is strict: it accepts exactly canonical
principal text, the same strings validate accepts, so an @icp-sdk/core
Principal converts once with principal(sdkPrincipal).
What a Candid message looks like#
You do not need the spec to use the codec, but the four parts explain most of its behaviour. A message is:
- The magic — the four ASCII bytes
DIDL(44 49 44 4c). - A type table — a length-prefixed list of the composite types the message uses. Entries reference each other by index, which is how recursion is expressed, and references may point forward.
- An argument type list — one type reference per argument.
- The values — all the argument values, concatenated, with no separators.
Only composites get table entries. The eighteen primitives are negative opcodes written inline:
null is -1, nat is -3, text is
-15, principal is -24. The block -18 to
-23 is the composites — opt, vec, record,
variant, func, service — and a primitive opcode appearing
in the table, or a composite opcode appearing inline, fails with
malformed_type_table.
Three shapes have no opcode of their own, so the codec spells them out:
- blob is emitted as
vec nat8; - a tuple is a record whose field ids are exactly
0..n-1; - unit is the empty record.
Records and variants identify fields by a 32-bit number, not by name. Schema objects carry
readable keys, so both directions derive the number by two rules: a key spelled
_N_ is the id
N, and every other key is the Candid label hash of its UTF-8 bytes,
h(name) = fold(h * 223 + byte) mod 2^32. Two keys on one node that derive the same
id fail closed with duplicate_field_id rather than one silently winning —
c.record({ a: …, _97_: … }) is that collision, because the label hash of
a is 97.
Emission order is fixed by the format, not by your object: record fields are read and validated
in declaration order — so the first issue's path matches what validate reports —
and then the buffered bytes are reordered into ascending wire-id order. Service method names go
out sorted by raw UTF-8 byte sequence. Identical inputs therefore produce identical bytes, and
nothing depends on time, environment, or map iteration order.
The bytes also do not depend on how the schema objects were built. A generated module mints a
fresh c.opt(c.blob()) for every field that spells one; a schema loaded with
schemaFromContract shares one node per Contract type; a hand-written schema does
whatever its author did. The encoder walks whichever graph it is given, then rewrites the type
table into a structural canonical form before writing it: equal Candid types become one entry,
numbered in first-visit order from the argument types. So one call is one byte string — and a
sound cache or deduplication key — whether its schemas were generated, loaded or hand-built,
however they share nodes, and in whatever key order the schema or the value spells a record.
blob and vec nat8 are the same wire type and share an entry. A schema
with no repeated structure writes exactly the table it always did.
Recursive types are canonical per knot. Every schema generated from or loaded from a Contract has
one knot per recursive type, because the Contract canonicalizer has already minimised the graph,
so those agree. Two separately hand-built knots for one recursive type — Even and
Odd written as two c.rec calls when each is
record { next : opt Self } — are not merged: the bytes are larger, valid, and
decode to the same value, but they are different bytes. Minimising cyclic graphs in the encoder is
a recorded non-goal.
Schema-directed decoding#
Decoding is driven by the schema you expect, not by whatever the bytes claim. The message carries its own type description; the decoder checks that description against your expected schema using Candid's coercion relation, and then reads the value bytes as your schema says to read them. That is the difference between a decoded value whose shape you can rely on and one you have to inspect before you can use it.
It is also what lets an old client keep working against a canister that has since grown. The coercion rules are the spec's, and the asymmetry between options and variants is deliberate:
-
An expected
optabsorbs content mismatches tonull. If the wire carriestextwhere youroptwanted anat, you getnull. A bare wire value at an expectedoptof its type auto-wraps. -
An unknown variant tag is a hard error (
unknown_variant_tag) — unless an enclosing expectedoptabsorbs it. -
A missing non-optional record field is a hard error
(
missing_field); a missing field whose expected type is opt-like decodes asnull. - Extra wire record fields and extra trailing arguments are skipped — and charged the same budgets as decoded values.
-
A boxed
optfollows the same rules. For anoptwhose inner type admitsnull(opt opt T,opt null,opt reserved), a present value decodes as{ some: v }. A mismatch the inneroptabsorbs is{ some: null }, and one the outeroptabsorbs isnull, exactly as thecandidcrate decodes them.
Each behaviour below is an executed assertion in the coercion-matrix test
(ts/tests/codec.test.ts). Here a small helper unwraps the bytes, and the expected
results are in comments:
import { c } from "@candid-core/schema";
import { decode, encode, type EncodeResult } from "@candid-core/schema/codec";
// The bytes of an encode that must have succeeded.
function bytes(result: EncodeResult): Uint8Array {
if (!result.ok) throw new Error(result.issues[0].message);
return result.bytes;
}
// A wire nat is accepted at expected int; the reverse hard-errors.
decode(c.int, bytes(encode(c.nat, 5n))); // { ok: true, value: 5n }
decode(c.nat, bytes(encode(c.int, 5n))); // type_mismatch
// Extra wire record fields are skipped.
const wide = bytes(encode(c.record({ a: c.bool, b: c.text }), { a: true, b: "x" }));
decode(c.record({ a: c.bool }), wide); // { ok: true, value: { a: true } }
// A field the wire does not carry: null if opt-like, hard error otherwise.
const narrow = bytes(encode(c.record({ a: c.bool }), { a: true }));
decode(c.record({ a: c.bool, b: c.opt(c.text) }), narrow); // { a: true, b: null }
decode(c.record({ a: c.bool, b: c.text }), narrow); // missing_field
// Expected opt absorbs a constituent mismatch; a bare value auto-wraps.
decode(c.opt(c.nat), bytes(encode(c.text, "hi"))); // { ok: true, value: null }
decode(c.opt(c.text), bytes(encode(c.text, "hi"))); // { ok: true, value: "hi" }
// Variants trap on an unknown tag — unless an enclosing opt absorbs it.
const later = bytes(encode(c.variant({ ok: c.null, later: c.nat32 }), { tag: "later", value: 7 }));
decode(c.variant({ ok: c.null }), later); // unknown_variant_tag
decode(c.opt(c.variant({ ok: c.null })), later); // { ok: true, value: null }Absorption never hides a broken message. When a constituent mismatch is absorbed, the cursor rewinds and the value is skipped — but malformed bytes and exhausted budgets stay hard failures, and the element charges already spent on the rewound walk are not refunded. Refunding them would let nested options multiply the traversal budget by the absorption depth.
The budgets, and what a caller observes#
CodecOptions is the complete set of bounds, each optional and each defaulting to an
exported constant. The same object is accepted by encode and decode.
Options are code, not input, so they are checked strictly and up front. An own key that is not
one of the five below throws TypeError — a misspelled maxDeph used to
apply the default without a word, and maxIssues belongs to validate,
not here. So does a limit that is not a non-negative safe integer: NaN (which used to
switch its bound off, since nothing is greater than NaN), a negative, a fraction, a
string, null, or Infinity. A trusted host that wants no practical bound
passes a large integer. undefined means the default, and 0 is a valid
limit that refuses any input consuming that resource. Each option is read exactly once, and the
call runs on that checked copy: a getter or Proxy cannot pass the check with one value and run
with another. An options object whose getter or trap throws while being read raises a
TypeError too, carrying the original exception as its cause.
export interface CodecOptions {
/** Input size ceiling for decode, in bytes. */
readonly maxBytes?: number;
/**
* Wire type table entry cap, mirroring `Limits::max_type_nodes`. Decode
* refuses a wire table claiming more entries. Encode charges one entry per
* distinct composite schema node its walk meets, before repeated structure
* is merged, so the table it writes is never larger than this.
*/
readonly maxTypeTableEntries?: number;
/**
* Traversal depth cap, mirroring `Limits::max_value_depth`. Encode's
* type-table walk charges it too, for Candid nesting depth: once per
* combinator level at which the table gains an entry, never for a `rec`
* hop, so a static schema nested deeper than this is refused.
*/
readonly maxDepth?: number;
/** Traversal element budget, mirroring validate's accounting. */
readonly maxElements?: number;
/** Byte cap for one unbounded `nat`/`int` encoding. */
readonly maxNumericBytes?: number;
}| Option | Default | resource reported | What it bounds |
|---|---|---|---|
maxBytes |
10_485_760 (10 MiB) |
bytes |
The input length, checked before any parsing begins. |
maxTypeTableEntries |
100_000 |
type_table_entries |
Decode: the declared type-table size, checked before entries are read. Encode: the distinct composite schema nodes the walk meets, before equal types are merged. |
maxDepth |
256 |
value_depth |
Traversal depth. Every descent, every rec hop, and every skipped nesting level counts; so does every combinator level at which encode's type table gains an entry (Candid nesting depth, not rec hops). |
maxElements |
1_000_000 |
value_elements |
Total elements walked. A blob or vec nat8 charges one per byte plus one for the node. |
maxNumericBytes |
1_048_576 (1 MiB) |
numeric_bytes |
The byte length of one unbounded nat or int encoding. |
When a bound trips, nothing throws and nothing is allocated speculatively. You get
{ ok: false, issues } whose first issue has code
resource_limit_exceeded and a resource_limit field:
export interface CodecResourceLimitInfo {
readonly resource:
"bytes" | "type_table_entries" | "value_depth" | "value_elements" | "numeric_bytes" | "stack";
readonly limit: number;
readonly observed: number;
}
The codec's walks keep their work on explicit stacks, not the JavaScript call stack, so these
limits are the only bounds and the answer never depends on the engine: at the default
maxDepth of 256 a hostile nesting a million levels deep is refused with
value_depth after work proportional to 256, and a host that raises
maxDepth encodes and decodes a 100,000-level linked list or ICRC-3 value the same way
on every call. encode's type table charges maxDepth for Candid nesting
depth only — once per combinator level at which it opens an entry, never for a rec
hop, since an alias adds no depth in Candid either — so any type the candid-core compiler
accepts encodes through its generated module at the default limits, and a hand-built schema
nested deeper than the limit is refused with value_depth at Candid depth 257.
stack is the one resource that is not an option you set. It means the host
JavaScript stack ran out mid-walk, which no depth of message, value or schema causes any more:
what does is your own code that a walk calls — a getter or Proxy trap on a value being encoded,
a rec thunk — recursing too deeply by itself. limit is the call's
maxDepth and observed the depth the walk had reached when the engine
refused. Tell it apart from value_depth by resource, not by comparing
the two numbers.
That triple is what lets a caller tell "this message is malformed" from "this message is legitimate but larger than the budget I set". Budgets are charged per element actually walked, so a five-byte length prefix claiming 2^28 elements cannot cost you 2^28 allocations — the walk stops at your ceiling and reports it.
Wire fields your schema ignores, trailing arguments it does not expect, and the rewound walk
behind an absorbed opt all charge the element and depth budgets. A large but
entirely legitimate blob needs maxElements raised above its byte length: the
suite pins a 2,000,000-byte blob decoding at maxElements: length + 1 and failing
at length.
Decode stops at the first hard error, so its issue list is short — usually one entry. That is a
property of the format, not an omission: once a value is misread the byte cursor cannot be
resynchronized, so there is nothing truthful left to report. validate, which walks
a value structurally, is the API that collects many issues at once.
Strictness where the format is ambiguous#
Non-minimal LEB128 is rejected#
A number encoded with more continuation bytes than it needs — 0x80 0x00 for zero —
fails with overlong_leb128, in type-table counts and in values alike. Be honest
about the standing of this rule: the Candid specification does not state it in so many words.
Rejecting it is this implementation's reading of the format as a strict inverse — one
byte sequence per value, so that decoding cannot accept two spellings of the same number and
re-encoding cannot pick a different one. It was recorded as a deliberate decision, not derived
from the text. The consequence to plan for is local: a hand-rolled encoder that pads its LEB128
output is refused here, whatever another implementation would make of the same bytes.
float32 is exact-or-refuse#
Encoding 0.1 at c.float32 fails with
unrepresentable_float32. It does not store the nearest float32 for you. What that
means as a caller: you either pass a value that survives the conversion, or you round it
yourself and take responsibility for the loss.
import { c } from "@candid-core/schema";
import { encode } from "@candid-core/schema/codec";
encode(c.float32, 0.1); // unrepresentable_float32
encode(c.float32, Math.fround(0.1)); // ok
encode(c.float32, Number.NaN); // ok
The reason is the round trip. With exact-or-refuse, decode(encode(v)) is structural
identity for every value the encoder accepts. With silent rounding it would be identity for most
values and a quiet approximation for the rest, and you would have no way to tell which.
Text and principals#
Candid text is a sequence of Unicode scalar values, so encode refuses lone surrogates
(invalid_text) and decode refuses malformed UTF-8 (invalid_utf8),
overlong forms included.
Principal text is handled inside the package, with no SDK dependency: the canonical form is
lowercase base32 of a CRC-32 checksum concatenated with the id bytes, grouped in fives with
dashes. principalBytesFromText decodes, then re-renders what it decoded and
compares it to the input — so any difference at all returns undefined. Uppercase
text, wrong dash grouping, a bad checksum, non-zero padding bits and an id longer than the
29-byte maximum all fail the same way. There is no lenient parse. Encode applies exactly the
check validate and isPrincipal apply, and refuses with
validate's code, invalid_type, at the same path:
import { c } from "@candid-core/schema";
import { encode } from "@candid-core/schema/codec";
encode(c.principal, "AAAAA-AA"); // invalid_type (case)
encode(c.principal, "aaaaa-ab"); // invalid_type (non-zero padding bits)
encode(c.principal, { toText: () => "aaaaa-aa" }); // invalid_type (not a string)Encode validates as it goes#
encode performs its own complete validation walk rather than trusting a prior
validate call — with one difference: it reads each property exactly once, so a
hostile getter cannot return a valid value to the checker and a different one to the emitter.
The codes and paths it produces match validate's on the same value, which the suite
pins across a matrix of invalid cases.
What the codec does not cover#
Three limits are worth knowing before you point this at arbitrary traffic.
-
Opaque reference values are refused. A
principal,funcorservicevalue whose tag byte is0— the opaque reference form — fails withinvalid_principal, including when the value is merely being skipped over. Skipping is not a validation exemption. -
External reference sequences are refused. After the value section is read,
the input must be exhausted; any bytes remaining fail with
trailing_bytes. -
Wire
futuretypes have no schema counterpart. They are skippable wherever skipping is legal, but a future value at a position the expected schema actually needs fails withtype_mismatch.
func and service values themselves are supported, in the
transparent form: a func value encodes and decodes as
{ principal, method }, and a service value as the principal of the running service.
This module makes no claim about any bytes outside the Candid message: agent envelopes, the request or response wrappers of the Internet Computer's HTTP interface, certificates, and reject encodings are all somebody else's format. Moving these bytes to a canister is the job of whatever call layer sits on top, and the agent it uses never sees a schema.
Issue codes#
CodecCode is a closed union of 24 members: adding one is an API change. Eleven are
shared with validate and mean exactly the same things there —
invalid_type, not_integer, out_of_range,
missing_field, unexpected_field, unknown_tag,
invalid_length, uninhabited_type, unsupported_schema,
unreadable_value and resource_limit_exceeded. The other thirteen are
wire-specific.
| Code | Fires when |
|---|---|
invalid_magic | The input does not start with DIDL. |
malformed_type_table | A type index is out of range, or a primitive appears in the table, or a composite appears inline, or a service method does not denote a func type. |
overlong_leb128 | A LEB128 or SLEB128 number uses a non-minimal encoding. |
truncated | The input ends mid-value. |
trailing_bytes | Bytes remain after the value section. |
type_mismatch | The wire type is not a subtype of the expected schema, with no enclosing opt to absorb it. |
unknown_variant_tag | The wire carries a variant arm the expected variant does not name. |
invalid_tag_byte | A bool, opt, principal, func or service value starts with a byte other than 0 or 1. |
invalid_utf8 | Decoded text or a decoded method name is not well-formed UTF-8. |
invalid_text | A string being encoded contains a lone surrogate. |
invalid_principal | Decoding only: a principal id longer than 29 bytes, or an opaque reference form. Encoding refuses non-canonical principal text with invalid_type, as validate does. |
unrepresentable_float32 | A float32 value is not Math.fround-exact. |
duplicate_field_id | Two keys of one record or variant derive the same wire label id. |
Every issue also carries path, rendered identically to validate's, so
one value produces one path text across the whole runtime. The only addition is the argument
root: an args-level walk reports $args[1].owner where the single-value API reports
$.owner.
Conformance evidence#
The codec is checked against the upstream candid Rust crate as the reference
implementation, and the differential runs in both directions. One test encodes every
vector case with candid and pins the hex into
tests/goldens/wire/<fixture>.wire.json, which the TypeScript decoder must
accept and interpret to a pinned domain value. A second test reads the TypeScript encoder's own
checked-in bytes (<fixture>.wire-ts.json), decodes them with
candid, and asserts they equal the case's value — so the reference implementation
validates our encoder, not only the reverse. The reference type environment is built by
candid_parser reading the fixture .did source directly, deliberately
not through candid-core's own compiler, so the reference path shares no code with the model under
test.
Each vector is also decoded twice: once under the generated golden schema, and once under the
schema schemaFromContract builds from the same fixture's Contract document. Both
must produce the same value and re-encode to byte-identical output. The ledger fixture is in the
set precisely because its generated and loaded schemas share nodes differently; for its vectors
the TypeScript bytes are also exactly the bytes candid writes. A seeded property
test goes further over every golden fixture: the same value, with its keys shuffled, must encode
to the same bytes through the generated schema, the loaded one, and rebuilt copies that duplicate
every node, intern every node, or wrap nodes in extra rec indirections.
Guarantees and verification sets out what each gate actually
establishes and what is deliberately not claimed. A gate existing is not the same as a property
being proven, and the codec is pre-1.0 like everything else here. The module is
crates/candid-core-ts/ts/codec.ts;
its header records each strictness decision and the issue it was decided on, and the vectors
live in
tests/wire_vectors.rs.
Next#
The eleven shared issue codes, path rendering, and the budgets the codec mirrors.
Guide Schemas at runtimeThe same schemas built from a Contract document, ready for encode and decode.
The complete mapping table, including the declarations that are left out.