The binary codec

Schema-directed Candid encoding and decoding in TypeScript, fail-closed on every input.

Published as a beta

This page describes the surface the repository builds, published as @candid-core/schema 0.3.0-beta.1 under the npm beta dist-tag. latest is still 0.2.0, which has the old one: a decoded principal is an object with toText() and encoding accepts one, the encoded type table depends on how the schema was built, and a misspelled limit is ignored. Migrating from 0.2.0 lists every difference with before and after code, and each is recorded in the package changelog under 0.3.0-beta.1.

@candid-core/schema/codec turns JavaScript values into Candid wire bytes and back. It is written in TypeScript, has no runtime dependencies, and exports exactly six functions: encode, encodeArgs, decode, decodeArgs, principalTextFromBytes and principalBytesFromText. Everything it needs to know about your data comes from a schema — the plain object a c.* builder returns, the module the generator emits, or a schema built at runtime from a Contract document.

Nothing here throws for a value or a byte string. Every call returns a discriminated result you branch on, and every failure carries a machine-readable code and a $-rooted path into the offending value. That holds for hostile bytes and hostile values alike: a getter that throws mid-encode becomes an unreadable_value issue, not an escaping exception. The one thing that does throw is a malformed options object — see the budgets below.

The four entry points#

A Candid message is an argument sequence, not a single value, so encodeArgs and decodeArgs are the primary pair. encode and decode wrap the one-argument case and re-root the issue paths so they read exactly like validate's.

encodeArgs ./codec
tscrates/candid-core-ts/ts/codec.ts
export function encodeArgs(
  schemas: readonly AnyFieldSchema[],
  values: readonly unknown[],
  options: CodecOptions = {},
): EncodeResult
tscrates/candid-core-ts/ts/codec.ts
export type EncodeResult =
  | { readonly ok: true; readonly bytes: Uint8Array }
  | { readonly ok: false; readonly issues: readonly CodecIssue[] };

One schema and one value per argument. Mismatched lengths fail with invalid_length before anything is written.

decodeArgs ./codec
tscrates/candid-core-ts/ts/codec.ts
export function decodeArgs(
  schemas: readonly AnyFieldSchema[],
  bytes: Uint8Array,
  options: CodecOptions = {},
): DecodeResult
tscrates/candid-core-ts/ts/codec.ts
export type DecodeResult =
  | { readonly ok: true; readonly values: readonly unknown[] }
  | { readonly ok: false; readonly issues: readonly CodecIssue[] };

One decoded value per expected argument. Extra trailing wire arguments are skipped; a trailing argument the wire does not carry follows the record-field rule below.

encode / decode ./codec
tscrates/candid-core-ts/ts/codec.ts
export function encode<T>(
  schema: Schema<T>,
  value: unknown,
  options: CodecOptions = {},
): EncodeResult

export function decode<T>(
  schema: Schema<T>,
  bytes: Uint8Array,
  options: CodecOptions = {},
): { ok: true; value: unknown } | { ok: false; issues: readonly CodecIssue[] }

The one-argument convenience wrappers.

decode gives you unknown, not T

The type parameter on decode<T> constrains the schema, not the payload: a successful result is { ok: true; value: unknown }, and decodeArgs returns readonly unknown[]. The bytes came from outside your program, so the runtime does not hand you a static promise about them. Narrow the value yourself, or pass it through validate — the reference-vector suite pins that each decoded value validates ok against the schema it decoded under.

Here is the whole surface in one round trip:

tscrates/candid-core-ts/ts/README.md
import { c, principal, type Infer } from "@candid-core/schema";
import { validate } from "@candid-core/schema/validate";
import { encode, decode } from "@candid-core/schema/codec";

const Account = c.record({ owner: c.principal, balance: c.nat });
type Account = Infer<typeof Account>; // { owner: Principal; balance: bigint }

const value: Account = { owner: principal("aaaaa-aa"), balance: 5n };

validate(Account, value); // { ok: true } | { ok: false, issues }
const encoded = encode(Account, value); // { ok: true, bytes } | { ok: false, issues }
if (encoded.ok) {
  decode(Account, encoded.bytes);
}

A decoded principal is its canonical text, a plain string typed as the branded Principal — the codec never constructs an SDK class — so a decoded value serializes with JSON.stringify, survives structuredClone, and compares with ===. In the other direction encoding is strict: it accepts exactly canonical principal text, the same strings validate accepts, so an @icp-sdk/core Principal converts once with principal(sdkPrincipal).

What a Candid message looks like#

You do not need the spec to use the codec, but the four parts explain most of its behaviour. A message is:

  1. The magic — the four ASCII bytes DIDL (44 49 44 4c).
  2. A type table — a length-prefixed list of the composite types the message uses. Entries reference each other by index, which is how recursion is expressed, and references may point forward.
  3. An argument type list — one type reference per argument.
  4. The values — all the argument values, concatenated, with no separators.

Only composites get table entries. The eighteen primitives are negative opcodes written inline: null is -1, nat is -3, text is -15, principal is -24. The block -18 to -23 is the composites — opt, vec, record, variant, func, service — and a primitive opcode appearing in the table, or a composite opcode appearing inline, fails with malformed_type_table.

Three shapes have no opcode of their own, so the codec spells them out:

  • blob is emitted as vec nat8;
  • a tuple is a record whose field ids are exactly 0..n-1;
  • unit is the empty record.

Records and variants identify fields by a 32-bit number, not by name. Schema objects carry readable keys, so both directions derive the number by two rules: a key spelled _N_ is the id N, and every other key is the Candid label hash of its UTF-8 bytes, h(name) = fold(h * 223 + byte) mod 2^32. Two keys on one node that derive the same id fail closed with duplicate_field_id rather than one silently winning — c.record({ a: …, _97_: … }) is that collision, because the label hash of a is 97.

Emission order is fixed by the format, not by your object: record fields are read and validated in declaration order — so the first issue's path matches what validate reports — and then the buffered bytes are reordered into ascending wire-id order. Service method names go out sorted by raw UTF-8 byte sequence. Identical inputs therefore produce identical bytes, and nothing depends on time, environment, or map iteration order.

The bytes also do not depend on how the schema objects were built. A generated module mints a fresh c.opt(c.blob()) for every field that spells one; a schema loaded with schemaFromContract shares one node per Contract type; a hand-written schema does whatever its author did. The encoder walks whichever graph it is given, then rewrites the type table into a structural canonical form before writing it: equal Candid types become one entry, numbered in first-visit order from the argument types. So one call is one byte string — and a sound cache or deduplication key — whether its schemas were generated, loaded or hand-built, however they share nodes, and in whatever key order the schema or the value spells a record. blob and vec nat8 are the same wire type and share an entry. A schema with no repeated structure writes exactly the table it always did.

Recursive types are canonical per knot. Every schema generated from or loaded from a Contract has one knot per recursive type, because the Contract canonicalizer has already minimised the graph, so those agree. Two separately hand-built knots for one recursive type — Even and Odd written as two c.rec calls when each is record { next : opt Self } — are not merged: the bytes are larger, valid, and decode to the same value, but they are different bytes. Minimising cyclic graphs in the encoder is a recorded non-goal.

Schema-directed decoding#

Decoding is driven by the schema you expect, not by whatever the bytes claim. The message carries its own type description; the decoder checks that description against your expected schema using Candid's coercion relation, and then reads the value bytes as your schema says to read them. That is the difference between a decoded value whose shape you can rely on and one you have to inspect before you can use it.

It is also what lets an old client keep working against a canister that has since grown. The coercion rules are the spec's, and the asymmetry between options and variants is deliberate:

  • An expected opt absorbs content mismatches to null. If the wire carries text where your opt wanted a nat, you get null. A bare wire value at an expected opt of its type auto-wraps.
  • An unknown variant tag is a hard error (unknown_variant_tag) — unless an enclosing expected opt absorbs it.
  • A missing non-optional record field is a hard error (missing_field); a missing field whose expected type is opt-like decodes as null.
  • Extra wire record fields and extra trailing arguments are skipped — and charged the same budgets as decoded values.
  • A boxed opt follows the same rules. For an opt whose inner type admits null (opt opt T, opt null, opt reserved), a present value decodes as { some: v }. A mismatch the inner opt absorbs is { some: null }, and one the outer opt absorbs is null, exactly as the candid crate decodes them.

Each behaviour below is an executed assertion in the coercion-matrix test (ts/tests/codec.test.ts). Here a small helper unwraps the bytes, and the expected results are in comments:

ts
import { c } from "@candid-core/schema";
import { decode, encode, type EncodeResult } from "@candid-core/schema/codec";

// The bytes of an encode that must have succeeded.
function bytes(result: EncodeResult): Uint8Array {
  if (!result.ok) throw new Error(result.issues[0].message);
  return result.bytes;
}

// A wire nat is accepted at expected int; the reverse hard-errors.
decode(c.int, bytes(encode(c.nat, 5n)));   // { ok: true, value: 5n }
decode(c.nat, bytes(encode(c.int, 5n)));   // type_mismatch

// Extra wire record fields are skipped.
const wide = bytes(encode(c.record({ a: c.bool, b: c.text }), { a: true, b: "x" }));
decode(c.record({ a: c.bool }), wide);            // { ok: true, value: { a: true } }

// A field the wire does not carry: null if opt-like, hard error otherwise.
const narrow = bytes(encode(c.record({ a: c.bool }), { a: true }));
decode(c.record({ a: c.bool, b: c.opt(c.text) }), narrow); // { a: true, b: null }
decode(c.record({ a: c.bool, b: c.text }), narrow);        // missing_field

// Expected opt absorbs a constituent mismatch; a bare value auto-wraps.
decode(c.opt(c.nat), bytes(encode(c.text, "hi")));  // { ok: true, value: null }
decode(c.opt(c.text), bytes(encode(c.text, "hi"))); // { ok: true, value: "hi" }

// Variants trap on an unknown tag — unless an enclosing opt absorbs it.
const later = bytes(encode(c.variant({ ok: c.null, later: c.nat32 }), { tag: "later", value: 7 }));
decode(c.variant({ ok: c.null }), later);         // unknown_variant_tag
decode(c.opt(c.variant({ ok: c.null })), later);  // { ok: true, value: null }

Absorption never hides a broken message. When a constituent mismatch is absorbed, the cursor rewinds and the value is skipped — but malformed bytes and exhausted budgets stay hard failures, and the element charges already spent on the rewound walk are not refunded. Refunding them would let nested options multiply the traversal budget by the absorption depth.

The budgets, and what a caller observes#

CodecOptions is the complete set of bounds, each optional and each defaulting to an exported constant. The same object is accepted by encode and decode.

Options are code, not input, so they are checked strictly and up front. An own key that is not one of the five below throws TypeError — a misspelled maxDeph used to apply the default without a word, and maxIssues belongs to validate, not here. So does a limit that is not a non-negative safe integer: NaN (which used to switch its bound off, since nothing is greater than NaN), a negative, a fraction, a string, null, or Infinity. A trusted host that wants no practical bound passes a large integer. undefined means the default, and 0 is a valid limit that refuses any input consuming that resource. Each option is read exactly once, and the call runs on that checked copy: a getter or Proxy cannot pass the check with one value and run with another. An options object whose getter or trap throws while being read raises a TypeError too, carrying the original exception as its cause.

tscrates/candid-core-ts/ts/codec.ts
export interface CodecOptions {
  /** Input size ceiling for decode, in bytes. */
  readonly maxBytes?: number;
  /**
   * Wire type table entry cap, mirroring `Limits::max_type_nodes`. Decode
   * refuses a wire table claiming more entries. Encode charges one entry per
   * distinct composite schema node its walk meets, before repeated structure
   * is merged, so the table it writes is never larger than this.
   */
  readonly maxTypeTableEntries?: number;
  /**
   * Traversal depth cap, mirroring `Limits::max_value_depth`. Encode's
   * type-table walk charges it too, for Candid nesting depth: once per
   * combinator level at which the table gains an entry, never for a `rec`
   * hop, so a static schema nested deeper than this is refused.
   */
  readonly maxDepth?: number;
  /** Traversal element budget, mirroring validate's accounting. */
  readonly maxElements?: number;
  /** Byte cap for one unbounded `nat`/`int` encoding. */
  readonly maxNumericBytes?: number;
}
OptionDefaultresource reportedWhat it bounds
maxBytes 10_485_760 (10 MiB) bytes The input length, checked before any parsing begins.
maxTypeTableEntries 100_000 type_table_entries Decode: the declared type-table size, checked before entries are read. Encode: the distinct composite schema nodes the walk meets, before equal types are merged.
maxDepth 256 value_depth Traversal depth. Every descent, every rec hop, and every skipped nesting level counts; so does every combinator level at which encode's type table gains an entry (Candid nesting depth, not rec hops).
maxElements 1_000_000 value_elements Total elements walked. A blob or vec nat8 charges one per byte plus one for the node.
maxNumericBytes 1_048_576 (1 MiB) numeric_bytes The byte length of one unbounded nat or int encoding.

When a bound trips, nothing throws and nothing is allocated speculatively. You get { ok: false, issues } whose first issue has code resource_limit_exceeded and a resource_limit field:

tscrates/candid-core-ts/ts/codec.ts
export interface CodecResourceLimitInfo {
  readonly resource:
    "bytes" | "type_table_entries" | "value_depth" | "value_elements" | "numeric_bytes" | "stack";
  readonly limit: number;
  readonly observed: number;
}

The codec's walks keep their work on explicit stacks, not the JavaScript call stack, so these limits are the only bounds and the answer never depends on the engine: at the default maxDepth of 256 a hostile nesting a million levels deep is refused with value_depth after work proportional to 256, and a host that raises maxDepth encodes and decodes a 100,000-level linked list or ICRC-3 value the same way on every call. encode's type table charges maxDepth for Candid nesting depth only — once per combinator level at which it opens an entry, never for a rec hop, since an alias adds no depth in Candid either — so any type the candid-core compiler accepts encodes through its generated module at the default limits, and a hand-built schema nested deeper than the limit is refused with value_depth at Candid depth 257.

stack is the one resource that is not an option you set. It means the host JavaScript stack ran out mid-walk, which no depth of message, value or schema causes any more: what does is your own code that a walk calls — a getter or Proxy trap on a value being encoded, a rec thunk — recursing too deeply by itself. limit is the call's maxDepth and observed the depth the walk had reached when the engine refused. Tell it apart from value_depth by resource, not by comparing the two numbers.

That triple is what lets a caller tell "this message is malformed" from "this message is legitimate but larger than the budget I set". Budgets are charged per element actually walked, so a five-byte length prefix claiming 2^28 elements cannot cost you 2^28 allocations — the walk stops at your ceiling and reports it.

Skipped data still costs budget

Wire fields your schema ignores, trailing arguments it does not expect, and the rewound walk behind an absorbed opt all charge the element and depth budgets. A large but entirely legitimate blob needs maxElements raised above its byte length: the suite pins a 2,000,000-byte blob decoding at maxElements: length + 1 and failing at length.

Decode stops at the first hard error, so its issue list is short — usually one entry. That is a property of the format, not an omission: once a value is misread the byte cursor cannot be resynchronized, so there is nothing truthful left to report. validate, which walks a value structurally, is the API that collects many issues at once.

Strictness where the format is ambiguous#

Non-minimal LEB128 is rejected#

A number encoded with more continuation bytes than it needs — 0x80 0x00 for zero — fails with overlong_leb128, in type-table counts and in values alike. Be honest about the standing of this rule: the Candid specification does not state it in so many words. Rejecting it is this implementation's reading of the format as a strict inverse — one byte sequence per value, so that decoding cannot accept two spellings of the same number and re-encoding cannot pick a different one. It was recorded as a deliberate decision, not derived from the text. The consequence to plan for is local: a hand-rolled encoder that pads its LEB128 output is refused here, whatever another implementation would make of the same bytes.

float32 is exact-or-refuse#

Encoding 0.1 at c.float32 fails with unrepresentable_float32. It does not store the nearest float32 for you. What that means as a caller: you either pass a value that survives the conversion, or you round it yourself and take responsibility for the loss.

ts
import { c } from "@candid-core/schema";
import { encode } from "@candid-core/schema/codec";

encode(c.float32, 0.1);               // unrepresentable_float32
encode(c.float32, Math.fround(0.1));  // ok
encode(c.float32, Number.NaN);        // ok

The reason is the round trip. With exact-or-refuse, decode(encode(v)) is structural identity for every value the encoder accepts. With silent rounding it would be identity for most values and a quiet approximation for the rest, and you would have no way to tell which.

Text and principals#

Candid text is a sequence of Unicode scalar values, so encode refuses lone surrogates (invalid_text) and decode refuses malformed UTF-8 (invalid_utf8), overlong forms included.

Principal text is handled inside the package, with no SDK dependency: the canonical form is lowercase base32 of a CRC-32 checksum concatenated with the id bytes, grouped in fives with dashes. principalBytesFromText decodes, then re-renders what it decoded and compares it to the input — so any difference at all returns undefined. Uppercase text, wrong dash grouping, a bad checksum, non-zero padding bits and an id longer than the 29-byte maximum all fail the same way. There is no lenient parse. Encode applies exactly the check validate and isPrincipal apply, and refuses with validate's code, invalid_type, at the same path:

ts
import { c } from "@candid-core/schema";
import { encode } from "@candid-core/schema/codec";

encode(c.principal, "AAAAA-AA"); // invalid_type (case)
encode(c.principal, "aaaaa-ab"); // invalid_type (non-zero padding bits)
encode(c.principal, { toText: () => "aaaaa-aa" }); // invalid_type (not a string)

Encode validates as it goes#

encode performs its own complete validation walk rather than trusting a prior validate call — with one difference: it reads each property exactly once, so a hostile getter cannot return a valid value to the checker and a different one to the emitter. The codes and paths it produces match validate's on the same value, which the suite pins across a matrix of invalid cases.

What the codec does not cover#

Three limits are worth knowing before you point this at arbitrary traffic.

  • Opaque reference values are refused. A principal, func or service value whose tag byte is 0 — the opaque reference form — fails with invalid_principal, including when the value is merely being skipped over. Skipping is not a validation exemption.
  • External reference sequences are refused. After the value section is read, the input must be exhausted; any bytes remaining fail with trailing_bytes.
  • Wire future types have no schema counterpart. They are skippable wherever skipping is legal, but a future value at a position the expected schema actually needs fails with type_mismatch.

func and service values themselves are supported, in the transparent form: a func value encodes and decodes as { principal, method }, and a service value as the principal of the running service.

The scope is the DIDL message

This module makes no claim about any bytes outside the Candid message: agent envelopes, the request or response wrappers of the Internet Computer's HTTP interface, certificates, and reject encodings are all somebody else's format. Moving these bytes to a canister is the job of whatever call layer sits on top, and the agent it uses never sees a schema.

Issue codes#

CodecCode is a closed union of 24 members: adding one is an API change. Eleven are shared with validate and mean exactly the same things there — invalid_type, not_integer, out_of_range, missing_field, unexpected_field, unknown_tag, invalid_length, uninhabited_type, unsupported_schema, unreadable_value and resource_limit_exceeded. The other thirteen are wire-specific.

CodeFires when
invalid_magicThe input does not start with DIDL.
malformed_type_tableA type index is out of range, or a primitive appears in the table, or a composite appears inline, or a service method does not denote a func type.
overlong_leb128A LEB128 or SLEB128 number uses a non-minimal encoding.
truncatedThe input ends mid-value.
trailing_bytesBytes remain after the value section.
type_mismatchThe wire type is not a subtype of the expected schema, with no enclosing opt to absorb it.
unknown_variant_tagThe wire carries a variant arm the expected variant does not name.
invalid_tag_byteA bool, opt, principal, func or service value starts with a byte other than 0 or 1.
invalid_utf8Decoded text or a decoded method name is not well-formed UTF-8.
invalid_textA string being encoded contains a lone surrogate.
invalid_principalDecoding only: a principal id longer than 29 bytes, or an opaque reference form. Encoding refuses non-canonical principal text with invalid_type, as validate does.
unrepresentable_float32A float32 value is not Math.fround-exact.
duplicate_field_idTwo keys of one record or variant derive the same wire label id.

Every issue also carries path, rendered identically to validate's, so one value produces one path text across the whole runtime. The only addition is the argument root: an args-level walk reports $args[1].owner where the single-value API reports $.owner.

Conformance evidence#

The codec is checked against the upstream candid Rust crate as the reference implementation, and the differential runs in both directions. One test encodes every vector case with candid and pins the hex into tests/goldens/wire/<fixture>.wire.json, which the TypeScript decoder must accept and interpret to a pinned domain value. A second test reads the TypeScript encoder's own checked-in bytes (<fixture>.wire-ts.json), decodes them with candid, and asserts they equal the case's value — so the reference implementation validates our encoder, not only the reverse. The reference type environment is built by candid_parser reading the fixture .did source directly, deliberately not through candid-core's own compiler, so the reference path shares no code with the model under test.

Each vector is also decoded twice: once under the generated golden schema, and once under the schema schemaFromContract builds from the same fixture's Contract document. Both must produce the same value and re-encode to byte-identical output. The ledger fixture is in the set precisely because its generated and loaded schemas share nodes differently; for its vectors the TypeScript bytes are also exactly the bytes candid writes. A seeded property test goes further over every golden fixture: the same value, with its keys shuffled, must encode to the same bytes through the generated schema, the loaded one, and rebuilt copies that duplicate every node, intern every node, or wrap nodes in extra rec indirections.

Guarantees and verification sets out what each gate actually establishes and what is deliberately not claimed. A gate existing is not the same as a property being proven, and the codec is pre-1.0 like everything else here. The module is crates/candid-core-ts/ts/codec.ts; its header records each strictness decision and the issue it was decided on, and the vectors live in tests/wire_vectors.rs.

Next#