interscript-ml contract · v1
Specification
This page normatively defines what every secryst crystal implements: the model index, the artifact format, the tokenizer, and the resolution algorithm. The canonical text lives in the contract repo as interscript-ml/SPEC.md; this page renders it for crystal users. The key words MUST, MUST NOT, SHALL, SHOULD, and MAY are to be interpreted as described in RFC 2119.
Conformance
An implementation conforms to interscript-ml v1 when it:
| # | Requirement |
|---|---|
| C1 | MUST resolve model ids against a models.yaml index of version: 1, honoring SECRYST_INDEX. |
| C2 | MUST verify the whole-artifact sha256 before installation and re-verify every cache hit; a mismatch MUST fail loudly. |
| C3 | MUST verify every .onnx member against the manifest sha256 map on load; members not covered by the manifest MUST NOT load. |
| C4 | MUST implement the byte tokenizer exactly (§4); id sequences MUST NOT be treated as raw bytes. |
| C5 | MUST produce byte-identical outputs to the reference crystal on the shared golden sets. |
| C6 | SHOULD install artifacts atomically (temp file + rename) so a partial download never masquerades as a model. |
§1
The models.yaml index
A single YAML document is the catalog. Model ids are stable public identifiers (e.g. khm-latn-1.0).
index schema
version: 1
models:
<id>:
filename: string # artifact file name
url: string # single-file channel (http(s):// or file://)
sha256: string # whole-artifact digest
size: int
precision: fp32 | fp16 # default fp32
task: string
parts: # OPTIONAL: split artifacts
- url: string
sha256: string # per-part digest, verified as it lands
size: int
The parts mechanism exists for artifacts exceeding GitHub’s 2 GiB per-asset cap: parts stream into one file in index order, each verified on arrival; the assembled file is then checked against the entry-level sha256 exactly as a single-file model — the cache contract is identical.
§2
Resolution algorithm
Every crystal resolves an id identically:
- Resolve
<id>inmodels.modelsof the index. Unknown ids MUST raise an error enumerating known ids. - If a cached copy exists at
<cache>/models/<id>/<filename>whose whole-file sha256 matches, use it — cache hits are re-verified, never trusted blindly. - Otherwise download (single URL, or parts in order) to a temporary file in the target directory, verifying digests as data lands.
- Verify the assembled artifact against the index sha256.
- Atomically rename into place, then load.
§3
IMF v1 model zips
The Interscript Model Format v1 is a zip containing at minimum metadata.yaml, encoder.onnx, and decoder.onnx. Zips MUST NOT rely on zip-level integrity; integrity is the manifest’s job.
metadata.yaml
format: imf-v1
tokenizer: bytes
id: khm-latn-1.0
task: transliteration
decoder: plain | kv # kv iff decoder-kv.onnx is present
precision: fp32
opset: 14
sha256: # every .onnx member MUST be covered
encoder.onnx: <hex>
decoder.onnx: <hex>
A conforming loader MUST reject: any other format, any tokenizer other than bytes (this runtime family is byte-level only), missing required members, and uncovered or mismatched member digests.
§4
Byte tokenizer
| Concept | Rule |
|---|---|
| Encoding a string | UTF-8 bytes b → ids b + 3, then one trailing eos |
| Decoding ids | stop at eos; skip pad/unk; (id − 3) mod 256 per byte; reassemble as UTF-8 |
| Special ids | pad = 0, eos = 1, unk = 2 |
Warning — the classic silent-garbage bug: ids are offset by 3 and carry a trailing EOS. Feeding text.bytes directly, or forgetting the EOS, produces plausible-but-wrong outputs that pass shape checks. Interop tests MUST cover both encode and decode round-trips, including multi-byte scripts.
§5
Environment
| Variable | Meaning | Default |
|---|---|---|
SECRYST_INDEX | Index URL or local path | the ml-models models.yaml on GitHub raw |
SECRYST_CACHE | Artifact cache directory | ~/.cache/secryst |