Spec Workflow
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
How to change spec/ and leave the repository consistent. For what the fields
MEAN, read Spec System first; this page is
the procedure.
Every command here is written out. If a step here disagrees with what the tools do, the tools are right and this page is a bug.
Before and after any spec change
just spec-status # what state the spec system is in, derived from the gates
Run it before you start, so you know what “unchanged” looks like, and again at the end. A change that moves the “deferred” or “failing” counts in the wrong direction is worth a second look.
Before treating deferred examples as implementation work, inspect their live
claim review with cargo run --manifest-path spec/Cargo.toml --bin spec_status -- --deferred.
Legal controls and subsumption claims remain regression obligations, not a request
to recreate a diagnostic that no longer applies. Resolve contradicted claims from policy and
source evidence; do not flip registry status merely to reduce the deferred count.
Adding a construct spec
A construct spec is a VALID fragment plus the tree it must parse to.
1. Write the file under the right spec/constructs/ subdirectory
(header/, main_tier/, tiers/, utterance/, word/):
# my_example
Description of what this example demonstrates.
## Input
```utterance
*CHI: hello world .
```
## Expected CST
```cst
(utterance
(main_tier
...))
```
## Metadata
- **Level**: utterance
- **Category**: main_tier
The fence label (utterance here) names a template in spec/tools/templates/
that wraps the fragment into a full CHAT file. If no template matches, create
one; the generator fails rather than guessing.
2. Get the real CST rather than writing one by hand:
cd grammar && tree-sitter parse <a file containing your input>
Copy the tree, dropping byte positions and field names.
3. Regenerate and verify (see “Regenerating” below).
Adding an error spec
An error spec is INVALID CHAT plus the codes it must produce.
1. Write the file in spec/errors/, named E###_<slug>.md. Everything
declared goes in +++ TOML frontmatter; the prose goes in the body.
+++
code = 'E301'
name = 'Empty speaker code'
[[example]]
source = 'E3xx_main_tier_errors/E301_empty_speaker.cha'
level = 'utterance'
claim = 'violates'
chat = '''
@UTF8
@Begin
@Languages: eng
@Participants: CHI Target_Child
@ID: eng|corpus|CHI|||||Target_Child|||
*: hello .
@End
'''
+++
## Description
Empty speaker code.
A misspelled or unrecognised key is a LOAD ERROR, so you find out from
just spec-check rather than from a field that silently did nothing.
Four things decide whether your spec asserts anything, and each is easy to get wrong. They are covered in full in Spec System; in short:
claimis the field that asserts, and it is REQUIRED.violates(the spec’s code must appear),legal(it must not), orsubsumed_by <code(s)>(the targets appear and the spec’s code does not). Extra emitted codes still pass; the exact per-stage sets are the snapshot’s business.- There is no
layerfield. Which stage catches a rule is observed, not declared: every example is a fixture whose runner checks both stages, and the per-stage record lives in the observation snapshot. statusandkindare NOT yours to declare. They are facts about the CODE, and they live inspec/codes/error-codes.toml, one entry per code. A spec naming a code that file does not declare does not load, andstatus = 'not_implemented'THERE still defers every example of that code and#[ignore]s its generated tests. Writing either key in a spec file is a load error naming the key. The separate deferred-spec regression check still verifies explicitlegalandsubsumed_byclaims against the live parser and validator. Deferring implementation of a code does not excuse a stale claim about accepted input or the alternative diagnostic that rejects it. Plannedviolatesexamples remain deferred; satisfying a subsumption claim does not implement the deferred code or change its registry status. There is no default forstatus: an invented answer to “is this rule live” is the kind of wrong value nothing notices, and a per-file copy of a per-code fact is the kind that several files could disagree about.source’s stem names the transcript, which is what rules about the file’s own name (E531) compare against.
Write the failing case first. A new error spec should fail before the rule exists; that is what proves the fixture actually triggers it.
A brand-new code needs one bootstrap step. The ErrorCode enum is
generated by spec_gen, and spec_gen (like schema-gen, the first step of
just regen) links against talkbank-model. So a validation rule that names
the new variant cannot compile until the enum exists, and neither command can
produce the enum while the rule is in place. Add the registry entry and the
spec first, run just spec-gen with the rule not yet referencing the variant,
then write the rule and run just regen in full; the second run rebuilds the
observation snapshot against the rule’s real behaviour.
Regenerating
Mutation candidates are not golden tests
For source-bound main-tier terminator deletions, use:
just spec-mutation-candidates corpus/reference/content/terminators-standard.cha
This runtime command admits one seed only after diagnostic-free parsing and
alignment-aware validation, including rules about its actual transcript stem.
Warnings also refuse admission. Add --strict-linkers to select the optional
cross-utterance rules; admission does not claim validity under unselected rules.
It produces one candidate per present typed main-tier terminator, verifies that
the model’s span matches the complete token, and preserves all other bytes.
Missing optional terminators produce no candidate. It never rewrites the model
or guesses a token from line endings.
Output is JSON on stdout, containing the exact seed, transcript identity, selected linker policy, and each candidate’s deleted byte range, removed text, and resulting CHAT. Every candidate is explicitly unreviewed, with no expected diagnostic. No output files are created by the command. Seed or span rejection occurs before candidate output; stdout I/O can still fail partway through a write. Retain the seed record when saving candidates so provenance does not depend on an input path continuing to hold the same bytes.
The API lives in spec-runtime-tools::mutation, where live parser/model
dependencies belong. AdmittedSeed privately owns its model and borrows its
immutable source; candidates can only be constructed from that model’s checked
spans. This is AST-source association, not an assertion that the generated CST
API has already migrated every mutation family. Only terminator deletion is
implemented here; the legacy mutators below are not migrated.
just spec-perturb produces unreviewed candidates, not diagnostic oracles.
Its JSON records have assessment: "unreviewed" and no expected_error field.
The removed diagnostic table had drifted from canonical rules; a mutation’s
name cannot establish which diagnostics a particular input should produce.
The legacy text mutators do not admit a valid seed through the parser/model. They can affect continuation lines, normalize final newlines, select a speaker code already declared by a seed, or introduce multiple defects. Do not use their outputs directly as goldens or treat a syntax-error census as validity. Before running them, review output-name collisions and the intended input scope.
Candidate output has an explicit planning phase: duplicate destinations
and existing output paths are refused before candidate writes begin. Execution
uses exclusive file creation as well, so a file created after preflight cannot
be overwritten. This protects output, not seed validity. A mid-run I/O failure
may leave earlier newly created candidates; preserve and review them rather
than assuming the batch was atomic or automatically deleting them. The legacy
--mine summary path is separate and is not covered by this candidate contract.
For promotion, retain the valid control, describe the precise mutation and its
source relationship, justify a claim from the specification or a recorded
ruling, then add the reviewed example to spec/errors/ and regenerate through
the normal owner. Review observations separately from expectations. A changed
parser output is evidence to adjudicate, not permission to rewrite the claim.
Generate owned artifacts
One command, from anywhere in the checkout:
just spec-gen # rewrite every generated artifact from the specs
just spec-check # or ask whether the committed copies are current
It regenerates every artifact in the registry, in dependency order (the
observation snapshot first, since the tree-sitter corpus derives its
membership from it); the generated artifact
table included in the
spec-system chapter is the live list. There is nothing to choose and no path to type:
every destination is a constant in spec/tools/src/artifacts.rs, so a
generator cannot be aimed at the wrong tree.
just spec-check writes nothing and is exactly what the
every_generated_artifact_is_current gate runs, so a green check means a green
gate.
The published error-reference pages under docs/errors/ are part of
spec-gen like every other artifact, and spec-check gates them.
Never hand-edit anything under a generated/ directory. An artifact that owns
its directory wipes it wholesale and refuses to clear one lacking its
.generated-output-dir marker.
Verifying
just spec-status # the derived summary
cargo test --manifest-path spec/Cargo.toml --workspace # every spec-side gate
just test # the main workspace
If your change touched the grammar, follow the full
Grammar Workflow as well: a
grammar.js edit needs tree-sitter generate before any parser behaviour can
be trusted.
Updating a registry
Two closed vocabularies live under spec/, each generating every site that
names it. Neither is edited anywhere but its registry.
just symbols-gen # spec/symbols/symbol_registry.json
just form-markers-gen # spec/form_markers/form_marker_registry.json
flowchart TD
registry["Edit the registry JSON"]
gen["Run its generator\n(loading validates; there is no separate check step)"]
fmt["Generator runs rustfmt on Rust output"]
gate["Drift gate compares committed output\nagainst what the generator produces"]
registry --> gen --> fmt --> gate
The generators format their own Rust output deliberately: otherwise just fmt
and the generator each rewrite the same bytes and the drift gate fails forever,
with both sides correct.
Each registry’s README covers its authorities and the follow-ups its generator
cannot do:
spec/symbols/README.md,
spec/form_markers/README.md.
Common mistakes
- Editing generated files. Change the spec or the registry, then regenerate.
- Wishing for an example that asserts nothing. There is no such state:
claimis required, and an example that cannot honestly sayviolatessayssubsumed_by(the worklist) orlegal(the boundary). - Flipping
statustoimplementedwithout regenerating. The fixture manifest still carries the old status, so the runner keeps skipping what you just enabled. - Regenerating reflexively. Regeneration is for artifacts that genuinely changed, not a substitute for deciding what the change needs.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).