Introduction
Status: Current Book last changed: 2026-10-07 (commit 5e895791) This page last changed: 2026-09-01 (commit b1db0410)
TalkBank is the world’s largest open repository of spoken language data. This repository (TalkBank/chatter) is the standalone home of the CHAT format authority and the chatter tool family: the chatter CLI, the Rust crates for parsing/validation/transformation, the tree-sitter-talkbank grammar, the talkbank-lsp language server, and the desktop validation app.
chatter is publicly released. To get it right away:
- Command-line tool (macOS / Linux):
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/TalkBank/chatter/releases/latest/download/chatter-installer.sh | sh(Windows and other options: Install). - Desktop app: download for your platform from the latest release.
- Full installation guide (all platforms, package details): Install.
The Rust crates are source-available from this repository (not yet published to crates.io). As a 0.x release, APIs and flags may change before 1.0.
Choose the right surface
| Task | Recommended Surface |
|---|---|
| CHAT validation, normalization, or conversion | chatter CLI |
| LSP integration in editors | talkbank-lsp standalone |
| Build CHAT tooling in Rust | Rust crates (talkbank-model, talkbank-parser, etc.) |
| Reuse grammar in other tools | tree-sitter-talkbank |
| Standalone desktop GUI for CHAT validation | Chatter Desktop (apps/chatter-desktop/) |
What’s In This Repo
chatterCLI: validate, convert, normalize, and analyze CHAT files from the command line, with an interactive TUI for corpus-scale workflows- Language Server (LSP): works with any LSP-compatible editor (Neovim, Emacs, Helix, Zed, etc.) to provide live validation and cross-tier alignment
- JSON data model: every CHAT structure as typed JSON with lossless roundtrip fidelity, backed by a published JSON Schema
- Rust API: parse, validate, inspect, and transform CHAT files programmatically via library crates
Who This Book Is For
| Audience | Start Here | Then Go To |
|---|---|---|
| CLI users validating, normalizing, or converting CHAT | Install | chatter Quick Start, CLI Reference |
| Rust library consumers parsing or transforming CHAT | Library Usage | crate-root rustdoc for talkbank-model, talkbank-parser, and talkbank-transform |
| Grammar / format consumers embedding CHAT parsing in other tools | CHAT Format Overview | tree-sitter-talkbank docs and the grammar/reference chapters |
| Contributors / maintainers working in this repo | Contributing setup | CI and release |
Repository Layout
grammar/ Tree-sitter grammar for CHAT
spec/ Source of truth: CHAT specification + error specs
crates/ Rust crates for model, parser, transform, cache, CLI, LSP, tests, and FFI support
apps/ Tauri v2 desktop app (`chatter-desktop`)
corpus/ Reference corpus (must stay 100% valid under the regression gate)
schema/ JSON Schema for the CHAT AST
tests/ Integration tests and fixtures
docs/ Strategy docs, proposals, and investigations for this repo
book/ This documentation (mdBook)
Data flows: spec (source of truth) → grammar (tree-sitter) → Rust crates (parsers, model, validation, CLI, LSP) → applications (chatter, desktop app).
This page last changed: 2026-09-01 (commit b1db0410). The whole book last changed: 2026-10-07 (commit 5e895791).
Install
Status: Current Last modified: 2026-09-28 08:30 EDT
Everything here comes from the latest release.
Just want to check CHAT files, and you are not a programmer? Get the desktop app below. You never need a terminal.
Chatter desktop app (recommended for most people)
The Chatter app checks CHAT transcripts in an ordinary window: open a file, see the problems highlighted, fix them, and re-check. No terminal and no setup, and it updates itself when a new version comes out.
macOS
The Mac app is signed and notarized by Apple, so it opens normally (no security warnings).
-
Download Chatter for your Mac:
- Apple Silicon (M1/M2/M3/M4, essentially every Mac sold since late 2020): Download Chatter for Apple Silicon.
- Intel (older models): Download Chatter for Intel Mac.
Not sure which you have? Apple menu () then About This Mac: if it says “Apple M…”, it is Apple Silicon.
-
Open the downloaded
.dmgfile. -
Drag Chatter onto the Applications folder in the window that appears.
-
Open Chatter from your Applications folder (or Launchpad).
Windows (Intel/AMD 64-bit, “x64”)
Download Chatter for Windows and run the installer. Windows binaries are not code-signed yet, so SmartScreen may warn on first run: choose More info, then Run anyway.
Linux (Intel/AMD 64-bit, “x86_64”)
Download Chatter
(AppImage)
(make it executable, then run it) or the .deb
package
(install with your package manager).
(The desktop app is x86_64-only on Windows and Linux today; macOS has both
Apple Silicon and Intel builds. The chatter command-line tool below also ships
a Linux ARM build.)
chatter, the command-line tool (for programmers and automation)
If you are comfortable in a terminal, the chatter CLI validates, normalizes,
converts (JSON), watches, and batch-processes CHAT files, and is
the right tool for scripting and CI.
macOS / Linux:
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/TalkBank/chatter/releases/latest/download/chatter-installer.sh | sh
Windows (PowerShell):
irm https://github.com/TalkBank/chatter/releases/latest/download/chatter-installer.ps1 | iex
Then run chatter --help. Full reference: CLI
installation and CLI
Reference. chatter self-updates with
chatter update.
talkbank-lsp language server (editor integration)
For live CHAT validation, hover, go-to-definition, and cross-tier alignment inside an LSP-aware editor (Neovim, Emacs, Helix, Zed, VS Code, and others), install the standalone, code-signed language server:
macOS / Linux:
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/TalkBank/chatter/releases/latest/download/talkbank-lsp-installer.sh | sh
Windows (PowerShell):
irm https://github.com/TalkBank/chatter/releases/latest/download/talkbank-lsp-installer.ps1 | iex
Or download the per-platform archive (talkbank-lsp-<target>.tar.xz, or .zip
on Windows) from the release and point your editor’s LSP client at the
talkbank-lsp binary (it speaks LSP over stdio on .cha files, language id
chat).
Editor quick fixes do not supply missing participant names, roles or languages. Enter those facts yourself: an undeclared speaker or missing/empty header does not establish the right value. E308, E504 and E507 therefore offer no guessed edit, consistent with the CLI fix policy. Diagnostics remain visible, including parser recovery diagnostics when a malformed header cannot be lowered. Other supported quick fixes remain available.
Rust crates and the grammar (embed in your own program)
To embed CHAT parsing / validation / transformation in your own program, depend
on the talkbank-* crates and the tree-sitter-talkbank grammar. They are
source-available from this repository (not yet published to crates.io). See
Library usage and the CHAT format
overview.
As a 0.x release, APIs and flags may change before 1.0; see the Release
Notes. For audio + ML pipelines (transcribe, force-align,
morphotag), see the upstream batchalign3 project, which has its own
installation flow.
This page last changed: 2026-09-28 (commit 2cb42a45). The whole book last changed: 2026-10-07 (commit 5e895791).
Quickstart
Status: Current Last modified: 2026-06-21 21:33 EDT
Task-driven entry points. Pick the row that matches what you want to do today; each path starts at the narrowest useful documentation surface instead of dropping you into the whole book.
| Today’s goal | Best first page | Surface |
|---|---|---|
| Validate / normalize / convert existing CHAT | chatter Quick Start | CLI |
| Add CHAT parsing/validation to a Rust program | Library Usage | Rust crates |
To download and install chatter, see Install.
For audio + ML workflows (transcribe / align media → CHAT), see the
upstream batchalign3 project, outside the chatter repo.
This page last changed: 2026-06-21 (commit a05df6e4). The whole book last changed: 2026-10-07 (commit 5e895791).
Changelog
All notable changes to this project are documented in this file.
The format is based on Keep a Changelog, and the project follows Semantic Versioning. Before 1.0, breaking changes to the CLI or library APIs bump the minor version and are listed under “Changed” / “Removed”.
Unreleased
0.29.0 - 2026-10-07
This release lets a program rewrite part of a transcript (regenerate
%mor/%gra or %wor, or split an utterance) while everything it keeps is
still checked against the source it came from. Each of these operations parses
once, never reparses its own output, and never certifies more than it checked.
One validation rule changes: a %wor word bullet that ends before it starts is
E362. The typed CST traversal is regenerated, which is breaking for code that
names its types.
Added
Dependent-tier replacement (talkbank-parser)
-
TreeSitterParser::admit_replacing_tiersparses once, removes the selected dependent tiers (ReplacementTiers::Morphosyntax:%morwith%gra;ReplacementTiers::WordTiming:%wor) before lowering, and validates every retained domain with the normal rules and tier-alignment checks. Headers, main tiers, retained dependent tiers and unclassified parser recovery still refuse. The result,AdmittedReplacement, retains the producing parse and aRemovedTierreceipt (tier, span, syntax recovery) for every concrete tier it removed, without declaring the original source valid. Refusals areReplacementFailure(Internal,SourceUnavailable,RetainedParse,Validation); an internal tool failure is never reported as invalid CHAT. -
TreeSitterParser::admit_planned_tiersdecides the removal from the document’s own typed headers (all of them, misplaced ones included) within the same parse; headers and utterances each lower once. Selecting nothing preserves and validates every tier.AdmittedDispositiondistinguishesPreserved(anAdmittedPreservation, which carries the original bytes together with theValidChatFileadmitted from them),ReplacedandRegenerating. -
TreeSitterParser::admit_word_timing_planchooses, per document, betweenWordTimingPlan::PreserveandWordTimingPlan::PreferRetained. UnderPreferRetained, valid word timing is kept, partial timing included. When complete admission fails, only the%wortiers the evidence names are removed (a tier whose own lowering fails, or that holds a validation error inside its source span, such as a reversed word bullet or a duplicate%wor), and the reduced document is validated again; errors located in no word tier remove every remaining word tier, and retained faults still refuse.RemovedTier::causereturns theRemovalCause(Selected,OwnLowering,LocatedValidation,UnattributedValidation) with the diagnostics that cost the tier. When removing recorded word timing leaves linked media without timing, the result isWordTimingAdmission::Regenerating, anAdmittedTimingRegenerationthat holds aPendingTimingChatFileand can never become aValidChatFile.
Timing regeneration (talkbank-model)
-
ChatFile::validate_for_timing_regenerationadmits retained structure for an operation that will produce timing. It returnsTimingRegenerationAdmission::Readywhen every requirement is already met, orPendingwith aPendingTimingChatFilewhoseMediaTimingObligationnames the linked-media declaration (E544) still awaiting timing. A pending file cannot be converted to aValidChatFile.PendingTimingChatFile::document_mutwrites regenerated timing into the document without separating it from its obligation, andPendingTimingChatFile::dischargechecks that document with the rule that issued the obligation: it hands the document back once E544 no longer fires, and the payload back while it would. Complete admission of the returned document still runs every rule. -
ChatFile::timing_evidencereturnsTranscriptTimingEvidence:Absent, orRecordedwith aRecordedTranscriptTimingwitness bound to the actual bullet and its document. E544 asks the same question through the same code. Presence of a bullet does not certify that the timing is valid. -
ValidationFailure::into_rejectionreturns the rejected document and the diagnostics that rejected it, both by move, or hands an internal tool failure back unchanged, since that is no evidence about the document. -
InternalFailure::tool_faultrecords a fault a tool detected in itself as one internal-error diagnostic.
Source-bound admission and utterance splitting (talkbank-transform)
-
parse_source_with_parserparses once and returnsParsedSourceChat, which keeps the source bytes and the parse product together so a policy can inspect the parsed model before deciding.ParsedSourceChat::admitadmits that same parse, tier alignment included and without reparsing, into anAdmittedSourceChat: the original bytes with theValidChatFileadmitted from them. No constructor pairs separately supplied text and model;AdmittedSourceChat::from_preservationtakes the parser’sAdmittedPreservation. -
utterance_split::UtteranceSplitPlanadmits a child assignment against the very utterance it will rebuild, in the morphology/extraction word domain (for_morphology) or the%worword domain (for_word_timing). It refuses (SplitRefusal) a slot-count mismatch, a child that reappears after another, a boundary inside one indivisible content item (such as a multi-word replacement), and a boundary that would strand a separator at the start of a child.executereturnsSplitOutcome::Unchangedwhen nothing splits, so the caller keeps the utterance it already holds and nothing is copied, orSplitOutcome::Split(SplitChildren): the rebuilt children, with anInvalidatedTierreceipt (TierInvalidationReason) for every dependent tier that could not be carried over. Children are not validity proof: a document holding them must still pass complete admission before it is written. -
utterance_split::WordSpeakerSplitPlansplits an utterance by the measured speaker ownership of each word. A diarization timeline is projected through complete, count-matched, lexically corroborated%wortiming, with the same held-time policy as whole-turn rediarization; a tie or an uncovered word refuses (WordSpeakerSplitRefusal) instead of inventing a track.WordSpeakerSource::admitadmits the source timing before any acoustic inference is run, andWordSpeakerSource::bind_timelinelater binds the timeline to that same utterance without rematching it. The outcome’sWordSpeakerPartitionisRelabeled { speaker }when one track owns every word, orSplitwith relabeled children and their tier-loss receipts;ownershipreports every word’s measured distribution. Track labels are not claims of personal identity; participants and IDs remain the caller’s responsibility. -
utterance_split::build_word_to_content_mapmaps each extracted word to the main-tier content item it came from, andextract::count_utterance_contentcounts whatcollect_utterance_contentwould extract, through the same walk, without building the words.
Changed
-
Breaking:
talkbank_parser::generated_traversalis regenerated from tree-sitter-grammar-utils’ redesigned carriers. Each carrier and choice is declared once, generic over a range phaseR(defaultRaw) and, when it reaches a slot the compiled grammar proves is never MISSING, a kind proofK(defaultBroad) before it, soX<'tree>still names the readingextract_<rule>returns.AdmittedXis now an alias of theKindAdmittedreading, emitted only for a shape with the kind axis; for any other shape, and for everyAdmittedXBoundView, writeX/XBoundView, which is the same type under both proofs. A narrowed slot is aNarrowedKindSlotwhoseMissingpayload isNarrowedMissing<K, KindMissing<T>>: the placeholder underBroad, uninhabited underKindAdmitted, so a by-value match on an admitted slot omits the arm. Code generic over a kind slot bounds the payload by what it reads (KindPlaceholder<T>,RecoveryNode,SourceRecovery) rather than namingKindMissing<T>. A placeholder at a narrowed slot under the admitted proof is the newReconstructionFault::ContradictedKindProof. Parsing, validation and serialization are unchanged. -
The runtime that
generated_traversalembeds no longer contributes its examples astalkbank-parserdoctests: they are written against tree-sitter-grammar-utils’ own crate path, so in a consumer one failed to compile and thecompile_failones passed only because their import did not resolve. They are nowignoreexamples; the generator’s own crate still runs them. -
A
%worword bullet that ends before it starts is now E362, reported at the bullet’s own source span. A transcript containing one, which validated before, is now rejected. Zero-duration, overlapping and out-of-order word bullets remain legal (CLAN CHECK checks no word bullets; chatter rejects only the self-contradictory case); timing consumers refuse zero-duration word intervals at their own boundary. -
The free-text dependent tiers (the nine bullet-payload tiers, the fifteen raw-text tiers and
%xtiers) read their body under the admitted kind proof. The body is a nonterminal tree-sitter never inserts as a MISSING placeholder, so the reader no longer has a branch that parsed a placeholder’s empty text as content. -
Dependency updates: comrak 0.56, cc 1.6 and insta 1.49; for the desktop application’s tooling, WebdriverIO 10,
@vitejs/plugin-react6.1.2 and Vite 8.3.3. The desktop’s “reveal in file manager” action goes through its typed desktop capability instead of a component-local Tauri import.
0.28.0 - 2026-10-02
This release is about telling the truth about a validation run: how it ended, what it read, what it did with the cache, and what each output promises. Most of it is breaking for library callers and for scripts that relied on a command silently ignoring something; each such change is marked.
Changed
-
Closed-stdout subprocess tests construct a pipe with its reader already closed before launching the child. A consuming endpoint capability removes the parent/child scheduling race without sleeps, retries or weaker assertions.
-
Cache-write test doubles retain
ResolvedPathinstead of erasing it to a raw path, so unresolved spellings cannot construct expected cache identities. Path-resolution contracts explicitly follow native Windows and POSIX rules for..through missing directories; no platform’s semantics are emulated. -
Workspace tests and native-platform CI collect all failing test binaries with
--no-fail-fast. Failures still fail the command; no tests are retried or ignored to obtain a passing result. -
Cache handles offer
close(self), which consumes the query capability and waits for pooled connections and SQLite workers to shut down. Callers that remove an owned database can close every owned handle explicitly instead of relying on background destructor timing. Cache lifecycle tests close both writer and inspector before deletion, then obtainAbsentby fresh inspection; they no longer query an unlinked database or assume it can be removed while open on Windows. Non-current schema inspection, ledger-read errors, failed migration and refused validation-scope admission also await shutdown of unretained pools. Observations without a query handle no longer leave inspection workers holding a database file. -
Breaking (Rust API):
RunEnding::Completerequires producer-admittedCompleteStats;StoppedandIncompleterequirePartialStats. Read counts throughsnapshot()and shortfalls throughmissing_files(); the separateunprocessedandlost_filesfields are removed. Cloning stopped-run counts cannot promote them to complete-run success. CLI JSON and desktop event formats are unchanged. -
Documentation publication metadata follows its own Git history: rendered book headers use the existing Git-date preprocessor, and other affected documents link to their own history. Content-preserving release squashes do not require fresh content-review dates. The date check admits the header itself, so a placeholder example in the body or a history link to another document cannot mask a stale handwritten date. The existing stale-document baseline is unchanged.
How a validation run ends
- Breaking (CLI): a
chatter validaterun that did not cover every file fails the command (exit 1) and says so. A run told to stop (Ctrl-C, or--max-errors) reports how many files it left,Stopped after reaching the error limit (N); M file(s) were not validated.(orValidation cancelled; ...), and a stop that came after the last file is no stop at all: the run is complete. A run that found no transcript saysError: no .cha files found in PATHand exits 1. - Breaking (CLI): the TUI’s exit status is the run’s, as the other
outputs’ are: closing it (
qor Esc) exits 0 only after a complete run with no invalid, unreadable or tool-failed file, and 1 otherwise, including when it is closed before the run ends or the terminal fails. Ctrl-C stops the run and leaves the TUI open on its ending; a second Ctrl-C force-quits at once with exit 130, as it always did, whatever the run found. It lists every file that failed, a file it could not read, a failed roundtrip and a tool failure among them, says “No .cha files found” for an input with none, and shows the--suppressand cache notes the other outputs print. An unreadable argument used to show “no errors found” and exit 0. - Breaking (CLI):
--max-errors Ncounts errors only, never warnings, and the validation runner enforces it: each worker counts its file’s errors before taking the next file, so with--jobs 1the stop is exact, and the text, JSON, audit and TUI surfaces all honour it. The CLI used to count every rendered diagnostic, so a warnings-only file could spend the limit and the run still exit 0 with files never validated; the TUI ignored the limit.--max-errors 0is a usage error (exit 2). - A run that lost files says why: workers that panicked, workers that could not create their parser, or a worker thread the system refused, every one observed and none hidden behind another, or no explanation (a validator defect). A worker that could not create its parser used to return an empty tally, so a run in which every worker failed that way read as a stop, or as a loss with nothing to explain it. Text, audit and TUI endings and the desktop’s incomplete state say the cause.
- Breaking (CLI JSON):
validate --format jsonreports the ending on stdout: the summary’scancelledis replaced byoutcome("complete","stopped"or"incomplete"); a stopped run emits astoprecord (reason"max_errors"withlimit, or"cancelled", andunprocessed_files) before its summary; a run that lost files emits anincompleterecord (lost_files,total_files,cause("worker_faults"or"unexplained") anddetail); a run that died emits anabortedrecord. All of these were stderr text, which JSON mode promises to keep empty. - JSON mode keeps stderr empty by construction.
--suppressand the deprecated--check-xphonarenoticerecords, as is a Ctrl-C handler that could not be installed (interrupt_unavailable); cache maintenance is acacherecord, Ctrl-C prints nothing (the run ends with acancelledstop record), and an unwritable--auditfile is reported by the command instead of exiting from inside the renderer. validate --format jsonwrites each record straight to stdout, with no intermediate string, and a consumer that closes the pipe ends the stream: the run exits 1 with stderr empty. It panicked (exit 101, a panic message on stderr).- Breaking (CLI JSON): a run that found no transcript ends with a summary
whose
outcomeis"nothing_found"and which carries no counts; it was a"complete"summary oftotal_files: 0. validate --format jsonrecords and--auditJSONL lines are serialized from typed models: a diagnostic’sseverity(Error,Warning) is spelled by the model rather than taken from aDebugrendering, and keys come in a fixed order (typefirst).- Breaking (CLI JSON):
validate --format jsonemits exactly one file record per file. A file with warnings and no error is onevalidrecord carrying awarningsarray; it was aninvalidrecord followed by avalidone. Aninvalidrecord’serror_countcounts errors only (it counted warnings too), while itserrorsarray still lists every diagnostic with its severity. Aroundtrip_failedrecord carries itsdiffand anywarnings, and the contract lists that status.
What a run presents, and what it reads
- Breaking (CLI):
chatter validatedecides its output once, from all of--format,--quiet,--auditand--tui-mode: plain text, quiet text, JSON, an audit file, or the TUI. The TUI is chosen automatically only for plain text with stdout a terminal;--format jsonon a terminal used to open it. Flags that name two outputs are usage errors (exit 2):--auditwith--formator--quiet(it printed text and exited 0),--format jsonwith--quiet(--quietwas ignored), and--tui-mode forcewith any of them. - Breaking (CLI):
chatter validatereports an argument it cannot read, a nonexistent one included, as a read error in its results (✗ PATH (read error: ...), aread_errorrecord in JSON mode, counted among the invalid files), validates the rest, and exits 1. It used to refuse the whole run with stderr text a JSON consumer could not see. - One walk finds the input of
validate,to-json,fix, thedebugcommands and the desktop app’s directory runs. A directory or entry that cannot be read is reported, never skipped:validateand the desktop record it as a file that could not be read, and the other commands report it and exit 1 before processing anything. Links are followed (to-jsonused to skip them), a directory reached twice is walked once, and a link whose target is gone is a failure whatever its name, since it may have been a directory (a link to an unmounted volume’s subcorpus). A file named on the command line offixor adebugcommand is used whatever its extension, asvalidatealready did. - A transcript’s stored name is resolved once, where it is found: a walk
takes it from the listing that found the file, and a file argument is
resolved once before any worker starts.
fixand thedebugtools refuse an argument whose stored name cannot be resolved before processing anything (exit 1;fixused to skip it with an error line), andto-jsonon a directory refuses a transcript whose stem is not UTF-8 the same way. - A transcript named twice on the command line (two spellings of one file,
dir/a.chabeside./dir/a.chaordir/sub/../a.cha, or a file inside a directory also given) is processed once, under the first spelling given. fixand everydebugtool refuse an argument list that names no transcript (exit 1);debug overlap-auditanddebug linker-auditused to analyze nothing and exit 0.- Breaking (CLI): a
validate --auditfile that could not be written whole fails the run (exit 1), and the writer stops at its first failed write; a failed write was a warning, the file kept going with a hole in it, and the run could exit 0. - Breaking (CLI JSONL): every
validate --auditrecord carriesseverity("Error"or"Warning", spelled as the--format jsonrecords spell it), aftercode:{"file", "code", "severity", "message", "line", "column"}. A warning record had nothing to tell it from an error, though a warning does not fail its file. The record shape is now documented in the diagnostic contract. - Breaking (CLI): the
validate --auditsummary counts what its labels say.Files that failed(wasFiles with errors) counts the files the run failed (invalid, unreadable, a failed roundtrip or a tool failure), so a valid file with warnings is not one, and a file that failed with no diagnostic (unreadable) is.Total errorscounts errors only and the newTotal warningscounts warnings;Total errorscounted every diagnostic.Diagnostics by code(wasErrors by code) gives each code’s errors and warnings separately (E301: 2 error(s), 1 warning(s) in 2 file(s)).
The validation cache
chatter cache clear --dry-runwrites nothing: it opens the cache read-only, as an audit does (SQLite’s-shmand-walfiles may appear beside an existing database, as for an audit), and a cache with no database has nothing to clear. It created the cache directory, its lock file and the database, and ran migrations, to count zero rows.- Breaking (CLI):
chatter cache statsonly reads the cache, as the dry run does. With no cache database it saysNo cache database at PATHand exits 0 (no cache yet is a legal state, as the dry run and an audit treat it), and it creates nothing; it created the directory, its lock file and the database, and migrated it, to report zero entries of a cache it had just made. A database an older build left is reported, not migrated. - Breaking (CLI):
chatter cache clear --dry-runandchatter cache clearagree, because both start from one look at the cache directory. Over a database of an older schema the dry run exited 1 (“a schema this build does not read”) while the clear migrated it and went ahead; the dry run now says the clear would first migrate the database, and that the number of entries is known only after the migration (which can remove duplicates), and writes nothing; the clear says it migrated, then what it cleared. With no cache database, a clear creates none and says so (Cleared 0 cache entries: no cache database at PATH); it created and migrated an empty database to clear nothing. A database a newer build wrote failsstats,clearand its dry run alike (exit 1). - Cache keys are a specified hash. Every row was keyed by
DefaultHasher, whose algorithm Rust does not promise to keep, so a toolchain update could have turned the persistent cache into misses with nothing to say so. Rows are now keyed by blake3 over a written-down encoding of the path’s components and the row’s namespace (the book’s validation-cache chapter gives it). The key scheme is part of the cache version, so rows written by earlier releases are never served and leave through the existing prune of superseded versions: the first run after upgrading revalidates every file once. - The cache knows a transcript by its location with every directory
resolved by the operating system and its own name kept as stored, so
every spelling of a file is one row: relative or absolute, through
..or a linked directory, and/tmpbeside/private/tmpon macOS.validate --forceclears the cached verdicts of the files it validates whatever spelling the paths were given in; with a relative argument (validate --force a.cha) it cleared nothing and the stale verdict was served.cache clear --prefixresolves its prefix the same way (a directory since deleted through its deepest existing ancestor), and the missing-file purge no longer depends on where it is run from. - Breaking (CLI):
validate --auditonly reads the cache: it opens an existing cache read-only, creating, migrating, pruning, clearing and writing nothing (with no cache, or one an older build left, it runs without one and says so), and--audit --forceis a usage error (exit 2; it cleared the files’ rows). SQLite may still create its shared-memory and write-ahead-log files (talkbank-cache.db-shm,talkbank-cache.db-wal) beside an existing database, which any reader of a WAL database needs; the database itself is not written. - Validation reads each transcript once and keys the cache by the hash of
those bytes, so a verdict is stored for exactly the content validated even
if the file changes during the run. A cache read or write that fails is
counted and reported (a stderr warning,
cache_errorsin the JSON summary, a note in the desktop summary) instead of looking like a cold cache. - A run reports cache hits and misses only for files that consulted a cache,
and
cache_hit_rate(the JSON summary’s, the text summary’sHit rate,ValidationStatsSnapshot::cache_hit_rate()) is hits over those files (hits plus misses); it was hits over every file, so an unreadable file lowered it. A run whose cache did not open reports no misses, its JSONcache_hit_rateisnull(it was 0.0, with every file a miss), and the text summary saysCache: not consulted. A cached roundtrip verdict counts as a hit. - Breaking (CLI JSON):
chatter cache stats --format jsonis tagged bydatabase:{"database": "absent", "cache_dir": ...}when the directory holds no database,{"database": "older_schema", "cache_dir": ...}for a database an older build left (its entries are not counted), and"database": "current"withtotal_entries,cache_dir,cache_size_bytesandlast_modified. It reports what it could not find asnull, never as a stand-in:cache_dirisnullfor an in-memory cache (it was the string"in-memory");cache_size_bytesis the database file’s length, ornullif the file is missing when the statistics are read (it was0, the same as an empty file);last_modifiedis UTC with exactly three fractional digits (2026-03-09T13:05:31.000Z),nullexactly whencache_size_bytesis (it was the current time). A size or time that cannot be read, a time outside the years -9999 to 9999, and a cache directory that is not UTF-8 fail the command instead of being reported as 0,nullor a lossy path. The text output says(in memory)and(no cache file). - Breaking (CLI):
chatter cache clearwith neither or both of--alland--prefixis a usage error (exit 2; it exited 1).--dry-runwith--prefixreports how many entries it would clear, and a clear reports the number its delete removed.--prefixtakes any path the system accepts, UTF-8 or not, and selects the rows the cache wrote for it.
Commands
- Breaking (CLI): a command whose standard output closes (
chatter ... | head) exits 1, silently, from every text writer, asvalidate --format jsonalready did. Text output caught theprintln!panic and exited 0, so avalidaterun that found invalid files exited 0 when its summary had nowhere to go; a missed catch was a panic (exit 101). Every CLI text writer goes through one fallible writer, and the panic hook that matched “Broken pipe” in panic messages is removed. - Breaking (CLI):
to-jsonon a directory with no.chafile (an empty tree, or a mount point with nothing mounted) exits 1 withERROR: no .cha files found in DIRbefore converting or pruning anything. With--pruneit exited 0 and deleted every.jsonunder--output-dir, each read as an orphan of the empty input; a prune now needs a non-empty population of transcripts by type. - Breaking (CLI): flags are parsed into what they select, and a
combination a command would have ignored is a usage error (exit 2):
normalize --skip-alignmentwithout--validate;to-json --skip-validation --skip-alignment;validate --list-checksbeside a path; andto-jsonwith an option for the other kind of input (--output-dir,--force,--pruneor--jobsfor a file;-o/--outputfor a directory). A file input used to ignore the directory options. A directory input without--output-diris a usage error (exit 2; it exited 1). - Breaking (CLI): a
--code(fix) or--suppress(validate) value that names no known error code (or, for--suppress, no known group) is a usage error (exit 2) naming the value; the commands printed their own error and exited 1. - Breaking (CLI):
--jobs 0(onvalidateandto-json) is a usage error (exit 2); it ran one worker with a logged warning. Omit--jobsto use every CPU. - Breaking (CLI):
chatter cache statstakes-f/--format text|json, asvalidatedoes, in place of--json. chatter fixexits 1 when a file could not be read or--applycould not write a fix, and says how many on stderr; it exited 0 after reporting each failure.chatter to-jsonandchatter normalizereport a failure the same way: a lineERROR: <path>: <failure>, then the rendered diagnostics of a parse failure, a validation failure, an incomplete validation or an internal failure.to-jsonon one file printed only a one-line summary for some of these, under its own✗ Validation errors found:and✗ JSON error:headlines;normalizeprinted✗ Validation errors found:for a validation failure and dropped the diagnostics of every other failure behindError: <failure>.chatter to-json --prunestill never follows a link in the output tree, while the walk that reads transcripts now follows them: its deletions stay inside--output-dir.--prunealso no longer deletes a JSON file whose transcript merely could not be checked, and reports a file or empty directory it cannot remove (the run then exits 1).chatter to-json <dir>warns when it cannot remove the stale JSON of a transcript that now fails to convert, runs on the shared worker pool (a worker that panics is reported with the files it left unconverted, and the run exits 1), and counts files queued on its progress line.chatter watchvalidates each changed file through the same one-file pipeline asvalidate(stored name, cache, rules), withvalidate’s default cache identity, and opens the cache once when it starts (it reopened and pruned it on every save). It keeps watching when a watched file cannot be read (a rename or a lock mid-edit) instead of exiting, and prints a file’s diagnostics asvalidate --quietdoes.
Library API
- Breaking (library, talkbank-cache):
RulesVersion::for_testingexists only with the newtest-supportfeature, which no production build enables; a production caller could name a version no rule set produced. The crate’s examples useRulesVersion::current(). - Breaking (library, talkbank-model): a file stem is checked where it is
made.
FileStem::from_stemreturnsResult<FileStem, FileStemError>, refusing an empty stem and one with a path separator, andOwnedTranscriptName::Namedholds anOwnedFileStem(built from a path, a checked string or aFileStem) instead of aString;TranscriptName::to_owned_namemakes one. The LSP names a document by its decoded file path (TranscriptName::for_path), so a file name with a space or an accent is matched against@Mediaas stored, not percent-encoded. - Breaking (library): a validation run takes
DistinctTranscripts(talkbank_transform::paths, built only byDistinctTranscripts::new), stored transcripts each named once by resolved location, sovalidate_files_streaminghanded one file under two spellings validates and counts it once, as the argument path already did; it validated it twice.ExpandedArguments::into_partsreturns them. - Breaking (library): a validation run’s ending is typed. A stream ends
with one
ValidationEvent::Finished(RunEnding)(it ended withFinished(stats),FinishedIncompleteorAborted), whereRunEndingisComplete(stats),NothingFound(no transcript at all; it was a complete run of zero files),Stopped { stats, unprocessed, reason },Incomplete { stats, lost_files, cause: LossCause }(LossCauseisWorkerFaults, never empty, orUnexplained) orAborted(AbortReason), with counts asNonZeroUsize.RunEnding::passed()is the one answer to whether a run vouches for its input, the answer the CLI, the TUI and the desktop all give.AbortReasongainsNoEnding, for a consumer whose stream closed without an ending, and the newCancelReason(ErrorLimit { limit }orRequested) implementsDisplay, the wording every surface uses for a stop.ValidationStatsSnapshothas private fields read through accessors, with a non-zerototal_files, and only the runner makes one (itscancelledfield is gone: a stop is theStoppedending);RunCoverageis no longer public. The streaming entry points return aCanceller(wasSender<()>) and are no longer generic over the cache (they take aValidationRun). - Breaking (library):
ValidationConfig:check_alignment: boolisalignment: AlignmentValidation;roundtripis aRoundtripCheck(SkiporRun; was abool);jobsisOption<NonZeroUsize>;error_limit: ErrorLimit(UnlimitedorStopAfter(n)) is new; anddirectoryandcacheare removed (every walk descends every level, and the cache is a parameter).DirectoryModeandCacheModeare removed. The atomicValidationStatsis removed: each worker counts its own files and the runner sums them after the join. - Breaking (library): every runner entry point
(
validate_directory_streaming(directory, &run),validate_files_streaming(files, &run),validate_arguments_streaming(input, &run)) takes aValidationRun, aValidationConfigbound to the cache the run may use; they took the configuration and the cache as two arguments, so a library caller could hand a run under one rule set a cache opened for another (strict linkers, another parser) and get that cache’s verdicts.ValidationRun::new(config, cache)refuses a cache whose identity is notconfig.cache_identity()(CacheIdentityMismatch),ValidationRun::uncached(config)binds none, andValidationRun::config()reads the bound configuration.VerdictReadergains a requiredidentity(). The CLI opens the cache from the run’s configuration and binds it; the desktop binds the pool it memoizes per identity, andvalidate_target_streaming_with_config(target, &run)takes the bound run. - Breaking (library): a run’s cache is a
RunCache(Absent,ReadOnly(Arc<dyn VerdictReader>)orReadWrite(Arc<dyn ValidationCache>)), and how a file used the cache isFileCompleteEvent::cache: CacheUse(Hit,Miss,NotConsulted);FileStatusloses itscache_hitfields, andcache_hit_rate()returnsOption<f64>. - Breaking (library): the cache traits.
VerdictReader(get,get_roundtrip) andValidationCache: VerdictReader(set,set_roundtrip), all required, take aResolvedPath(fromtalkbank-model: the parent directory resolved, the name kept as stored; everyStoredTranscriptcarries its own,StoredTranscript::resolved), theContentHashof the bytes validated, and anAlignmentValidation, and returnResult<CacheLookup<V>, CacheError>(a failed lookup is an error, never a miss). Validation verdicts areCacheOutcome; roundtrip verdicts are the newRoundtripOutcome(Passed,Failed).ReadOnlyCacheopens an existing, current cache read-only and writes nothing; it refuses a missing one (CacheError::NoCacheDatabase) or one whose schema is not this build’s (CacheError::SchemaNotCurrent).CachePool’s own get/set methods,clear_prefixandclear_allare removed: maintenance iscount(&CacheScope)andclear(&CacheScope)withCacheScope::AllorUnder(ResolvedPrefix), andclear_paths(any iterator of&ResolvedPath) clears by key in every namespace.CacheStatsholds aStorageStats(InMemory, orDirectory { cache_dir, database }withDatabaseFile::MissingorPresent { size_bytes, modified });SpaceReclaimedgainsVacuumedSizeUnknown;VersionPruneOutcomegainsFreshDatabase(an in-memory cache, where no prune ran);VersionPruneReport::versions_deletedis ausize;purge_nonexistentfails on a path it cannot check instead of deleting its entry. NewCacheErrorvariants:CorruptColumn,CountOutOfRange,ModifiedOutOfRange,NoCacheDatabase,SchemaNotCurrent. - Breaking (library, talkbank-cache): administration starts from
CacheOnDisk::inspect()(orinspect_directory(dir)), one look that creates, migrates and writes nothing and returns what is there:Absent(NoDatabase),OlderSchema(OlderSchema)orCurrent(InspectionCache). AMaintenanceCacheis reached only from that value, byOlderSchema::migrate()orInspectionCache::into_maintenance():MaintenanceCache::openandopen_directory, which created and migrated a database in any directory they were given, andInspectionCache::openandopen_directoryare removed. A database a newer build migrated is the newCacheError::SchemaNewer(forReadOnlyCachetoo, which reported it asSchemaNotCurrent, whose message said a writing run would upgrade it);SchemaNotCurrentnow means an older schema only. - Breaking (library): one event per file.
ValidationEvent::ErrorsandRoundtripComplete, withErrorEventandRoundtripEvent, are removed: a file’s diagnostics travel inside itsFileStatus, in the variants that can have them:Invalid { diagnostics: InvalidDiagnostics }(never empty, with aNonZeroUsizeerror_count()),Valid { warnings }andRoundtripFailed { warnings, diff }(Option<FileDiagnostics>), andInternalFailure { failure, attempt }, whoseFailedAttemptsays whether the failure’s diagnostics point into the file’s text (Validation { source }) or into the roundtrip’s serialized text (RoundtripReparse).FileStatus::shown()gives the diagnostics to show against the file, with that text;FileStatus::failed()says whether the file failed. A consumer renders each file once, and the file’s text moves into its status without a copy. - Breaking (library):
StoredTranscriptcarries itsResolvedPath(StoredTranscript::resolved), made when it is admitted, so a transcript whose directory cannot be resolved is refused byresolveandfrom_entry. - Breaking (library, talkbank-model): the coordinated splice
MorTier::splice_range_coordinated(gra, item_range, block, root, redirects)(andsplice_coordinated(gra, item_idx, block, root, redirects), the same over one item) takes its replacement as aSplicedBlock, says where the block’s root attaches with aSpanRoot, and where host dependents of the replaced items go with aHostRedirects; it adds no cycle and no second root, so a host%gratree stays a tree. Every type is re-exported besideMorTierintalkbank_model::model::dependent_tier::mor, withCoordinatedMutationError.SplicedBlock::new(mors, relations)replaces thenew_morsandnew_relationsarguments. It parses the block-relative relations once and refuses (SplicedBlockError) a relation count that differs from the chunk count (wasCoordinatedMutationError::CountMismatch), a head outside the block (wasHeadOutOfNewBlock), a relation whoseindexis not its chunk (IndexOutOfOrder), a block with no root or several (RootCount), and a cycle (Cycle). The splice wrote a rootless cyclic block into the host: the1 -> 2,2 -> 1reparse ofpor@s favor@sbecame a cycle in the host’s%gra.SplicedBlock::root_chunk()exposes the admitted root as aBlockChunkso callers can direct host dependents to it without a separate root scan.- A host
%grarelation that depended on a replaced item keeps depending on that word. The splices rewrote such a head to the first chunk of the new block, by position: in morphotag’s L2 output forich glaube it's@s:eng working@s:eng und don't@s:eng stop@s:eng .that madeunddepend ondoandstoponit, where they depend onstopandworking. The caller states the correspondence:HostRedirects::ByItem(equal item counts, each item following its counterpart) orHostRedirects::PerItem(Vec<ItemTarget>), one target per replaced item:ItemTarget::Chunk(BlockChunk), a chunk of the block, orItemTarget::Counterpart, the block item at the same position (chunk for chunk when the two items have the same chunk count, otherwise the new item’s head chunk), so a caller can state one target and let the rest follow. A statement that does not fit, or a dependent whose item has no unique head chunk, is refused before either tier changes (RedirectItemCountsDiffer,RedirectCountMismatch,RedirectOutOfBlock,NoCounterpart,NoUniqueHeadChunk). SpanRootreplaces theroot_anchor_override: Option<usize>argument:UtteranceRoot(head0, relationROOT), orHostChunk { chunk, relation }, a host chunk (SemanticWordIndex1) numbered as before the splice and translated by it, with the relation the span root takes under it, anAttachmentRelation, whose constructor refuses a root label (RootRelationUnderHost), soROOTunder a host head cannot be written; the block’s own label for its root is not kept.SpanRoot::from_gra_headbuilds one from a host relation’sGraHeadRefand the relation. The old anchor was a pre-splice index never shifted, andNonetook the head of the range’s first chunk: indont@s:eng mal geh .the span’sdocame to depend onmal. A span root at the utterance’s root while a host relation outside the range is already the root is refused (UtteranceRootTaken; the splice wrote two roots), as is a host chunk inside the replaced range or past the host (SpanRootInReplacedRange,SpanRootOutOfHost), or one whose own chain of heads reaches the range (SpanRootDependsOnSpan: withx@s y@s z .andz -> x, anchoring the spanx yatzreturned a cycle with no root).- A host whose
%gradoes not number each relation by its chunk (relationk, from 1, with indexk) is refused (HostIndexOutOfOrder), since every head is read as a chunk number. Such a host was spliced anyway, its misnumbered relations’ indices shifted by arithmetic that could wrap below zero on a shrinking splice. Inside the splice, the numberings before and after it, of the block, and of the replaced range are distinct types with one translation between the host’s numbering before and after, so a pre-splice index cannot be written into the result. CoordinatedMutationErrorderivesClone,PartialEqandEq, and its host-chunk fields areSemanticWordIndex1. It stays exhaustive, as the crate’s error enums are: a new refusal is a compile error for a caller that matches them. talkbank-tools adapts when it moves to this release.
- Breaking (library):
talkbank_model::ParseValidateOptionshas private fields; its publicvalidate,alignmentandstrict_linkersbooleans could say “alignment without validation”. The level is the newCheckLevel(ParseOnlyorValidate(AlignmentValidation)), set withwith_leveland read withlevel();should_validateis removed (usevalidation_policy().is_some()).RuleSelectionholds aLinkerChecks. - Breaking (library, talkbank-model
async):validate_asyncandvalidate_with_rules_asynctake anOwnedTranscriptName(Named(String)orAnonymous) in place offilename: Option<String>, whoseNonesilently meant “skip the rules about the transcript’s name”. - Breaking (library):
chat_to_json,chat_to_json_named,chat_to_json_with_schema_policyandchat_to_json_unvalidatedtake aJsonLayout(PrettyorCompact; waspretty: bool), andJsonSchemaPolicy::from_skip_flagis removed: the libraries keep no bool-to-mode constructors, and the CLI parses each presence flag straight into the mode it selects. - Breaking (library): timestamps in the TOML decision files are
talkbank_transform::recorded_time::RecordedTime, replacingchrono::DateTime<Utc>(PendingEntry::created_at,MergeOverride::decided_at, and the matching parameters ofMergeOverride::auto_decision,operator_decisionandjudgment_to_pending). It is written as New York time in whole seconds with its offset ("2026-05-27T08:41:00-04:00"), whichever host writes it, and reads every RFC 3339 time with an offset, so existing files load unchanged, and in TOML also a native datetime.
Desktop
- The desktop app shows a stopped (cancelled) run as its own state, with the
number of files it never reached, and an incomplete run with its cause;
FrontendStats.cancelledis replaced by astoppedevent. - The desktop’s all-valid claim is the runner’s verdict, carried as
passedon thefinishedevent, so it agrees with the CLI’s exit status: a run whose files have only warnings is “All N files valid; K warnings”. A target with no transcript is its ownnothingFoundstate, with Re-validate offered; it was a finished run of zero files. Stop, loss and abort reasons are the runner’s own wording. - A validation cache that will not open is said in the run’s summary, with the reason, and the run goes on without it; it was a line on the app’s stderr, which a desktop user never sees.
- The
validate,open_in_clanandexport_resultscommands each take onerequestargument (ValidateRequest,OpenInClanRequest,ExportResultsRequest), the value the backend already used, in place of loose ones. A text export writes each file’s status as the app shows it, one label from one owner, where the backend kept its own copy of the sentences. ValidateRequest’sjobsisOption<NonZeroUsize>(wasOption<u32>), sojobs: 0is refused when the request is deserialized. The Parallel jobs field parses its text into a positive whole number (or empty for all CPUs), says when an entry is not one and which value the next run uses, and never sends0or1.5on.
Internals
- One worker pool,
talkbank_transform::worker_pool::fan_out, runs the validation runner andchatter to-json <dir>. Its workers run on the same 16 MiB stack as the CLI’s program thread (they used the 2 MiB thread default),--jobs 1is the same pool at width one, and each worker returns its own counts instead of updating shared counters. - chatter no longer depends on
chrono; it usesjiff.
Removed
- Breaking (CLI):
chatter fix --dry-run. A barefixalready reports without writing, and--apply --dry-run(meaning “do not apply”) was confusing. Passing--dry-runis an unknown-argument usage error (exit 2), so a script using it fails instead of writing. - Breaking (CLI):
chatter to-json’s hidden--validateand-a/--alignmentflags, deprecated no-ops since validation and alignment became the default; passing one is a usage error (exit 2) instead of being ignored. - Breaking (library, JSON, desktop protocol): the parse-error file
status, which nothing produced (both parsers always return a model with
diagnostics, and an unparsable file is
Invalid):FileStatus::ParseError,ValidationStatsSnapshot::parse_errors, the JSON summary’sparse_errors(always 0), theparse_errorrecord status, the text summary’sParse errors:line, and the desktop’sparseErrorstatus andparseErrorscount. The desktop summary counts “invalid or unreadable files”, which is whatinvalid_filesholds. - chatter’s
RoundtripValidationMode, a second copy ofRoundtripCheck. - Breaking (CLI):
chatter debug overlap-audit -f/--format, which the command never read (it prints TSV, with--databasefor JSON lines); passing it is a usage error (exit 2). - Breaking (library):
LossCause::tag; a JSON consumer reads the CLI’sincompleterecord, whosecauseis serialized from its own wire type.
Added
- W110 warns when the
@Mediafilename and the transcript’s own name differ only in letter case (Session.chadeclaring@Media: session). CHECK’s comparison ignores case, so this stays out of E531, but such a name finds its recording on a case-insensitive filesystem and not on a case-sensitive one. Case and Unicode spelling are reported independently, so a name can carry both W110 and W109. PipelineError::diagnostics()returns the located diagnostics a pipeline failure carries, orNone;to-json,normalizeand speaker identification report failures through it.talkbank_model::LinkerChecks(Lenient,Strict) andRuleSelection::with_linkers, so a caller holding the linker choice as a value selects the strict linker checks without a branch of its own.ChatFile::validate_at(policy, errors, name)validates under aValidationPolicy: its rules, at its alignment coverage.talkbank_transform::pathsfinds input the one way every command uses:walk_files(aWalkofFoundFiles andWalkFailures, with aLinkspolicy,FolloworSkip),walk_transcripts(aTranscriptWalkofFoundTranscripts, each a relative path and aStoredTranscript) andexpand_transcript_arguments(ExpandedArguments, oneStoredTranscriptper resolved location).StoredTranscript::into_path.talkbank_model::ResolvedPathandResolvedDirectory, a file’s location with its directory resolved by the operating system and its own name kept, the identity the cache keys by, andResolvedPrefix, a resolved directory or file that a cache scope selects under (all re-exported bytalkbank-cacheandtalkbank-transform).validate_arguments_streaming, the runner entry point for expanded command-line arguments, which reports each unreadable argument as a read error in the run’s results.talkbank_transform::worker_pool(fan_out,PoolRun,PoolOutcome), the one bounded worker pool (see Internals), andPoolOutcome::faults, the one reading of how its workers ended, asPoolFaults (Unwound,ThreadRefused); the runner’sWorkerFaultwraps one asPool.
0.27.0 - 2026-09-28
Changed
-
Cross-platform verification keeps Windows workspace doctests on Cargo’s normal C runtime, avoiding a Tauri static-runtime library-search collision. Desktop release packaging retains its static-runtime configuration.
-
Experimental re2c conversion preserves multiword participant names and treats lexer continuation tokens as separators in logical gem labels. Diagnostic differences remain explicitly assessed, not copied from the default parser; re2c is still not a production validity authority.
-
Pseudonymization review stores proposed morphology and refused-plan payloads behind owned boxes, reducing inline variant and error sizes without changing admission or output policy.
MorphologyOutcome::Proposed::proposedis nowBox<Mor>for Rust callers. -
Updated the Rust and JavaScript desktop dependencies together, including Tauri and its plugins, and refreshed jsonschema, cc, and Vite.
-
Participant parsing admits selected source ranges before decoding names and roles. Recovery states are retained; unreadable producer ranges reject the entry as an internal failure rather than admitting a partial participant.
-
Breaking (low-level Rust parser API):
parse_postcode_nodeaccepts a generatedSourceBound<PostcodeNode>instead of a node and separate source string. Final-code lowering preserves source association and propagates range failures as internal failures.ChatParserAPIs and CHAT/JSON behavior are unchanged. -
The primary parser preserves postcode source spans, allowing incremental relocation and diagnostics to retain the original token location. CHAT and JSON output are unchanged.
-
Breaking (Rust API):
Utterance::newstarts with unknown provenance;parse_healthis private and read throughparse_health(). Parser adapters finish through accumulatedParseHealth::finish_utterance. Appending dependent tiers withdraws provenance and alignment caches. Programmatic producers useChatFile::validate_construction_with_policyfor checked typed construction, without reparsing CHAT.ParseHealthState::Constructedpermits model checks but never source-byte splicing or a claim of parser-backed cleanliness. JSON remains unchanged and carries no runtime admission evidence. -
English decade generation now admits only multiples of ten in 0-90 shorthand or 1100-2990 full-year form. Other suffix-bearing inputs retain their exact spelling instead of acquiring a guessed plural year phrase.
-
Generated typed-CST carriers support opt-in range admission, preserving recovery while making subsequent payload text reads infallible.
@Mediabody lowering uses this boundary; CHAT and JSON policy are unchanged. -
Breaking (generated Rust API): selected repeat/optional elements carry an uninhabited
Absentpayload. Empty repetitions and optionalNoneremain possible; an element already selected cannot independently become absent. -
LSP quick fixes no longer invent participant or language facts for E308, E504 or E507, matching the shared fix catalog. Diagnostic prose is no longer parsed to construct participant declarations; enter the actual facts explicitly.
-
English ordinal generation preserves tokens outside its supported 0-9999 range unchanged, instead of rewriting their suffix to
th. -
English ordinal generation omits prose commas in thousands with remainders, retaining the existing conjunction convention. For example,
1234thexpands toone thousand two hundred and thirty-fourth, a spoken word sequence. -
Breaking (Rust API):
ReplacementWords::newandTryFrom<Vec<Word>>reject empty lists; JSON admission enforces the same invariant.Replacement::newaccepts admittedReplacementWords. Useinto_vecand checked reconstruction instead oftake/retain, or edit elements through mutable slices. Single-word construction remains infallible. -
re2c replacements use the existing word grammar directly, retaining source spans instead of splitting/reparsing text or inventing plain words on failure. Empty and unclosed replacements are parser errors. Glued replacements receive a diagnostic on re2c’s opening token, without copied tree-sitter recovery diagnostics. Historical E208 is deprecated; primary-parser E376/E342 rejection and the empty-replacement regression remain.
-
Breaking (Rust API):
PauseTimedDuration::Parsednow carries a checkedParsedPauseDurationwith private fields and read-onlyseconds(),millis()andas_str()accessors. Construct throughPauseTimedDuration::new; CHAT and JSON formats are unchanged. -
Diagnostic display enrichment preserves legitimate zero-width EOF locations instead of moving them back onto the preceding byte, including when the file ends with a newline.
-
E306 empty-turn deletion now declines turns with dependent tiers, preventing their content from being silently reassigned to the preceding turn.
-
E241 marker-spelling repairs now require an exact source-bound word. The existing spelling vocabulary remains authoritative; comment text and omission/shortening notation are not rewritten as plain markers.
-
E750 group-edge repairs now require source-bound whitespace at the owning annotated group’s content boundary instead of neighboring delimiter bytes. Complete space runs and nested groups retain recovery-repair admission; groups with unrelated structural recovery decline a proposal.
-
E244 now reports adjacent stress markers once per affected word. Longer runs no longer create duplicate diagnostics and overlapping repair proposals.
-
Duplicate-comma proposals now require exact source-bound comma tokens in a clean tier body, sharing token admission with semantic comma deletion.
-
Duplicate-primary-stress repair now traverses source-bound stress tokens. An earlier isolated mark no longer hides a later duplicate run, and all separate primary-duplicate runs in the word are repaired together. Only duplicate tokens are removed; other stress positions and diagnostics remain.
-
E259 comma-deletion proposals now use source-bound comma, tier-body and whitespace nodes. Removing an initial comma consumes its complete separator instead of leaving leading spaces; interior commas preserve their separator. The proposal remains a semantic change requiring user review.
-
Missing-terminator fix alternatives now use a source-bound grammar ending for insertion, before final postcodes on main tiers. This prevents misplaced terminators and splitting CRLF line endings, and refuses tiers whose structural boundary cannot be established. Choosing the terminator remains a user decision; these proposals are not automatic repairs.
-
E501 fix proposals now require source-bound, byte-identical repeated headers with no conflicting declaration of that kind. Conflicts, formatting differences and recovered structure decline a proposal. Complete CST header ranges replace physical-line guessing; the existing header-write restriction is unchanged.
-
Structured dependent-tier recovery now retains generated source ownership through recursive traversal instead of accepting raw nodes and independent text. Utterance recovery slots likewise preserve their source-bound fields through read admission. Recovery classifications, conservative alignment taint and placeholder behavior are unchanged.
-
Adding recovery taint no longer promotes unknown parse provenance (such as JSON-imported utterances) into partially clean provenance. Alignment remains unavailable until parser-backed provenance has actually been established.
-
Unreadable CST source ranges in generic, dependent-tier and utterance recovery now report internal failure, not invalid CHAT or an encoding repair suggestion. Readable recovery diagnostics and conservative alignment taint are unchanged.
-
Morphology tier-body, main-word and post-clitic parsing now consumes the generated canonical-grammar admission proof. Impossible missing-composite states are removed from these boundaries; lexical and structural recovery, absent children and internal-failure reporting remain distinct.
-
Fix planning no longer invents participant declarations, roles or default languages for E308/E504/E507. Supply those facts explicitly. E604 removal uses the complete typed GRA tier, preserving intervening dependent tiers and handling continuation lines; multiple targets refuse selection. E306 likewise uses the typed main-tier boundary rather than a line-prefix guess and refuses grammar-recovered main tiers.
-
splice::catalog_fixnow borrowsParsedSourceinstead of accepting raw source text. Retain the owner fromTreeSitterParser::parse_chat_file_with_sourceand pass diagnostics from the same input. W109 planning reuses this CST instead of creating a parser and reparsing for each diagnostic. Recovery admission and changed-output verification are unchanged. -
Speaker-identification input-rejection reports now use the distinct
incomplete_validationfailure category when parser provenance prevents complete validation. Previously this was collapsed intovalidation. Consumers of the typed enum or JSON failure category must handle the new variant; no match evidence or CHAT-invalidity claim is inferred from it. -
Word spelling is derived from typed structure, including the JSON
raw_textfield. Imported computed spelling cannot override content. RustWord::newandnew_uncheckedno longer accept a separate raw argument;raw_text()returns an owned string andset_raw_textis removed. Use typed builders for markers and streamingWriteChat/Displaywhen appropriate. Original source bytes must be read from the source, not from derived spelling. External token producers should handle fragment-parser refusal instead of falling back to unchecked construction. Shortening/embedded-marker errors retain grammar diagnostics without duplicate raw-spelling rescans. -
Number conversion uses exact lexical entries for languages without a language-specific composer. Unsupported numerals, currency and number groups remain unchanged instead of using generic multiplication or partial rewrites.
-
Transcript construction rejects malformed media source spelling before reducing local paths, including malformed discarded directory components. Admitted remote URLs remain unchanged.
-
Generated source-bound
extract_admittedAPIs accept compiled-language evidence and exclude Missing only for proven nonterminal slots. Participant and language headers, document/utterance and main-tier/body/ending reconstruction use this admission; lexical Missing, Error, applicable Absent and source-read failures remain supported. Raw extraction stays broad.ReconstructionFaultaddsGrammarBindingfor failed producer admission; downstream exhaustive matches must handle this as an internal tool failure. Low-level pre-begin and dependent-tier dispatch now consume generated admitted choice types. Document recovery source-ownership failures report E001 rather than classifying a producer fault as invalid CHAT. -
Generated repeat elements and present optional elements now use selected-slot types with uninhabited
Absentpayloads. Consumers can eliminate impossible absence diagnostics while retaining MISSING/ERROR recovery, nested fixed positions, empty repeats and optionalNone. Kind-preserving projections retain the slot’s absence type. -
Low-level
%pho/%modparsing now requires source-bound generated nodes and returnsCstFailurefor source/reconstruction failures. Compound spelling, grouped content, recovery and string fragment APIs are unchanged. -
Low-level
%sinparsing now requires a source-bound generated node and returnsCstFailurefor source/reconstruction failures. Sign-group fallback and empty-token policies remain; string fragment APIs are unchanged. -
Low-level
%graparsing requires a source-bound generated node and returnsCstFailurefor source/reconstruction failures. String fragment APIs, numeric admission, recovery and relation completeness accounting are unchanged. -
Low-level
%mortier parsing now requires a source-bound generated node instead of a node plus independent text. Bind through the existingParsedSourceowner; string fragment APIs are unchanged. Recovery and placeholder handling are retained. Word/post-clitic/feature readers retain the same source association; producer faults propagate separately from missing morphological content. -
Generated
ChoiceSlotmakesUnexpecteduninhabited: selected-choice extraction consumes its retained match plan and returns producer faults throughReconstructionFault. Parser consumers no longer invent CHAT diagnostics for that impossible choice state. Missing, Error, Absent, displaced nodes and supertype-classification recovery remain supported. -
Source-bound bullet/text tier adapters return
CstFailure, preserving source-binding failures alongside reconstruction faults without substituting empty content. The nested bullet-content reader propagates the same failures rather than returning a shortened segment list. High-level fragment parser signatures are unchanged. -
ChatDate::Validnow holdsCheckedChatDatewith private fields. Construct dates throughChatDate::from_textornew; matchValid(date)and useday(),month(),year()andas_str()instead of field access. JSON remains a string; format/day-range admission is unchanged (not calendar validation). -
Generated concrete CST extraction now returns
Result<_, ReconstructionFault>; ERROR-root extraction returnsResult<Option<_>, _>. Source-bound match plans retain selection decisions for consumption instead of independently rematching. Affected public header and tier adapters also returnResult; callers must propagate producer failure, not substitute empty carriers.DocumentRoot::classifynow returns the publicly exportedCstFailure. -
E001 is classified as
DiagnosticKind::InternalFailure, not CHAT invalidity. Validation/admission and transformation pipelines preserve this distinction. CLI file records addinternal_failure, summary records addinternal_failures, and desktop events carryinternalFailure/internalFailures. Failed attempts are neither valid nor invalid and never enter the validation cache. Producer faults during optional roundtrip reparsing also remain internal failures, rather than mismatches, and write neither validation nor roundtrip cache entries. Legacy checked node/source reads report E001 for incompatible ranges or UTF-8 boundaries instead of misclassifying those API faults as CHAT errors. -
Flagged reference-first merge drafts return ordering uncertainty exclusively through
DraftOrderReviewrecords; they no longer insert generated review@Commentlines. Contributor comments, source ordering and strict-policy refusals are unchanged. Reviews are also accessible on unvalidatedMergeDraft.before_output_utterancenow has typeOutputUtteranceBoundary; useutterances_before()for its count rather than treating it as a line index. -
ParseErrorBuilderrequires message and location in its type state beforefinish(), which now returnsParseErrordirectly. RemovedParseErrorBuilderErrorandtry_finish(); callers must supply both fields and remove result/option handling around finishing. Streaming source-range rejection now returns the original parse diagnostic, not a parser-creation error or a fabricated empty document. -
Spanish cardinal expansion uses bounded hundreds/thousands composition, correcting
100000tocien mil. Unsupported noun scales and unresolved agreement preserve the original input rather than inventing a phrase. -
Lexical Unicode validation now rejects private-use scalars and noncharacters across all planes instead of copying CHECK’s high-BMP blacklist and private markup exemption. Ordinary compatibility characters are no longer rejected by that range rule; control-character checks are unchanged.
-
Retrace joining carries admitted utterance ownership into mutation instead of raw line indices. A forward scan preserves header barriers and chain repairs without repeated removal of later lines; repair policies are unchanged.
-
Vector and small-vector semantic differences share one sequence comparison policy, preserving bounded-report ordering and unmatched-tail diagnostics.
-
ErrorContext::from_source_with_spannow returnsResult<_, SourceExcerptError>. Invalid source slices and unrepresentable snippet coordinates are refused, rather than replaced by an empty context. Valid zero-width ranges remain valid. -
Context label fields are private. Construct
SampleTypeLabel,RoleLabelandConsentTierLabelthroughTryFrom<String>; read them withas_str(). Corrected sample-type verdicts share this nonblank admission boundary. -
Standalone-word conversion and
%worparsing require producer-bound nodes, not separate node/source pairs. Main-tier words, replacement words, timed words and standalone fragments preserve their source association to conversion. -
Main-tier CST conversion now requires a producer-bound
MainTierNode, not a typed node paired with independent source text. Fragment, utterance, and EOF recovery callers retain that association through conversion. -
Main-tier contents parsing now consumes a bound
ContentsNode; the shared contents/group cycle retains generated source association through recursive group and quotation dispatch, including recovery placeholders. -
Speaker-mapping strings now refuse repeated source assignments, including identical repeats, instead of silently keeping the last assignment.
-
AdjudicationErroraddsInvalidDecisionMapping; exhaustive library matches must handle this refusal when a rename lacks its required adult role.
Fixed
-
Transcript construction preserves admitted HTTP/HTTPS media references verbatim instead of reducing them to local filesystem basenames.
-
Line-map end lookup clamps extreme out-of-range line indexes to source EOF instead of overflowing its next-line calculation.
-
Empty source spans no longer report overlap with a surrounding range; diagnostic insertion points remain distinct from byte-covering highlights.
-
ParseError::internalnow emits E001 (InternalError) rather than a CHAT syntax diagnostic, preserving failure-versus-invalidity admission even if presentation severity is downgraded. -
Display-span normalization clamps extreme out-of-source coordinates before interpolation, avoiding integer overflow while retaining UTF-8-safe spans.
-
Structured bullet timestamp readers now retain the producing source through all four consumers and distinguish failed source admission from missing fields.
-
Main-tier body lowering propagates failed content/ending reads as internal failures and preserves source ownership into utterance-ending extraction.
-
%worbullet handling retains internal reconstruction failures instead of silently treating them as unavailable word timing. -
Header lowering no longer substitutes an unknown comment or a partial participant name/role after an internal source-read failure.
-
Rebasing a
ParseErrornow preserves its independently owned context text and relative highlight instead of shifting that highlight out of its snapshot. Document locations and secondary labels continue to move together. -
Dependent-tier fragments reject trailing tiers or speech instead of returning the first tier and silently discarding the rest of the caller’s input.
-
Single grammatical-relation, phonological-word and participant-entry fragments reject extra items instead of silently discarding them. GRA and PHO fragments no longer append synthetic content; empty-input errors retain caller offsets.
-
Single MOR-word fragments reject multiple items and post-clitics instead of silently returning only the first main word. Refusal diagnostics retain caller coordinates; complete MOR-tier parsing still accepts those structures.
-
Lenient transform parsing no longer hides diagnostics on tiers whose names merely start with
mororgra. Suppression uses actual parsed generated-tier ownership and retains unlocated or cross-tier diagnostics and alignment taint. -
Shortening validation uses source-bounded, nonnegative nesting depth, avoiding signed-counter overflow on very large imported spellings while retaining unmatched-closing and unclosed-opening diagnostics.
-
Judgment context consumes typed header ages instead of reparsing raw text. Unsupported ages no longer yield a numeric age from a valid-looking prefix; bounded components prevent arithmetic overflow, and explicit ages skip fallback.
-
Speaker-sample head/tail limits no longer overflow for large budgets. Selection is bounded by available turns and overlapping windows never duplicate speech.
-
Adjudication library failures now preserve the failing request and every later pending request in order. Retrying does not reapply already accepted decisions.
-
Adjudication admits mapping conversion before recording a decision. A rename without its required role is refused without consuming the pending request; the CLI exits with status 2 and leaves its files unchanged.
0.26.0 - 2026-09-24
spec-perturbemits explicitly unreviewed candidates, not stale expected diagnostic labels. Its JSON replacesexpected_errorwithassessment: "unreviewed"; reviewed spec claims remain the golden authority.
Changed
-
Validation and roundtrip cache keys retain native path identity instead of hashing lossy display text. Equivalent separator spellings now reuse the same cache fact, including mixed Windows separators; existing cache entries may be relearned without deleting the cache.
-
English cardinal generation composes short-scale units correctly instead of multiplying complete table phrases (for example, 2000 now yields “two thousand”, not “two one thousand”). Authored reference controls cover scales through the
u64boundary; other languages’ generation policies are unchanged. -
Bullet-text tier adapters and
parse_bullet_contentnow require producer-bound nodes rather than separately supplied source text.BulletTextNodecarries the source lifetime as well as the tree lifetime. Generated child-range admission and recovery diagnostics remain; this is source association, not a guarantee of valid CHAT. Inner text/bullet/picture choices and leaves now preserve the same binding; the segment sink no longer accepts or stores independent source text. -
A generated missing-TAB recovery node in a header separator now reports specific E303 instead of generic E342, including multiword header names. Other recovery nodes remain diagnosed; recovered text is not treated as valid.
-
Overlap analysis retains each anchor’s original main-tier span; orphan diagnostics no longer reconstruct it through a second utterance-index lookup.
OverlapAnchorexposesutterance_span(). Its origin and the per-utterance origin are private, so callers obtain these records fromanalyze_file_overlapsrather than struct literals. Matching and diagnostic-location policy are unchanged. -
Prefix-marker language validation no longer guesses a disallowed language when word-language resolution is unavailable. Missing-header diagnostics remain, and an explicit disallowed word language still reports E763.
-
The public
LeafContenttraversal view now distinguishes underline opening and closing markers with their optional source spans. Exhaustive downstream matches must treat both as notation. Underline validation uses the shared structural owner, preserving replacement targets and source locations without maintaining a separate container list. CHAT and JSON formats are unchanged. -
Incremental editor parsing now uses parser-owned source revisions. Cached trees cannot be replaced independently of their source, and the parser derives edits from the actual previous revision before reuse. Raw-tree compatibility APIs retain their caller obligations and recovery checks.
-
Phon
%xphointmedia-bounds validation no longer overflows at maximum timestamps. Its 1 ms tolerance and separate invalid-interval diagnostics remain. -
E704 includes timed untranscribed speech (
xxx,yyy,www) when checking same-speaker overlap, matching CHECK. The 500 ms tolerance is unchanged. -
Media-header lowering retains producer-bound source identity through filename, type and status fields, borrowing checked text without temporary payload strings. Existing recovery and validated filename admission remain in place.
-
Scalar/text headers and the single-value number, recording-quality and transcription headers now retain the same source-bound payload association, with borrowed text and existing malformed-header recovery preserved.
-
E220 rejects bare numeral words even in languages that permit embedded tone or homonym digits. Omission notation and unresolved-language policy are unchanged; numbers must be written out according to their pronunciation.
-
Chinese number spelling no longer duplicates the zero between skipped four-digit groups (
100000001becomes一亿零一, not一亿零零一). -
Desktop keeps read, parse and roundtrip failures visible even without CHAT diagnostics. File details, completion summaries, window titles and notifications no longer misreport these failures or cancelled runs as valid. Concurrent update triggers share one check/prompt/install operation, and a failed native error dialog cannot reject the best-effort update command. The parser selector explicitly labels re2c experimental and incomplete. Plain-text exports now retain every file’s status and failure reason, including failures without diagnostics; malformed export records refuse before writing.
-
W109 now warns when both media and transcript names use the same non-NFC spelling. Disk validation and fixing resolve the stored directory basename, rather than trusting a normalization-equivalent argument spelling. Failed resolution is an explicit I/O error.
chatter fix --code W109 --applycan normalize only the media filename token; it never renames files or changes remote URLs, and a file-side warning may remain after a successful fix. -
@Time Startreports E541 for out-of-range clock components, matching the existing duration bounds: hours 0-23 and minutes/seconds 0-59. Both accepted and rejected values retain their original spelling. Diagnostic rendering requires model-issued refusal evidence. -
Bracketed content owns all inter-item spaces. A leading timing bullet no longer inserts a spurious space after an opening bracket; standalone bracketed-item serialization emits the bullet payload without a separator.
-
Main-tier timing bullets before any utterance material now report E770, including bullets inside retraced groups and after linkers. Explicit zero, words, events and pauses establish material. Parse-recovered main tiers do not produce this absence claim; timing evidence is retained.
-
Main-tier lowering preserves internal timing bullets before a terminator and alongside a distinct terminal bullet. Unterminated tiers transfer only their final bullet to the terminal slot, retaining earlier timing scopes instead of silently discarding them. Missing-terminator diagnostics remain unchanged.
-
Media filename comparison uses explicit equality, mismatch and canonical- equivalence outcomes. Unicode normalization advice always identifies an actual noncanonical side; existing filename-matching behavior is preserved.
-
Standalone slash words now report E243, including nested and replacement words. Repetition annotations
[/]and free-text%com:slashes remain valid; validation preserves the source rather than guessing a repair. -
Generated CST extraction preserves source identity through document recovery, child slots and choice projections. Document/header dispatch, the utterance entry point, dependent-tier dispatch and participant lowering use these capabilities; participant text reads no longer accept an independent source string. Recovery states and range admission remain checked.
-
Generated single-node CST choices preserve an admitted source range through variant selection. Dependent-tier attachment admits its choice once and shares that proof with dispatch and parse-health classification, replacing per-variant range checks without changing recovery policy.
-
Compound validation requires spoken material in every part. Stress or other nonlexical markers no longer hide an empty first, middle, or final part; E232/E233 report the defect without modifying the original word.
-
Unicode ellipsis in word text now reports E243, including nested and replacement words. Original source text is retained; the valid CHAT trailing-off terminator
+...is unchanged. -
Roundtrip difference reports only say “and more” when another difference exists. Present lines are quoted and kept distinct from missing lines.
-
Main-tier semicolons now report E769, matching current CHAT and CHECK 48. Legacy separator syntax remains parseable and roundtrippable; nested and retraced semicolons receive the same source-positioned diagnostic.
-
Removed an unused annotation argument and unreachable exclusion check from extraction’s ordinary-word helper; the scoped walker remains the exclusion owner. Canonical category specs retain extraction/validity separation.
-
Expanded E220 specs with mixed/ambiguous language controls and paired candidate substitutions, preserving the existing digit-permission policy. Scoped examples additionally verify language precedence through NLP extraction, with the observed CHECK discrepancy documented separately.
-
Corrected reserved bullet-rule examples and documentation to preserve the adopted default timing policy; optional CHECK continuity checks remain documented divergences, not implemented Chatter requirements.
-
Timed pauses no longer overflow when converting minutes to bounded seconds. Unrepresentable numeric projections retain their original spelling as unsupported values through CHAT and JSON roundtrips.
-
Comma licensing now follows nested content in document order, reporting commas before spoken content inside groups while retaining omission policy.
-
Cross-utterance quotation/completion checks now consume file-issued positions instead of independent indices, retaining real first/last-position behavior while eliminating silent invalid-index exits.
-
Expanded canonical language-context controls and mutations to verify bare shortcut resolution and unresolved metadata without fabricated fallbacks.
-
Expanded canonical underline controls and marker-deletion cases to grouped standalone markers and nested replacement text, including JSON provenance.
-
@Dateand@Birthreject leading plus signs in day/year components; fixed-width ASCII-digit admission replaces permissive integer parsing. Date construction and JSON decoding share this admission with validation, preserving malformed spellings as unsupported values. -
Replacement words retain their source wrapper span and reject a missing separator after the replacement or its trailing scoped annotations.
-
Bracket-to-word spacing validation now descends into retraces, groups and quotations, retaining sibling boundaries and exact source locations.
-
Timed pauses using minutes and seconds now parse on
%modand%pho, matching main-tier pause syntax. ASCII colons remain invalid in ordinary IPA words. -
Phon reconstruction checks share the alignment mapping’s independent
%mod/%phopositions, preventing false errors after one-sided pauses. -
OverlapMarkerIndex::newis now fallible and admits only 1-9; JSON decoding enforces the same range. The unused post-construction validator is removed. Scoped overlap token decoders reject malformed indices instead of silently changing them into unindexed markers. -
Ambiguous word-language markers validate every candidate ISO code, using the same E519 rules as explicit and mixed markers. Undeclared but valid word-level language codes remain permitted.
-
Opaque postcode labels no longer trigger quotation-balance errors when their text resembles quotation syntax. Actual quotation checks are unchanged.
-
Non-ASCII speaker IDs are consistently invalid. Typed validation and main-tier recovery report E307; recovery no longer mislabels this syntax fault as an undefined speaker (E522).
-
Bracket recovery recommends adding a separator only when its retained source actually lacks whitespace; separated malformed annotations retain their rejection without misleading spacing advice.
-
Invalid-control-character diagnostics acknowledge permitted underline markers rather than describing only standalone CHAT delimiters.
-
@Optionslowering uses generated typed flag slots, preserving ordered supported and unsupported values and retaining empty/missing-name recovery. -
Recovery diagnostics no longer recommend the retired
[x N]repetition notation. Recognizable legacy counts suggest explicit repeated speech with[/]; fragmented bracket errors retain their less-specific diagnostics. -
OffsetAdjustingErrorSinkprojects secondary labels as well as primary locations, clears stale line/column coordinates, and preserves independently indexed source context instead of replacing it based on text length.RebasedErrorSinklikewise clears cached primary line/column coordinates after translating document offsets; unlocated spans remain unlocated. -
The re2c backend retains scoped annotations and ordered retrace chains on quotations, including quotations nested inside other groups.
-
Bare
@Glazy gem markers now parse without a colon or label, as specified by CHAT. Labels remain optional typed groups; recovery diagnostics are retained. -
Splice admission refuses replacements crossing utterance boundaries; a clean starting point no longer licenses changes extending into another utterance. Multi-tier replacements within the same utterance remain admissible.
-
CHAT construction preserves recognized media types, including
missing, and refuses unsupported declared types instead of changing them to audio. Omitted media types retain the documented audio default. -
CHAT construction rejects partial timing pairs and timing attached to empty main-tier text instead of silently discarding supplied timing.
-
CHAT construction applies declared
@Optionsto utterance parsing, preserving CA omission semantics before serialization. Its private build context owns a nonempty language declaration and the contextual parsing operation; header and utterance construction borrow their inputs from that same description. -
Diagnostic contexts decode an omitted
expectedlist as empty, matching their existing serialization. Computed alignment metadata containing such diagnostics can now roundtrip through JSON without inventing parse provenance. -
Rediarization preserves header-only transcripts, including declared participants. Header pruning requires admitted nonempty track evidence, preventing an invalid empty
@Participantsheader when there is no speech. -
Breaking (transform API):
rediarizenow requires a diagnostic sink and rebuilds participant metadata through the canonical header join. Its returned model no longer loses participant records that reappear on parsing serialized output.rediarize_contentrefuses serialization on header-join errors. -
Sanitization uses delimiter-free placeholders for inline events, freecodes and other spoken events, preserving parseable CHAT instead of introducing nested bracket syntax. Documentation now distinguishes the supported fields from complete de-identification and identifies preserved metadata risks.
-
Sanitization now redacts the nine documented free-text header payloads beyond
@Comment, retaining header kinds and speaker references. Exhaustive typed dispatch requires an explicit policy for every future header variant. -
Coordinated
%mor/%grareplacement validates and exclusively borrows the host range before mutation; reversed/empty ranges and short grammatical tiers are refused atomically. Single-item replacement shares the same checked path. Breaking (model API):CoordinatedMutationErroraddsInvalidItemRange; single-item replacement now also refuses donor heads outside the new block. -
Diagnostic highlights stay within their display text at UTF-8 boundaries, including zero-width positions on empty lines and at end of text. Breaking (model API):
PlainDisplayResult::textis private; usetext()or consume the mapping withinto_text()so text cannot drift from offsets. -
Breaking (model API):
UnderlineMarkerstores an optional private source span exposed byspan(). Usefrom_span/with_spaninstead of struct literals; source-independent and JSON-decoded markers now explicitly have no location. Underline diagnostics still fall back to their enclosing word or tier span. -
Underline marker decoding and JSON Schema generation share an explicit wire type, so schema-checked conversion accepts the existing null word-marker payload as well as internally tagged markers without changing serialized output.
-
Language-switch declaration updates retain the selected header’s typed mutable collection instead of indexing and checking its kind again; missing headers and last-header selection keep their existing behavior.
-
Identity language retagging preserves declarations and returns unchanged statistics, including files with span notation that would block a real rename.
-
Media reconciliation retains an exclusive header borrow while proving it is the sole declaration, eliminating a second search and redundant missing-header path while preserving exact duplicate counts.
-
Media reconciliation recognizes internal main-tier bullets through the typed recursive content walker; timed documents can no longer be certified untimed merely because their bullet is not utterance-final.
-
Main-tier recovery collection belongs to its admitted source-bound fragment; callers can no longer supply independent nodes, source strings or offsets.
-
Participant entries own their canonical CHAT serialization; whole headers delegate to the same writer used by standalone fragment clients.
-
Legacy utterance probing uses checked source admission and propagates parser failure; classified input owns its complete envelope or required scaffolding.
-
Internal bullet-text carriers now enter only through generated typed nodes; removed their redundant raw-node classifier and unused raw projection.
-
Unclosed-delimiter findings retain their admitted recovery node and text; diagnostic conversion no longer accepts an independent node or span.
-
Re2c text-tier admission diagnoses and omits all-zero inline bullets, sharing the policy across full-file and offset-aware fragment entry points.
-
E360’s specification distinguishes undelimited text from actual media bullets and includes a parse-backed inline all-zero timestamp rejection fixture.
-
Top-level dependent-tier and unknown-header recovery consume producer-bound source slices; binding now precedes dependent-tier diagnostics and tainting.
-
Generic file-error analysis also requires a bound source slice, including the line-slot recovery route, rather than independently supplied node and text.
-
Document and line-slot recovery share one source-binding diagnostic boundary; a real-node regression rejects foreign trees even with identical source text.
-
Malformed dependent-tier routing admits nonempty labels without manual byte indexing; exact delimiters and conservative unknown-label taint are preserved.
-
Main-tier prefix decoding retains generated kind proofs for MISSING slots; speaker and colon diagnostics consume the corresponding typed nodes.
-
Inline-picture decoding admits a nonempty, delimiter-checked filename before ownership conversion; incompatible source ranges still reject diagnostically.
-
Inline bullet times retain their all-zero-pair policy through a private validated value, with separately parsed timestamp-boundary tests.
-
Phonology fallback retains its generated group type through checked source admission, preserving fallback text and rejecting incompatible ranges safely.
-
Word and main-tier fragments retain generated typed nodes and source identity in one sealed source binding, instead of independent node and slice fields.
-
Main-tier fragment admission derives original input from its parsed envelope, removing independent source/input pairing and checking capacity before allocation.
-
Main-tier rejection carries producer-issued evidence of an emitted diagnostic, eliminating the fragment consumer’s diagnostic-free rejection fallback.
-
Non-colon separators consume generated typed placeholder admission directly; unclassified-placeholder and other recovery diagnostics remain distinct.
-
Marked-token decoding checks source ranges before UTF-8 admission, reporting incompatible sources instead of indexing outside them; marker refusal remains explicit.
-
Language-list recovery retains offending nodes with their grammatical roles and shares one diagnostic constructor for code-slot and repeated-group faults.
-
Participant list and entry recovery likewise carry typed fault roles into a shared reporter, preserving their distinct diagnostic codes and contexts.
-
Sign-tier token decoding shares a typed word/group source transition; whole-group fallback retains its generated kind and checked source boundary.
-
Morphology feature values retain generated kind proofs through MISSING recovery and check source ranges before decoding, preserving refusal diagnostics.
-
Word-recovery fragment diagnostics derive their spans and display contexts from the admitted recovery value instead of independently paired node/text.
-
Generic recovery uses the same admitted fragment diagnostic constructor, preserving full-source and subspan policies at their separate boundaries.
-
Annotated-group recovery retains its typed main-tier body ownership, so a missing form suffix inside a retraced group reports E202 rather than E316. Paired specification seeds and deliberate mutations cover form suffixes and scoped annotations after replacements.
-
Phonology and sign groups use the same typed body-owned recovery path; paired specification mutations preserve E202 for missing form suffixes regardless of which group contains the word.
-
Timing-tier source spans now participate in derived rebasing for parsed, unsupported and empty content, preserving document coordinates through dependent-tier fragment APIs. Corpus-backed tests cover full-line and content-only dependent-tier parsing.
-
Document lowering consumes source-ordered document/recovery parts; duplicate
@Endretains E501 even without a final newline. Reconstructed error wrappers retain generic whole-input recovery for siblings whose document role is unknown. The internal lowering handoff replaces the publicDocumentRoot::into_childrenprojection, which discarded outer recovery. -
Generated repeat selection now retains its extraction cursor through a consuming capability; leading extras and all recovery states remain preserved.
-
Utterance construction admits a readable typed main tier before conversion; incompatible source ranges diagnose and taint the main tier instead of panicking.
-
Utterance recovery also requires readable source admission, conservatively tainting alignment domains when no tier label can be read.
-
CA element and delimiter decoding retains generated node types through a shared checked-source boundary; incompatible ranges reject without panicking.
-
Dependent-tier recovery consumes the shared readable-source carrier before classifying text, retaining its diagnostic families for malformed input.
-
First and repeated language codes share typed slot admission, preserving positional diagnostics and all producible recovery states.
-
Participant slots likewise share admission; only the first-slot position carries the enclosing header needed to diagnose a wholly absent entry.
-
Empty-POS morphology diagnostics retain the recognized token’s occurrence, avoiding an earlier identical split tail when selecting the error span.
-
Generic and word recovery diagnostics share a validated readable-source carrier; incompatible node ranges produce diagnostics instead of panicking.
-
Wrapped-header selection descends through producer-bound source slices. Header admission takes its parsed wrapper and ordinal, rather than an independently selected node; complete-input and header-counting policy remain.
-
Breaking: Generated wrapper/supertype child slots are now
KindSlot. Their MISSING payload retains its proven kind; useknown_or_placeholderfor typed recovery. Raw and composite slots retain fallible classification. Removed foreign-kind-placeholder branches only where the new producer type rules them out; ordinary ERROR, absence and displaced recovery remain. -
Breaking:
DocumentRoot::classifynow accepts a producer-ownedParsedSourcerather than a bare tree. UseTreeSitterParser::parse_source_incrementalto obtain it. Document lowering derives source text from this owner; raw tree access remains available through the consuminginto_treetransition andparse_tree_incremental. -
Other-speaker events consume generated typed slots rather than positional child assertions. Failed transitional text reads reject instead of constructing successful empty values; required-slot recovery remains explicit.
-
Standalone word and main-tier fragments retain producer-bound source slices through admission and lowering. They share the document parser’s 32-bit range admission and
ParseFaileddiagnostic when tree-sitter cannot produce a tree. -
Generated kind classifiers cover nested, kind-disjoint choices. Content recovery uses that producer instead of a hand-written alternative list, and content dispatch retains its typed node rather than rechecking a raw copy.
-
Wrapped header admission derives both its tree and its original-input mapping from one parsed wrapper owner; lowering retains a checked source slice.
-
Shared header field decoding retains typed nodes through checked range and UTF-8 admission. Language-code admission returns the validated model value, without an independently supplied node/text pair.
-
Main-tier separator dispatch retains its generated
SeparatorNode, removing raw-node kind checks and bare-leaf paths outside the producer’s alternatives. Recovery inside separator nodes is unchanged. -
Base-content, pause, overlap and standalone-word lowering retain generated node types through dispatch.
%woruses the generated word-wrapper extractor instead of a positional child read. Word conversion still rejects MISSING placeholders; nested recovery slots remain explicit. -
Bullet-capable text retains its generated carrier and segment choices through lowering, including nested trailing spaces and picture alternatives. Structured bullet readers require
BulletNodeand use generated timestamp fields with closed start/end roles; range failures reject without indexing panics. -
Unclosed-delimiter findings retain the text that established them. Removed a shadowed lone-bracket diagnostic and redundant empty-text checks without changing recovery priority.
-
Main-tier displaced-body reporting derives its sink and grammar context from sealed typed carriers, removing independent node-slice, region and label arguments. Raw slot-level recovery keeps its explicit region policy.
-
Linker decoding uses generated first/repeated groups and exhaustive typed alternatives. Recovered displaced linkers retain source order and token spans; the obsolete raw-kind membership helper and legacy wrapper path are removed. Postcode lowering retains
PostcodeNodethrough checked text admission. -
Main-tier speaker admission retains
SpeakerNodethrough checked text decoding. Its private admitted value keeps nonempty text with its source span; mismatched source ranges reject instead of panicking during prefix conversion. -
User-defined tier taint classification consumes the generated typed prefix slot instead of raw child zero. Only a readable Present
%xmodprefix names the model-alignment domain; recovery retains conservative taint policy.
Fixed
- Main-tier recovery classification admits a readable node/source slice before inspecting content, rejecting out-of-range and split-UTF-8 inputs without indexing panics. Existing marker priority and bracket recovery are preserved.
- Generated traversal cursors consume a bounded remaining-child iterator through their existing typestate transitions. Exhaustion cannot advance past EOF, and final recovery sweeps no longer substitute an empty tail for an invalid index.
- Grammatical-relation field reads reject out-of-source ranges without panicking; field roles retain their generated slot types and admitted indices are nonzero.
- Header-fragment parsing preserves
@Window,@Fontand@Color wordsas typed editor metadata, sharing the document parser’s exhaustive pre-@Begindispatch instead of silently lowering these headers asUnknown.
Removed
- Breaking: removed the experimental
chatter merge,chatter pipeline, andchatter batchCLI commands. Their structural transcript interleaving did not perform fuzzy event correspondence, repair diarization, or reconcile segmentation. Thetalkbank-transformstructural merge APIs remain available; their reporting typestate is unchanged.
0.25.0 - 2026-09-15
Added
- Source-bound donor selection for structural merging preserves original header boundaries after utterance removal or splitting. Selected donor origins map back to original parents; recorded header brackets constrain order without synthesizing utterance timestamps.
- Selected-donor merges accept source-bound relative order constraints
(
RelativeOrderConstraint,with_relative_order), an opt-in timed gem exterior policy (with_timed_gem_exterior; the placements it decided are reported asGemExteriorPlacement), an opt-in flagged draft order that serializes unresolved frontiers reference first with a review comment (with_flagged_draft_order,DraftOrderReview), and header-only references. merge_chat_files_with_donor_selection_draftreturns aMergeDraftbefore validation. Its only edit,set_terminal_bullet, replaces an end-of-line bullet and is recorded as aBulletEdit;Merged::bullet_editsreports the edits aftervalidate.talkbank_model::validationexposesSPEAKER_OVERLAP_TOLERANCE_MSandhas_transcribed_content.
Changed
MergeError::InvalidDonorSelectionreports inconsistent selection coordinates. Downstream exhaustive matches must handle this new variant.- A repeated
@Languagesheader is reported as E501. - Source-order merges serialize utterances with exactly equal starts reference first; section markers at the same instant are refused as ambiguous.
- Merges refuse a missing, repeated or empty
@Languagesdeclaration in either input (MergeError::InvalidLanguageDeclaration), and selected-donor merges refuse a selected child whose bullet lies outside its parent’s. MergeErroraddsInvalidLanguageDeclaration,InvalidRelativeOrder,RelativeOrderTimingConflictandInvalidGemExterior. Downstream exhaustive matches must handle them; the CLI reports each as a precondition (exit 2).- Releases publish only after the desktop installers are built and verified; the draft carries the CHANGELOG notes, and the app banner is added after publication.
- Release artifacts are built with cargo-dist 0.33.0. Its shell installer keeps
the
envPATH helper beside the install receipt (by default~/.config/chatter) for flat installs, and moves an existing helper there.
Security
- rustls is updated to 0.23.45 (RUSTSEC-2026-0285), which rejects TLS 1.3 handshake messages accepted across encryption level boundaries. It reaches the CLI’s self-update and LLM client and the desktop app’s updater.
0.24.2 - 2026-09-12
Added
- An explicit source-order merge policy derives a unique interleaving from source ordering and available timing evidence without synthesizing timestamps. Ambiguous ordering is refused; the existing timed merge policy remains the default.
Changed
MergeErrorincludesAmbiguousUtteranceOrder; downstream exhaustive matches must handle the new variant.- CLI release artifacts are published through an explicit release workflow dispatch.
0.24.1 - 2026-09-11
Fixed
- A
@Mediadeclaration with themissingmedium no longer triggers E544 for absent timing. An explicitly absent recording does not promise linked media; expected recordings still require timing or an appropriate status. - Added a specification example and generated regression fixture for the missing-medium declaration.
0.24.0 - 2026-09-10
Changed
- Structural merge now consumes ordered AST streams rather than collecting headers and sorting utterances. Interleaved body headers and dependent-tier order survive; selected utterances without timing or with reversed source starts are refused. Section markers use neighboring source timing bounds, and ambiguous cross-source placement is refused instead of guessed.
MergeErrorhas new variants for timing, section-placement, metadata-order, participant-join and output-validation failures. Exhaustive library matches must handle them.MergedandReportedretain a validated document;Reported::into_filerelinquishes that proof for subsequent edits.
Fixed
- Documentation-date checks now include pending commit/squash changes before publication. The commit hook checks the actual index, preventing an unstaged repair or an older gate receipt from masking stale staged date headers.
- Merge builds the derived participant map through the canonical header join and validates the assembled AST, including tier alignment, before returning success. Callers no longer need serialization and reparsing to obtain a consistent participant map.
- Donor IDs extend the contiguous opening ID block; donor body comments remain at their source position. An opening ID after a comment is refused rather than silently reordered.
Added
WorSlotMembershipPolicy::admits(&Word) -> bool, the public per-word%woradmission predicate.WorMainTierProjection::from_mainadmits its slots through it, so it is the projection’s own rule rather than a second statement of it. A downstream consumer that counts%wor-eligible words per content item (both Batchalign trees carried a hand copy ofcounts_for_tier(word, TierDomain::Wor)beside awalk_wordsfor this) asks the policy and deletes the copy.
0.23.0 - 2026-09-09
Removed
walk_overlap_pointsandOverlapPointVisitfromtalkbank_model::alignment::helpers: a visitor over overlap markers with no caller in any tree, carrying a third private walk of its own (and a position convention for intra-word closing markers that differed from the collector’s).extract_overlap_infois the one API.talkbank_transform’s corpus discovery and manifest API (discover_corpora,build_manifest,corpus_summary,format_manifest,CorpusManifest,CorpusEntry,FileEntry,CorpusFileStatus,FailureReason,ErrorDetail,ErrorLocation,ManifestError). No command in this repository and no known dependant used it; its only caller was its own test.TierContent’sto_content_string_no_bulletsandwrite_tier_content_no_bullets,ValidationContext::with_quotation_validationandwith_bullets_modewith thebullets_modefield they set (thebullets@Optionswas removed from CHAT and the flag was always false, so E362’s check on bullet monotonicity now simply runs), and the parser API’s always-falsebullets_mode(); none had a caller anywhere. Breaking for a library user who called any of them.- The
Utterancebuilder helpers with no caller:with_preceding_headers,with_user_definedand every per-tierwith_*exceptwith_mor,with_gra,with_sinandwith_com(add_dependent_tieris the one route they were sugar over and remains); the semantic-diff renderersshort_summary,short_summary_with_source,render_with_source,render_comparison,render_comparison_shortandrender_tree_diffonSemanticDiffReportwith theRenderModethey took and the tree renderer behind them, and theUtteranceaccessorsmor,gra(the cloning aliases;mor_tierandgra_tierstay, as dophoandsin),mor_tier_mut,computed_language_metadata,wor_alignable_word_count,pho_alignable_word_countandsin_alignable_word_count. None had a caller in this repository’s root workspace, in talkbank-tools or in the downstream Batchalign, and a whole-workspace coverage run showed every one unreached; the two cloning aliases had one caller in the spec runtime tools (a separate workspace), moved to the borrowing accessors. Breaking for a library user who called any of them;SemanticDiffReport::render(also itsDisplay) remains, and the alignable count that is used,mor_alignable_word_count, remains. TreeSitterParser::parse_utterance_cst. It forwarded to the freeparse_utterance_nodeand had no caller in this repository or in any repository known to depend on it; a whole-workspace coverage run showed it unreached. Breaking for a library user who called it; the public routes areTreeSitterParser::parse_utterance(one utterance from its text) andTreeSitterParser::parse_utterance_fragment.talkbank_parser::parse_dependent_tier(the free function) andtalkbank_parser::tiers::parse_mod_tier_from_unparsed. Neither had a caller in this repository or in any repository known to depend on it. The first returned an untypedUserDefinedTierfor every tier, where thetalkbank_model::ChatParser::parse_dependent_tierroute returns the typedDependentTier; a whole-workspace coverage run showed both entirely unreached,%modparsing having gone through the typed%photier parser for as long as the typed traversal has existed. Breaking for a library user who called either; use thetalkbank_model::ChatParsertrait’s methods.
Added
talkbank_model::content::word::Word::new(NonEmptyString, WordText): the checked constructor, taking the two proofs a word’s texts must carry. The tree-sitter parser builds through it;new_uncheckedremains for test support and the front ends not yet migrated.
Changed
-
talkbank_parser::generated_traversal::NodeSlotisNodeSlot<'tree, T, M, U, A>: the payloads ofMissing,UnexpectedandAbsentare the node (orNoChild) where the position can produce the state and the uninhabitedNeverwhere it cannot, and every generated accessor names its position’s kind through one of four aliases (ChildSlot,SeqSlot,ChoiceSlot,ClassifiedSlot).AbsentcarriesNoChild;RecoveryandSlotValuecarry the same parameters;NodeSlot::viewgives a borrowed slot’s states back by value as aSlotView, so a match through a reference can omit the arms the position kind rules out. Breaking for a library user matching the slot directly: an arm for a state the position cannot produce no longer compiles, which is the point. The parser’s own hand-written arms for those states, each carrying a diagnostic for a case that cannot happen, are gone with it. -
A content-bearing recovery node inside a
%mortier is E702 at any depth. It was E702 for a direct child of the tier and E316 for a node below one, because the utterance parser walked every dependent tier’s children before attaching it and reported them without the tier’s name, and the typed dispatch then reported the direct children again: a%worline with an unparsable word carried the same E316 twice at the same span. One reporter now walks the whole tier in the tier’s words. Three E316 examples and E711’s first are E702’s (subsumed_by), E702’s first example is aviolatesclaim, and the E342 text for a MISSING node inside a tier names the tier (“in gra tier”). -
Counting and extraction take
PositionalDomain(Mor,Pho,Sin), a new type with noWor:count_tier_positions,count_tier_positions_until,collect_tier_items,TierCountable,AlignableTier::DOMAIN, andtalkbank_transform::extract::{extract_words, collect_utterance_content}.TierDomainkeepsWorand stays the vocabulary of the walkers and ofcounts_for_tier;PositionalDomainconverts into it, andTryFrom<TierDomain>refusesWorwithNotPositional. The%worcount and pairing areWorMainTierProjection’s, and the twoWorarms in the counter and the one in the extractor were a second implementation of that count, agreeing with the projection only by test; a probe over every reference-corpus file and every spec example found them equal before they were deleted. The overlap-marker position walk inalignment::helpers::overlapstill counts on the%worscale with its own traversal and is not changed here.WorMainTierProjection::slotsis crate-private (pair throughbind_timing). A downstream caller passingTierDomain::Morto any of these writesPositionalDomain::Mor; one asking a%worcount callsMainTier::wor_projection().slot_count(). This is a breaking Rust API change. -
Two validity rulings (maintainer, 2026-09-08), both grounded on real CLAN CHECK. A bullet INSIDE a main-tier utterance is timing evidence for the media-consistency family:
hello \u{15}100_200\u{15} world .is E752 without an@Mediaheader (CLAN CHECK 112 fires on it) and satisfies a declared@Media(it was E544 before, and valid without the header). And whitespace-only content on a bullet-payload tier (%com,%add,%exp,%gpx,%int,%sit,%spa,%act,%cod) declares nothing: E756 beside E758, as%engalready was (CLAN CHECK 31 rejects the same lines). Transcripts that relied on either gap now validate differently. -
Diagnostics inside an angle group, a quotation, a pho group or a sin group are now the ones the tier body gives for the same material. The four constructs parsed their contents through a second, hand-written walker with its own generic ERROR analysis; the one typed
contentswalker serves both now. Visible changes: a curly single quote inside a quotation is E256 (it was E331); a stray[inside a construct is E316 at that byte (it was three E330 “expected X” messages); an ERROR fragment that opens a bracket or parenthesis and never closes it is E312 or E313 on the tier body as well as inside a construct (hel(lo .was E316; a parenthesis that opens the whole utterance still fails at file level), and “never closes” means no closer anywhere in the fragment, not merely not at its end. E331 (UnexpectedNodeInContext) has no known route from CHAT input any more and is recorded as unreachable. -
WordLengthening::count, itswith_countargument, and re2c’s AST lengthening count now useNonZeroUsizeinstead ofu8. JSON keeps the integer field and its omitted-one default, accepts longer runs, and rejects zero. This is a breaking Rust API and JSON-admission change.
Fixed
-
The phonological tiers accept superscript one, two and three (U+00B9, U+00B2, U+00B3), the only superscript forms those digits have, wherever the other superscript digits were already accepted, and the last three modifier tone letters of their block (U+A71D to U+A71F, the raised and low exclamation-mark letters,
ꜞamong them) as the rest of the block already was. A%phoor%modword carrying one was E316 unparsable content (TalkBank/chatter#6, #7): 52 of 52 sessions of one tone-language corpus, 18 of 96 in a second, 3 of 58 in a third, and 2 in a fourth. -
E715 and E734 no longer report a
%phoor%modtier one token long when the main tier carries a pause inside a<...>group: a pause was counted as a phonological token at the top level of the utterance but not inside a group, while the phonological tiers carry it in both places, as the Phon team’s French corpora show at scale (72 records in 49 sessions of one corpus, every one a pause inside an overlap or retrace group; TalkBank/chatter#5). The counting walker’s own 2026-08-08 note had held that arm open for exactly this evidence. -
E740 and E741 no longer report a
%modor%phoword that carries the linking tie‿(U+203F) as a mismatch against its%xphoalnreconstruction: the tie joins two symbols into one segment or marks the absence of a break, is never a phone, and has no alignment column, so the source word is compared modulo the tie as it already was modulo the stress and syllable-boundary marks (a pair side carrying a tie is not a bare segment and is compared as written). TalkBank/chatter#3, with the inventory in #4: until now every such word in two PhonBank corpora was a spurious mismatch. -
A group with no code after it (
<w> .) is E342 alone, and the model keeps the bare group that was written. The grammar requires a code there, tree-sitter inserts a MISSING placeholder, and the annotation decoder used to read the placeholder’s kind and build a full retrace nobody wrote: the validator then reported E757 and E370 against constructs the file does not contain, and the E342 spec example’s roundtrip diverged. The decoder skips MISSING nodes now. CHECK-parity for CHECK 51 expects E342; the E342 example leaves the backend-parity baseline, both parsers agreeing. -
E710 is reported only by the
%grarelation parser. The dependent-tier recovery analyzer had a branch that fired on the substring%gra:anywhere in an ERROR node’s text and called it E710, “non-numeric index”: an%engor%xbody mentioning%gra:, or junk after a well-formed%grarelation, was reported as an invalid relation. The branch is gone; such a node is the generic E316 (E258 for a double comma, E760 for a%moritem with an empty part of speech, as before). The E760 branch’s own gate accepted%mor:anywhere in the text for the same reason and now needs the line, the tier context, or a text that starts with the prefix. -
A MISSING node is reported once. The whole-tree recovery backstop suppresses a candidate already covered by a region diagnostic by span overlap, and widened only its own zero-width MISSING span to a byte, so a region’s E342 for the same point (itself zero-width) never covered it: every MISSING node inside a dependent tier or a header list carried two E342 texts at one span. A zero-width region diagnostic at the same point now covers the candidate when it reports the same code; a different code there (E376 for an empty replacement beside its MISSING word segment) still leaves the E342 reported, as E208.md documents.
-
Overlap-marker positions count a replaced word inside a group once. The collector behind
extract_overlap_info(and sotop_onset_fraction,estimate_onset_ms, and the cross-utterance E347 and E704 checks, which read the paired positions; E373 reads only the indices and was not affected) walked with two private traversals, and the bracketed one scanned a replaced word’s replacement words too, so<doggie [: dog]>under a marker counted two words where the%worprojection counts one and every later marker position, and the onset fraction, drifted. The collector now walks with the sharedwalk_contentat the%wordomain, the projection’s own leaf set, sototal_wordsis the projection’s slot count by construction; a snapshot of every marker-bearing utterance in the reference corpus and the spec examples was byte-identical across the change, and the group case is pinned. -
chatter debug sanitizeredacts%actand%codtiers. Both carry the same bullet payload as%com, and passed through the strict sanitizer with their text intact under a comment deferring their redaction; an action line is free text about the participant. A parse-backed table of every dependent-tier kind through the sanitizer is the pin. -
chatter debug sanitizeredacts the words on%wor. The tier repeats every main-tier word beside its bullet and passed through untouched, so a sanitized file with a%wortier still carried the whole utterance in clear. The tier is now judged the way timing recovery judges it, before the main tier is rewritten:WorMainTierProjection::bind_timingfor the counts, thencorroborate_wor_timingfor the words. A tier that corroborated the main tier has each word rewritten as its paired main-tier word’s display text, now that word’s placeholder (w1w1for a compound), so it corroborates the sanitized main tier exactly as before; a tier that drifted in count, or carried a word the main tier did not, takes fresh placeholders rather than a manufactured agreement, and disagrees after as it did before. Bullets are kept byte-exact, and sanitizing the output again reproduces it. The test that claimed to pin%woroffsets wrote them as bare1000_1100tokens, which the model parses as words; it passed only because nothing touched the tier. It now pins whole lines, drift and compounds included.CountMatchedWorTimings::pairsexposes the owner’s pairing. -
chatter validate --roundtripcounts a file’s roundtrip on every run. When the file’s validity and roundtrip verdicts were both served from the cache, the summary saidPassed: 0for a file whose roundtrip had passed, andFailed: 0for one whose roundtrip had failed: the cached branch built the file’s status and never touched either roundtrip counter. The status now records whether the roundtrip ran (FileStatus::Valid { roundtrip: RoundtripVerdict }, a new public field and type) and both counters are derived from it. -
E370 (a retrace marker with nothing after it to retrace) is labelled at the marker’s own bytes whatever the spacing around it. The rule used to find the marker by rendering the main tier back to CHAT text and taking the offset there, which was right only when the source was already canonical: on
<hello there> [/] .the label sat one byte early. The parser now records the marker token’s own span on the retrace (Retrace::marker_span, a new public field,Nonefor a retrace built without a source) and the rule reports there. -
Overlap markers inside an angle group, a quotation, a pho group or a sin group are kept in the model and written back. The constructs’ old contents walker handed each item to a second walk over the item’s children, and an
overlap_pointis a single token with none, so<hello \u{2308} there \u{2309}> [/] hello there .parsed clean, validated clean and wrote back with both markers gone. -
%grastructural diagnostics (E721, E722, E723, E724) no longer fire on a tier the parser had to shorten. A relation the model cannot hold is rejected and dropped, and the rules for sequential indices, root count and cycles describe the graph the author wrote, not what survived; E722 in particular reported “no ROOT relation” against a tier whose only surviving relation was a ROOT. A newJudgeableGrawitness is the sole route to those rules and asks both questions that decide it, parse recovery and prior alignment findings. -
The re2c backend records parse recovery on a dependent tier it could not build as written: a
%grathat lost a relation, a%morwhose conversion failed, a%worbody it could not re-lex. Cross-tier alignment previously compared such a tier against its neighbours and reported the difference its own recovery had created, so%morand%gracounts disagreed (E720) on a transcript where they agree. -
Both parsers preserve lengthening runs beyond 255 colons without integer overflow or truncation. The default model marker now consistently contains one colon, with no zero-count repair during serialization.
-
Re2c retains malformed form suffixes for specific E202/E203 diagnostics; repeated dangling markers no longer produce both errors for one defect.
-
Release lint checks application-version synchronization before compilation.
0.22.0 - 2026-09-06
Changed
talkbank_lsp::backend::utils::LineIndexborrows its source. Itsoffset_to_positionmethod accepts only the offset, preventing callers from pairing indexed line starts with another text. This is a breaking Rust API change.
Fixed
- LSP edits, formatting, semantic tokens, selection ranges, symbols and quick fixes consistently use UTF-16 coordinates. Multiline semantic captures split into individual lines, and whole-document formatting includes the final newline.
- Gem outlines use parsed header spans and matching labels, preserving CRLF positions and counting only actual utterances in the parent outline.
- Language-service initialization retains its result in
OnceCell; nested highlighter access returns an error instead of panicking. Execute-command services accept only their own request enums, removing routing panic branches.
0.21.0 - 2026-09-06
Changed
-
LSP backend cache fields are replaced by a private source-bound analysis. The unused public incremental-splice and validation-cache modules are removed. Syntax reuse remains incremental; models and validation results are rebuilt together using the shared model validator.
-
DocumentRootis a private-field classification with method accessors rather than a publicly constructible enum. It owns both document lowering and whole-source diagnostic scope. -
re2c parsed header lines carry
HeaderProvenancein place of a standalone separator field, and box their header payload. The owned lexer extent now reaches model header spans; boxing keeps the file-line enum compact. -
re2c pause tokens and parsed pause variants retain a
PauseLexemeinstead of discarding the full lexical extent. This changes their Rust payload types.
Fixed
-
A truncated document without a final newline retains its complete simple final main tier. Shared terminal recovery reuses the normal fragment parser and preserves caller coordinates while validation reports missing
@End. -
LSP diagnostics after edits now agree with fresh-open text, including deleted headers, recovery suffixes, Unicode edits and skipped debounce revisions. Tree-sitter edits use the cached tree’s own source and UTF-8 byte coordinates. Feature and pull-diagnostic requests cannot reuse another revision’s spans. Published diagnostics carry editor versions; obsolete analyses are discarded.
-
LSP validation includes shared file-level rules such as E752 instead of a separate incomplete validation sequence. Protocol regression tests share the existing executable harness, preserving interleaved responses/notifications.
-
Recovery before or after a complete document receives localized diagnostics without discarding the document. A complete final main tier stranded outside its line wrapper when
@Endis missing is retained through normal utterance construction. -
Documents missing
@UTF8retain their headers and utterances and report E503, without cascades claiming present headers are absent. Both parsers locate the diagnostic at the end of the file, including rebased fragments. -
Both parser backends admit complete fragment coordinate ranges before parsing. Origins above 2 GiB retain correct model and diagnostic spans; overflowing 32-bit ranges are rejected instead of truncated. Synthetic wrapper text no longer consumes the caller’s document range.
-
re2c header fragments reject extra headers, utterances and unsupported trailing lines instead of returning a partial result. Lowering consumes an admitted logical header; folded content and recovery diagnostics survive.
-
re2c preserves pause spans, including timed-pause parentheses, through nested content and fragment rebasing. Shared validation now owns pause spacing; duplicate token scans are removed.
-
Separator spacing validation visits nested groups, reporting E765 at the missing space just as it does for top-level content. Spaced group controls remain valid.
0.20.2 - 2026-09-06
Fixed
- Single-header parsing rejects extra headers instead of silently returning
only the first. Lowering consumes a complete
HeaderFragmentthat owns the selected node and its source; folded content and LF/CRLF remain accepted. - Standalone header and dependent-tier parsing derive diagnostic coordinates from their owned synthetic source rather than separately supplied prefix lengths. Header lookup failures carry the caller’s text and document origin instead of empty context. Existing malformed-header diagnostics are retained.
0.20.1 - 2026-09-06
Fixed
- The standalone LSP exits after the editor’s
exitnotification even when the editor keeps stdin open. A completed shutdown permits exit code 0; exit before a successful shutdown returns code 1. Protocol completion now stops the transport, and runtime teardown does not wait on Tokio’s uncancellable stdin reader. Process-level regression tests cover the actual binary, rejected shutdown, early exit, and EOF.
0.20.0 - 2026-09-06
Changed
-
Breaking: removed
CachePool::open_or_else; useCachePool::newand handle itsResultdirectly. Cache opening no longer splits failures between an optional handle and a callback; the CLI retains the concrete opening error. -
Validation caches include both parser implementation source fingerprints, closing stale verdict reuse after parser-only edits without a version bump. Shared build-only source hashing reads each crate’s own packaged files.
-
Breaking: re2c
Token::TierPrefixcarriesDependentPrefixToken, including the lexer-selectedDependentBodyKind; dependent parsing matches that enum. -
Breaking:
CacheStats::cache_diris optional: in-memory storage has no filesystem directory. File-backed statistics retain the opening directory. -
Breaking: validation-cache constructors require
CacheIdentity(rules and parser), and roundtrip cache methods use that bound identity instead of a parser string.ParserKindis shared fromtalkbank-modeland re-exported.MaintenanceCacheexposes administrative operations without verdict methods. -
Breaking: re2c prefix tokens carry
PrefixTokenpayload/separator state. AST header lines, main tiers andDependentTierEntryParsedretain separator provenance. Access a prefix payload withtext(); file AST snapshots reflect the new header and dependent-entry shapes. -
Breaking: re2c dependent-tier AST adds
RejectedMor, retaining raw input after failed morphology admission without fabricating a model tier. -
Breaking: re2c
Token::TierPrefixdenotes a complete colon-tab prefix; the newIncompleteTierPrefixvariant identifies recovery from a bare label. -
Breaking: re2c postcode tokens and main-tier AST postcodes carry checked payload state and lexer locations.
main_tier_to_modelandutterance_to_modelnow require an error sink; callers can no longer lower these structures without deciding where recovery diagnostics go. -
Breaking: re2c AST
ParsedAnnotationseparatesScopedannotations from retrace, replacement, language-code and postcode structures. MatchParsedAnnotation::Scoped(ScopedAnnotationParsed::...)for scoped kinds; their conversion to model annotations is now total. AST inspection snapshots reflect this category; serialized CHAT model output retains its shape.
Fixed
-
Release bumping updates and checks both desktop npm lockfile version fields. Dependency versions remain unchanged; CLI regression checks run in the fast app-version gate.
-
Owned fragment wrappers project secondary labels with primary locations and identify synthetic context by exact source text instead of a length heuristic. Independent diagnostic context is preserved even when longer than the input.
-
Main-tier fragments reject trailing material instead of silently accepting their first tier. Aggregate spans exclude a synthetic final newline, while retaining caller-supplied LF and CRLF line endings. Root-admission diagnostics now describe the required source shape without raw CST-kind wording.
-
Fragment APIs no longer subtract obsolete word/main-tier wrapper lengths or subtract caller offsets from diagnostics. Synthetic utterance, participant and dependent-tier wrappers own their input boundary for model and error projection. Complete CHAT documents are recognized by the utterance adapter.
-
re2c reports unsupported lines as E326 with their original source spans and preserves following utterances. Diagnostic rebasing keeps context highlights relative to their own source text.
-
Speaker-qualified
@Birth of,@Birthplace of, and@L1 ofheaders retain their separator spans, including non-CA whitespace violations and CA exemptions. -
Tree-sitter’s whole-file fragment API now rebases model spans along with streamed diagnostics when parsing embedded CHAT at a nonzero offset.
-
Tree-sitter recovery no longer reports E758 for spaces after rejected dependent-tier or header content. Separator provenance requires adjacency to the actual tab, preserving the original content diagnostics.
-
Gate receipts verify the actual committed trees in every pushed ref, including annotated tags. Uncommitted fixes cannot authorize an older commit, and a gate whose source changes during verification cannot issue a receipt.
-
re2c E760 highlights the original empty-POS morphology item, including across continuation lines and non-ASCII text. Recovery inspects source-owned items without rebuilding rich-token payloads or using a dummy diagnostic span.
-
re2c no longer dispatches longer Phon labels through
%modor%phobody parsers. Bare and x-prefixed syllabification, alignment and interval tiers retain their own grammar without false E316 diagnostics. -
Cache statistics report the directory actually opened instead of resolving the current default again, including for explicitly located maintenance pools.
-
Validation cache rows are isolated by parser in both CLI and desktop. Switching parser/rule combinations no longer risks serving another parser’s verdict or consumes extra retained generations. Maintenance opens do not prune rule generations or expire rows merely to display statistics.
-
re2c records trailing separator spaces across headers and tiers. The shared validator reports E758 outside CA, and serialization canonicalizes separators in either mode. This removes the separate main-tier scan and CA probe.
-
re2c enforces morphological lemma starts and nonempty features at lexing. Rejected
%morreports E316/E600 and preserves morphology taint instead of converting to an unsupported tier with E605. -
re2c rejects a replacement glued to its word with E375/E316, using the original bracket locations. A spaced replacement remains valid. The canonical parser’s malformed closing-bracket highlight excludes absorbed trailing whitespace and uses the original source for its diagnostic context.
-
re2c reports E602 for malformed dependent-tier separators even when content follows the label, including a space in place of the required tab. Recovery uses the lexer-classified prefix and locates the complete malformed line; empty content no longer participates in deciding whether its prefix is valid.
-
re2c rejects whitespace-only postcodes with E363 while preserving the rest of the tier. Valid postcodes preserve leading payload whitespace and trim only trailing whitespace, matching the canonical parser. Postcode diagnostics retain the complete token span through file, utterance and main-tier APIs.
-
re2c utterance fragments forward parse diagnostics, and file diagnostics honor the caller’s offset. A shared streaming adapter replaces temporary diagnostic collectors in header and participant fragments.
-
re2c reports E757 when rich bracketed annotations are glued to the following word, including
[!]thereand[= toy]there. The check uses the parser’s annotation categories and reports the following word’s original lexer span.
0.19.0 - 2026-09-05
Changed
-
Breaking: the LLM response cache has one owning handle per path across processes. Share the handle across threads and drop it before reopening.
CacheErrordistinguishes a busy cache and a visible replacement whose final directory sync failed. -
Breaking: removed
SinToken::new_unchecked; use checkedSinToken::new.SinTier::from_tokensnow returnsResult<SinTier, EmptyText>. re2c’sSinTierParsedandSinItemParsedown checkedSinTokenvalues and no longer take a source lifetime parameter. -
Breaking: re2c AST word
raw_textis nowCow<str>, distinguishing borrowed source from owned reconstructed text. Parser combinators separate source and token-storage lifetimes. -
Breaking: re2c’s
sin_tier_from_textreturnsParseOutcome<SinTier>; malformed fragments can no longer appear as successfully parsed empty tiers. -
Breaking: diagnostic enrichment takes a source-bound index. Replace
enhance_errors_with_line_map(errors, source, map)withenhance_errors_with_index(errors, &SourceIndex::new(source)), or retain oneSourceIndexfor repeated batches. Its immutable borrow prevents source edits while the index is used; callers cannot pair another file’s line boundaries with the source and trigger a UTF-8 slicing panic. The existingenhance_errors_with_sourceconvenience API retains its signature.
Fixed
-
LLM response-cache writes now prepare, flush, and atomically publish snapshots before updating memory. Failed writes preserve old entries; concurrent puts cannot publish stale snapshots. Unix builds also confirm directory durability.
-
Both parsers now share source-bound control-character checking before parsing. re2c no longer silently accepts forbidden controls in free-text tiers; the lexical diagnostic retains its original source and exact byte range.
-
E212 spec coverage now demonstrates its reachable CA-mode word-category boundary, alongside legal controls. Its existing implementation is marked implemented, replacing the misleading legal-only deferred fixture.
-
Utterance validation now reaches bare and grouped
%sintokens admitted through JSON, reporting empty text at the tier’s span. Both parsers construct tokens through the checked constructor; the duplicate unchecked path is gone. -
re2c parsing no longer leaks copied source, token arrays, recovery buffers, or reconstructed words. The lexer safely handles unpadded input at EOF, and word-fragment conversion retains spans from the caller’s original source.
-
re2c
%sinfragment parsing now uses the whole-file grammar, preserving single-token gesture groups and rejecting unclosed groups with a diagnostic. The duplicate whitespace parser and its independent group state are removed. -
Diagnostic line/column lookup no longer caches by source address and length, which returned stale positions after same-length edits or allocation reuse. One-off lookup scans without allocation; indexed batches retain logarithmic lookups without retaining a hidden source copy or thread-local cache.
-
Foundation publication checks now cover every workspace crate outside the approved first wave.
talkbank-llmis explicitly held back; a metadata-only mode checks manifests and dependencies without packaging or registry access. -
Generated fixture and documentation directories now retain unchanged files and prune only obsolete output through an ownership capability. Conflicting ownership, nested human content and symlinks are refused before pruning.
-
Spec regeneration preserves unchanged outputs in shared directories, including generated model code and Rust test bodies, while still removing explicitly retired files. Progress counts now report actual writes.
-
Tree-sitter generation now stages all grammar artifacts and preserves unchanged files. The grammar currency check no longer rewrites source files, and also checks generated C headers.
-
Node-type, traversal and conformance-inventory regeneration now preserves unchanged files and publishes changed output only after the generator succeeds. A failing generator no longer truncates those committed Rust files.
-
Removed an unnecessary schema rewrite that treated valid Draft 2020-12
$refsiblings as invalid and could modify literal schema data. Generated schemas now retain schemars’ structure; enum tags and referenced payloads remain jointly validated. -
Schema generation now preserves unchanged files and runs only when explicitly requested by
just schema-genorjust regen. Ordinary tests no longer rewrite a compile-time dependency and trigger avoidable recompilation. Schema currency failures report the repair command without dumping the complete schema. -
to-json --skip-schema-validationretains the transcript name and requested CHAT checks, including E531 for mismatched media filenames. Single-file and directory conversion now share one named parsing path before schema policy selects serialization. -
Directory JSON conversion prints individual parse/validation diagnostics and exits with failure when any file fails; successfully converted sibling files remain available.
Added
JsonSchemaPolicyandchat_to_json_with_schema_policylet library callers select JSON Schema validation independently of CHAT validation and transcript identity. Existing conversion functions retain their signatures.
0.18.1 - 2026-09-05
Added
- Timing-producing transforms now have a typed media-link transition.
reconcile_media_timingconsumes aChatFileand returns either anUntimedChatFileor aLinkedMediaChatFile. A timed document must have one usable@Mediadeclaration; the transition removesunlinked, accepts an already-linked declaration, and returns typed errors for missing, ambiguous, or contradictory media. Both states expose only an immutable document and post-transition serialization. This prevents a forced-alignment pipeline from writing fresh timing bullets while retaining the contradictory@Media: ..., unlinkedstatus rejected by E552. Validation verdicts are unchanged.
0.18.0 - 2026-09-04
Changed
-
Warm development tests no longer scan unpacked macOS codegen objects. The root and specification workspaces embed line-table debug information in linked artifacts instead of retaining every
.rcgu.o. This bounds the file count in both target directories and removes filesystem enumeration from the warm-test path while preserving source locations in diagnostics. Nine spec generator commands are also excluded as empty libtest harnesses; their library and integration tests remain in the suite. -
Breaking: validation returns owned evidence rather than a mutable phase marker.
ChatFileis no longer generic.validate_intoreturnsResult<ValidChatFile, ValidationFailure>; accepted payloads are read-only, errors retain the rejected model, and unknown/recovered tier provenance cannot pass.validate_with_policyrecords rules, alignment coverage and transcript name.parse_validated_with_parseradditionally requires error-free source parsing. Remove the oldNotValidated/Validated/ValidationStateimports and consumeinto_unchecked()before editing an accepted document. Serialized transcript fields remain unchanged.Required-validation compatibility APIs now use the same proof-producing transition before returning an explicitly editable model. Their streaming variants return
Errafter parse or validation failure even when the caller’s diagnostic sink discards messages. Merge preflight retainsValidChatFilewhile reading the accepted reference and borrows its document for the merge. -
Breaking: utterance builders can retain an utterance comment.
UtteranceDescaddscomment: Option<ComTier>; Rust struct literals must supply this field. A supplied comment is emitted as%com. A comment on an empty utterance is rejected instead of being silently discarded. -
%pho,%modand%sincount mismatches are reported by one algorithm. The utterance metadata path used its own copy of the positional alignment with the diagnostic codes passed in as parameters; it now uses theAlignableTierroute, which reads the tier’s own type (%phoor%mod) to choose E714/E715 or E733/E734. The codes are unchanged; the messages are the positional form the%sinroute already used (a per-position table instead of a bare count). -
A control character is a lexical error anywhere in the file. E315 is now decided over the whole input before any parse, so a forbidden control character in a word, a
%comline or a header value is reported at its own offset; previously a word’s surfaced as generic E316 and free-text tiers accepted it silently. Permitted are TAB, LF, CR, the bullet delimiter U+0015 and the CA underline pairs U+0002 U+0001 / U+0002 U+0002; CLAN’s italics pairs are reported (CHECK 102). -
E303 covers every header whose colon is not followed by a TAB. It used to fire only for
@Comment:, every other header fell to E316, and the message said “space” whatever followed the colon; it now names what was found (a space, nothing, or the character). -
Wordlexical content is read-only outside its owning type. Direct access to the former publiccontentfield is replaced bycontent(); callers that intentionally replace typed content use the named mutation APIs, which invalidate derivedcleaned_text. This prevents a content edit from leaving stale lexical text in JSON. Direct crate-internal access toraw_textis closed as well, so recovery spelling changes use the explicit setter rather than bypassing the field boundary. -
Speaker-code structure now has one typed assessment across every model surface. Direct
SpeakerCode::validatepreviously mislabeled an overlong code as undeclared (E308), mislabeled a reserved character as a missing CST node (E302), and enforced a different character policy from headers and main tiers. All three routes now consume the same producer-issued valid/invalid state and report E307. The seven-character limit counts Unicode scalar values rather than UTF-8 bytes, and diagnostic context records the offending code rather than an internal field label. -
Malformed regions now distinguish unpaired CHAT quotation delimiters from unrelated parser recovery. Structurally unpaired
“or”delimiters report E242 even when tree-sitter encloses them in a larger error node. Balanced quotation delimiters inside some other malformed region no longer produce a false E242, and ASCII straight quotes are not mislabeled as CHAT quotation delimiters.
0.17.0 - 2026-08-30
Changed
-
Reference-mode speaker identification now preserves typed lexical support.
DonorMatchReportretains the reference, donor, shared, and union token counts that derive each Jaccard score; its winner, evidence, and confidence margin are no longer independently constructible public fields. Thresholds are checkedConfidenceThresholdvalues, while confidence is explicitlyNoInformation,Finite, orUnboundedinstead of overloading0.0and infinity. Use the read-only report accessors in place of direct field access.chatter speaker-id --write-match-report NEW.jsonwrites the accepted, low-confidence, structural-refusal, or input-refusal evidence without replacing an existing report. -
%wortiming is now admitted through explicit typed evidence states.MainTier::wor_projection()defines the shared positional membership policy; count binding, canonical-token corroboration, and complete positive interval assessment are separate states, so equal word counts alone cannot be treated as trustworthy timing. The impossible tier-levelWorTier::bulletfield is removed: timing evidence exists only when an actual%worword carries a bullet. Callers of the former alignment and tier-bullet APIs must migrate to the projection, binding, corroboration, sequence-assessment, andWorTier::timing_evidence()APIs. -
rediarizeandrediarize_contentrequire aDiarizationTimeline. The windowed overlap algorithm needs turns ordered by start time, but the former&[DiarizationTurn]API let every library caller bypass that precondition and silently obtain a wrong winner.DiarizationTimeline::newowns the sorting transition, retains the longest-turn window bound, and keeps its ordered storage private.TurnsFilenow exposessource()andtimeline()accessors instead of independently public fields. -
Tree-sitter 0.27.0 now drives the Rust parser, highlighting runtime, grammar crate, spec tooling, and grammar-generation CLI. The Node binding is independently current at 0.25.1. Generated parser artifacts and the complete parser/backend parity gates are regenerated and checked with this toolchain.
-
The vendored Rust lexer is regenerated with re2c 4.6.
re2c-version.tomlis now the exact generator source of truth, andjust verify-vendored-lexerrefuses a differentre2rustbefore comparing generated bytes. The 4.6 output differs from 4.5.1 only in its generator provenance header; upstream’s 4.6 implementation change is Zig-only. -
Merged::reportis the only route from a merge to a file.into_fileandfileare gone fromMerged;report(sink)yields aReported, which owns them. Serializing a merge without asking what it dropped is no longer writable, which is what two commands did until each was fixed by hand. -
chatter mergeandchatter pipelinenow warn when a File 1 speaker is dropped. A speaker outside--retainloses every utterance while keeping its@Participantsrow, so the output declares someone who says nothing.AmbiguousSpeakerdoes not catch it: that fires only when a code appears in both files. Both commands print the same warning, naming the speakers and how many utterances each lost, from one shared reporter.chatter batchdrives the pipeline path, so the silent one was the path that runs whole corpora. -
merge_chat_filesreturns the provenance of every merged utterance, not only the merged file. It always knew this and threw it away: it walks each input in order building two lists, then stable-sorts the combination bystart_ms. A consumer joining the output back to its inputs had to reconstruct the mapping by matching(speaker, raw bullet), which is correct only while two facts hold that no caller can check, that the sort is stable and that inserted utterances are cloned unedited.The return type is now
Merged, carrying the file, oneMergeOriginper output utterance in output order, oneReferenceFateper File 1 utterance and oneDonorFateper File 2 utterance, each in its own input’s order. Ordinals areReferenceIdx/DonorIdx, separate types because two same-signature accessors over one index type answered confidently about the wrong file.merge_chats, the string wrapper, is removed: it had no non-test caller, and aStringreturn cannot carry the provenance, so every consumer that wants the report has to work on parsed files anyway.MergeError::Parsegoes with it, sincemerge_chat_filestakes files that are already parsed.Prefer
Merged::utterances_with_origin()to pairing the accessors by hand: zippingorigins()againstfile().linestype-checks and is wrong by the number of header lines.DonorFateis a partition rather than a list of exclusions, so “this donor utterance is unaccounted for” is not expressible. An earlier form returned only the excluded ordinals and proved completeness with arithmetic, which balances just as well when every ordinal is shifted by the donor’s header count.Both inputs are accounted for.
DonorFatecovers File 2;ReferenceFatecovers File 1, where an utterance whose speaker is not retained is dropped andAmbiguousSpeakerdoes not catch it, because that fires only when a code appears in both files. A reference-onlyMOTwithretain = [CHI]therefore passed every precondition, kept its@Participantsrow, and lost every utterance silently.DonorFate::Insertedcarriestiers_stripped, becausestrip_tiersapplies to the donor and only the donor: a bareInsertedclaimed “carried over” for an utterance that was carried over AND edited.Still not a complete account of everything a merge omits, which is why the accessors are named
excluded_by_retainanddropped_not_retainedrather thanexcludedanddropped: donor headers other than@IDand@Commentare not carried.
Fixed
-
chatter pipeline --override-filenow refuses invalid override files. Malformed TOML, unsupported schema versions, and read failures previously disappeared into “no override configured”, silently triggering a fresh automatic speaker match. Explicit operator input now travels through the existing typedOverrideFileErrorexit path and no merged output is written. -
chatter rediarizeno longer counts overlapping turns from one track twice. A track appearing in several turns is now measured by the UNION of its coverage: gaps remain gaps, while same-track overlaps count their shared interval once instead of manufacturing speaker time and distorting--contested-at. Cross-track overlap remains evidence for both simultaneous speakers, soownership.total_msis the sum of per-track union-held time and can exceed the utterance bullet’s duration. The JSON shape is unchanged; its corrected measurement semantics are documented in the user guide. -
chatter fix --applycan now repair E750 inside its recovered utterance. Ordinary splice edits remain barred from parser-tainted regions. The E750 catalog entry alone carries the typed state that it removes the delimiter whitespace responsible for that recovery, and the command still reparses and independently validates the resulting CHAT before writing it.
0.16.0 - 2026-08-27
Removed
-
E214is retired, and the reason is worth more than the code was. It began as “a bare[*]carries no error code” and was deliberately DISABLED as leniency Decision 1, because reference files use bare[*]as valid CHAT. Its number was then reused in the same file for a DIFFERENT rule, “the scoped-annotation list is empty”, while its spec file went on documenting the original. So one code carried a retired rule in its documentation and an unreachable one in its implementation, and its own spec example produced no diagnostic at all. Nothing detected the drift because neither rule could fire.ErrorCode::EmptyAnnotatedContentAnnotationsis gone; code matching on it will not compile. -
rules::should_skip_groupis absorbed into the descent module that was its only remaining caller.
Changed
-
Scoped annotations are NON-EMPTY by construction.
AnnotatedContentAnnotations::newreturnsOption<Self>,TryFrom<Vec<_>>replaces an infallibleFromthat skipped the check,Deserializerejects an empty list rather than accepting one off the wire, and there is noDefault.Annotated::new(inner, annotations)takes the annotations instead of starting empty;Annotated::with_one(inner, annotation)is the single-annotation path, andwith_scoped_annotationstakes the newtype.The
OptionIS the bare-versus-annotated decision, so sevenif scoped.is_empty()branches in the two parsers collapsed into it. -
UtteranceContentgainsAction, andBracketedItemgainsGroup. Both enums had a gap where their sibling had a bare variant, and the parser filled it by wrapping the construct in anAnnotatedcarrying nothing. That was 20,184,072 values across a 106,000-file corpus, 99.3% of allannotated_actionnodes, almost all of them a bare0marking silence in daylong audio. The two content enums are symmetric now: every annotatable construct has a bare and an annotated form on both sides. Exhaustive matches over either enum will not compile until they handle the new variant, which is the intended outcome. -
ErrorCodeis GENERATED fromspec/codes/error-codes.toml, a new per-code registry that is the single owner of a code’s variant name, its rustdoc, itskindand itsstatus, plus the retired numbers.kindandstatusare removed from all 236 files underspec/errors/; a spec’scodeis a foreign key resolved at load, so a loaded spec proves its code exists. Anything readingkindorstatusout of a spec file must read the registry instead. Exhaustiveness moved from a generator check to the compiler: a wrong match arm fails the build rather than a lint. -
GoverningMarkeris nowpub(crate); the public face isGoverningMark. Its variants were public, so a caller could construct one directly and resolve a word’s language without ever saying what enclosed the word, which is the question the type exists to force.GoverningMarkis opaque, with two constructors:of(word, enclosing)andwithout_own_marker(span, enclosing). This supersedes the 0.15.0 migration note below, which tells callers to useGoverningMarker::of(word, enclosing_span). That path is no longer public; useGoverningMark::ofwith the same arguments. -
talkbank_lsp::content_spanis removed, and with it the free functioncontent_span(&UtteranceContent). Deciding WHERE an item is now belongs to the model (WordRef::span,GroupRef::span), and deciding WHETHER the editor targets it belongs totalkbank_lsp::editor_target, which dispatches onContentStructurerather than on 28UtteranceContentvariants. A new variant is therefore classified once, in the model, instead of once there and once in the LSP where the two could disagree. -
Parser entry points take the generated typed node wrappers, not a node plus a kind string. Every route from a loose node into a typed one is
FromNodeKind::from_node, and the 411 sites that ASSERTED a node’s kind rather than testing it are gone. Callers passing(KIND_CONSTANT, raw_node)pairs pass the wrapper instead, so the proof is required where the value is born. -
JSON output changes, with no compatibility shim. An
annotated_actionorannotated_groupcarrying no annotations is nowactionorgroup. Previously-emitted JSON containing the empty annotated form will not deserialize. Regenerate rather than reading cachedto-jsonoutput. CHAT text is byte-identical either way, so no file changes validity.
Known limitations
-
--parser re2cis NOT READY to judge CHAT validity, and this release says so in the tool. A clean--parser re2crun is not evidence that a file is valid: the backend ACCEPTS constructs the default backend refuses. Measured 2026-08-27: an unrecognised scoped annotation on a quotation (“hello” [qq] .), on a pause (hello (.) [qq] .), and in utterance-initial position ([x 2] hey .) are all accepted, where the default backend reports E316 or E375.The cause is information lost before validation runs, not a missing rule. Both parsers build the same
talkbank_modeltypes and share one validator, but each has its own intermediate parse tree, and re2c’s does not carry annotations for every construct:ast::Grouphas anannotationsfield andast::Quotation, six lines below it, does not. Three of the five hosts that regressed here ARE fixed in this release, because their annotations do reach the model; these three do not reach it.--parserhelp andbook/src/architecture/parser-backends.mdnow say this at the point of use. That page also carried three claims this contradicts, including a parity table row reading “Re2c silent (misses error): 0”, and a recommendation to prefer re2c for batch and CI validation. All corrected; the parity figures are marked as not re-measured.Use re2c to COMPARE two implementations, which is what a specification oracle is for. Use the default backend to decide validity. The default backend is unaffected by any of this, and is what
chatter validate,normalize,to-jsonand the LSP use unless you ask otherwise.
Fixed
-
word@@reported one defect twice, the specific diagnostic buried under the generic one. The parser names a repeated@run as E203 with the run in hand (“a word may carry only one ‘@’ suffix, found ‘@@’”);check_inline_at_ markersthen added its own E202 (“dangling ‘@’ marker”) for the same word, because the suppression that stops a double already existed for the E203-against-E203 case and was never applied to the E202 branch four lines above it.word@c@was the same. Both now report E203 alone.A bare trailing
@(hello@) still reports E202: it carries no form type, so nothing else has named it, which is the case that branch exists for.Found by the release review; fixed by writing the spec example first, where the backend-parity gate stated it exactly:
tree-sitter [E202, E203] ... spec expects [E203]. -
chatter rediarizeassigned an utterance to the track of its single LONGEST TURN, not the track holding the most of it.best_tracktook the greatestoverlap_msover the turn list with no per-track accumulator, so three short turns of one track lost to one longer turn of another even when the first held twice as much of the utterance. Its own docstring and the CLI help both said “the track with the greatest overlap”, which is what it now computes.This is the shape a diarizer actually produces: pyannote emits short turns with gaps inside a single speaker’s run. A track appearing in several turns is accumulated now, and ties break on the track code rather than on turn order, so the winner is a function of the input rather than of how the diarizer sorted its file.
best_trackis replaced byTrackOwnership, which keeps the whole distribution (winner(),shares(),total_ms(),runner_up_share()) rather than computing it and returning one name. Returning only the winner is why the defect was invisible: nothing downstream could tell a track that held 95% of an utterance from one that held 34% of a three-way split.Breaking:
rediarizeandrediarize_contenttake a further argument (below), andRediarizeOutcomegains a field.
Added
-
chatter rediarize --contested-at SHAREreports utterances whose time is meaningfully split between tracks, in the stderr summary and in--summary-jsonunder a newcontestedlist. Each entry carries the utterance index, the track it was assigned to, and the full ownership distribution: every overlapping track with its summed milliseconds, descending, plus the total. The WHOLE distribution rather than a winner and a runner-up, because that narrower shape cannot tell a 55/45 split from 55/23/22.Contested utterances are still reattributed to their winner, so they are reported separately from
flagged, which keeps its narrower meaning of “declined to reattribute”. Placement is byte-identical with and without the flag; this is a reporting change.There is deliberately no default. Omit the flag and nothing is reported. What share makes an utterance genuinely mixed has not been measured against human listening, and a default would hand every user a constant wearing this tool’s authority. A value outside
0.0to1.0, orNaN, fails the command before any file is read, rather than silently meaning “nothing is ever contested”.Known limitation, stated in the book page: summed milliseconds per track cannot distinguish a speaker change INSIDE an utterance from crosstalk across the whole of it, and those want opposite remedies.
-
TimeSpanMs::start_ms()andend_ms(). The fields are private so thatnew()is the only route in and an inverted span cannot be built, which is right, but it left the type WRITE-ONLY through the public API:DiarizationTurn::spanis a public field of a type a caller could hold and could not read, so a downstream consumer ofparse_turns_jsonhad to re-declare the same concept to get the numbers back out. Reading cannot invert anything. -
THREE COMMANDS COULD DELETE A TRANSCRIPT AND REPORT SUCCESS. The worst class in this release, found by review rather than by any gate, and every case exited 0 with a green line.
chatter normalize notes.cha -o notes.cha, the documented in-place idiom, on a file of ordinary prose left a ZERO-BYTE FILE and printed✓ Normalized. v0.15.0 refused it, so this was a regression. On a transcript missing its@Endit deleted the LAST UTTERANCE, which is exactly the shape a file truncated mid-transfer has. On a malformed%gratier it emptied the tier, producing a file it then refused to read again. Those two are unchanged from v0.15.0 and were shipping in both.chatter debug retag-languageanddebug fix-swrote back a model that had DISCARDED an unparsable region:hello [[[[ test ]]]] world .becameworld ., in place, recursing over whole directories, with no--dry-runand no backup.debug join-retracehad the same shape and was found while fixing the other two.All four rewriters were the same three steps: parse,
to_chat_string(), write. The return type wasString, which cannot carry the one fact the caller needed, so no caller had it.chatter to-jsonrefused all threenormalizeinputs, because it happened to route through a stricter path, and the two commands disagreeing about the same model is what made this findable.Two different proofs, because the commands promise different things.
normalizereshapes and must lose nothing, sotalkbank_transform::Rewriterefuses when a source line has no counterpart in the output, compared with whitespace removed so the canonicalisation it exists for still passes. The three EDITING commands change content on purpose, so that test would refuse every legitimate edit they make; they require instead that the model reproduce the source BYTE FOR BYTE before the edit, which is the only point where faithfulness is a clean question for them. A refused file is left untouched and the message names the line, or tells the operator to runchatter normalizefirst.Six CLI subprocess tests pin all of it, including the case that must NOT refuse: the six reference-corpus files
normalizelegitimately rewrites. -
chatter debug retag-languageis new, and was missing from this section entirely. It retags a language code across all three notations it can reach (@Languages, the[- code]utterance precode, andword@s:code) and REFUSES a file naming the code in a<a b> [@s:code]span, which it cannot rewrite.--todeduplicates in@Languages. It is a tool and not a find-and-replace because a language code is also ordinary transcript content: its first use retaggedsuntofinacross a corpus wheresunis also colloquial Finnish for “your” and appears 27 times as real speech. -
A nested quotation stopped being detected as soon as either quotation carried an annotation, on the default backend only, so
“a “b” c” [//] hello .validated CLEAN while“a “b” c” .reported E372, and the two backends disagreed about a validity rule on identical bytes.[/],[*]and[% note]leaked the same way, and so did an annotation on the INNER quotation.A quotation has TWO spellings in the model, with and without its own scoped annotations, and each half of the rule named only the first.
descent.rsnamed both;main_tier.rsnamed one. The annotated spelling was introduced by the same release that gave quotations scoped annotations, and the nesting rule was never taught about it.Fixed as a type rather than as two more match arms:
GroupRef::QuotationandGroupRef::AnnotatedQuotationare folded into oneQuotation(QuotationRef)variant, mirroring theRetraceRefbeside it, so “is this a quotation” is a single arm that cannot be half-written. This is a breaking change toGroupRef; a caller matchingAnnotatedQuotationwill not compile, and the two spellings remain distinguishable one level down throughQuotationRef.QuotationRef::spanpreserves the distinction thatGroupRef::spandrew between the two, which folding them could have lost silently. The outer scan that looked for a quotation to test now descends throughContentStructureas well, so a wrapper cannot hide either side of the relation again. Spec examples 4 and 5 ofE372.mdare the two directions. -
@Locationand 14 other headers were REJECTED by the public fragment parser.parse_header_fragment, and theChatParser::parse_headertrait method behind it, dispatched through 19 hand-written arms plus a catch-all, while the grammar’sheadersupertype has 34 subtypes.@Activities,@Bck,@G,@Location,@Number,@Options,@Page,@Recording Quality,@Room Layout,@Time Duration,@Time Start,@Transcriber,@Transcription,@Thumbnailand@Unsupportedall reached it and came back as errors rather than asHeader::Unknown. The same headers parsed correctly inside a whole document, because that path matches the generatedHeaderChoiceexhaustively: two dispatchers for one job, one of them a drifted subset, and the tests covered only the arms that existed. -
Five call sites silently discarded a group and every word inside it.
convert_to_group_contentreturnedResult<BracketedItem, Group>where neither outcome is a failure, and the shape invitedif let Ok(item), which five call sites duly wrote. It is a TOTAL function returningBracketedItemnow, so there is no second case to ignore and those sites preserve the content. An intermediate two-variant enum was tried first and reverted: it moved the decision without closing it. -
The container descent rule had two implementations that had already drifted. The eight walkers and
count.rs’s four traversals each carried their own container arms, about thirty per side;walk/bracketed.rsshipped four ungatedAnnotatedQuotationarms whilecount.rsgated the same variant, so one node was walked by one and skipped by the other. Onehelpers::descentmodule owns it for every traversal now. -
A quotation could not carry a scoped annotation, which CLAN CHECK accepts. The grammar takes
quotation_with_optional_annotations. -
An
@IDage with a component too large for its field parsed as ZERO.AgeValue::from_textparsed each component with.parse::<u8>()behind an all-ASCII-digits guard, which leaves exactly one way to fail (a value above 255) and answered it with0.2;300.becameValid { years: 2, months: Some(0) }: two years and no months, presented as a successful parse, in the field that is the primary variable of most CHILDES research.Bounded honestly: no validation verdict changes. A component of three or more digits also fails the two-digit depfile pattern, so such a file was always reported invalid. What was wrong is what the typed model then SAID about it, which reaches library callers and anything reading the parsed age, not
chatter validate’s answer.age_componentreturnsResult<Option<u8>, Unrepresentable>now, so an ABSENT component (1;has no months) stays distinct from an unrepresentable one, and an unrepresentable one sinks the whole age toUnsupported, which preserves the original text byte for byte. -
ErrorCode’s string constructor with a silent fallback is REMOVED. It mapped any unrecognised string toUnknownErrorthrough a catch-all, soErrorCode::new("E7O5")with a letter O compiled and would have shipped E999 with nothing to catch it. Three of the%moralignment checks built their codes that way (E705, E706 and E716; the count mismatch among them is CLAN CHECK 140), which also meant nothing reasoning over the enum could see those three checks at all. The codes they emitted were correct, so no diagnostic changes; what changes is the API. All three return typed variants, andErrorCode::parse_exactreturnsOption<Self>and is now the only route from a string. Callers of the old constructor must handle theNone. -
A word carrying two
@suffixes was reported wrongly, and one such word was DELETED on the way out. Two distinct symptoms, which an earlier draft of this entry ran together:hello@@candhello@c@dnever formed a word at all, so the utterance fell to error recovery and reported the generic E316, “content could not be parsed”, while the model’s own rule for the shape could never fire.word@k@s:spaDID form a word and DID report E203 in v0.15.0. What was wrong there was the message, and what was dangerous waschatter normalize, which exited 0 and wrote the word back split in two:word@k@stbecameword@k@s t, silently, at exit 0. That is the deletion, and it is the reason this entry exists.Both shapes now parse, are refused as E203 with a message naming the actual defect (
A word may carry only one '@' suffix, found '@c@s:spa'), and serialize back verbatim, sonormalizeREFUSES the file instead of rewriting it.hello@c,dog@jandhola@s:spaare unchanged from v0.15.0.A word may carry at most ONE
@suffix. Ruled 2026-08-27, asked because CLAN CHECK accepts multiple suffixes and chatter does not: “Multiple suffixes might make logical sense, but it is computationally messy. So, let’s disallow that.”word@k@s:spais therefore invalid even though the form marker@kand the language suffix@s:spaare each fine alone. A documented divergence from CHECK, which passes these files; main-tier words with two@runs number zero across the ~106,000 kept files.A bare trailing
@(hello@) deliberately keeps its existing E202:@is the header sigil, and admitting a single one in word position moved the diagnostic for a doubled@End.--parser=re2cREFUSES all of these too, and agrees on the code for some of them. Measured:gumma@c@s:spaandbebe@k@streport E203 on both backends, because that lexer takes the two suffixes as separate tokens and the model’s own rule counts them. Not full agreement: on the@s:-bearing case the default backend now reports E203 alone where re2c reports E203 plus E255, because the word carries an undeclared form type and its@s:spano longer registers as a language marker. Both refuse the file.dog@b@c(E209 plus E253) andhello@@c(E321) it refuses for other reasons, since its lexer cannot form those words at all. The named message above is the default backend’s. The divergence is recorded in the parser-parity baseline as a Conflicting row forE203.md. -
E207’s message asserted that a KNOWN annotation marker was unknown. It read
"x" is not a known scoped annotation type. An annotation reaches the unknown path whenever no specific rule matched it WHOLE, which happens both when the marker really is unknown ([qq],[@ xyz]) and when a known marker carries content the rule refuses. Under--parser=re2c, whose rule set is narrower,[x 0]and[:]both land there, and the message then told the reader thatxand:are not known scoped annotation types, when the marker is not the thing at fault.[: replacement]is ordinary valid CHAT on both backends, so the old message was plainly false there. The message now names the annotation as written,could not read [x 0] as a scoped annotation, which is true in every case and shows more than the marker alone.Scoped honestly: this reaches the
--parser=re2cpath only. On the default backend the diagnostic is issued by the parser rather than the model, and its message is byte-identical to 0.15.0’s. An earlier draft of this entry said “affects both backends”; measured, it does not.And
[x 3]is NOT an example of a valid construct:hello [x 3] .is refused by both backends. Only the group and utterance-initial spellings parse, and only under re2c, which is its own divergence. -
Under
--parser=re2c, an annotation on a top-level word reported E207 at byte 0, on line 1, pointing at@UTF8.Annotated::newstarts at the dummy span and the tree-sitter parser follows it with.with_span(..); this converter never did. That path now takes the annotated construct’s span widened to cover every annotation whose text can be placed, and REFUSES rather than answering with the sentinel when nothing can be, so a span is either real or absent.SCOPED HONESTLY: this is one of eleven
Annotated::newsites. An annotation on a bracketed word, a group, an event, an action or a retrace still reports at byte 0 under this backend. Those need the same AST work as the retrace spans, which is the queued change that carries the lexer’s own token ranges through the parser instead of re-deriving them. -
validate --list-checksadvertised two checks that cannot fire. E361 (“invalid timestamp value in media bullet”) and E382 (“failed to parse%mortier content”) were markedimplementedin the code registry, and nothing can produce either. Both now list asPlanned:--list-checksgoes from224 checks (184 Active, 40 Planned)to223 checks (182 Active, 41 Planned), the one retired check beingE214above. The checks themselves are unchanged; the advertisement was wrong. -
Under
--parser=re2c, every rule keyed on a span was silently unreachable. Separators and words both reached the model atSpan::DUMMY, which is{0, 0}and therefore also a real position, and validation FILTERS on that value: a dummy span makes the model answer “there is no comma here”, so E258 (consecutive commas) and every other span-keyed rule never fired on that backend. Words and separators on the main tier now report the same spans on both backends, which is what was fixed and what was measured: 61 of 64 word and separator span sets match tree-sitter exactly.The backends are NOT span-identical in general, and this entry does not claim they are. Across the repository’s own
.chafiles, 434 (file, code) pairs are reported by both backends and 338 of them still differ, 329 because re2c answers at byte 0. Untouched here: the pause-glue mirror inparser/file.rs, every terminator, and every dependent-tier diagnostic. E370 is worse than byte 0, reporting an offset that is not a file position at all; that is unchanged from 0.15.0 and is not fixed here.This affects only the opt-in oracle parser; the default backend was never wrong.
-
Under
--parser=re2c, three diagnostics were reported TWICE. Giving words and separators real spans made three model rules reachable while the hand-written mirrors that existed BECAUSE they were unreachable were still in place, so E749, E764 and E765 arrived doubled. The parity gate stores codes in aBTreeSet, so multiplicity is structurally invisible to it and STILL is: a new six-utterance test (re2c_reports_no_diagnostic_twice) covers the mirrors it knows about instead. One doubled diagnostic survives that test’s case list, E307 on a bad speaker ID, unchanged from 0.15.0. -
Under
--parser=re2c, an interposed word lost its form marker.&*SPK:took a bare word body where the grammar defines the payload as a whole standalone word, so a form marker, an@slanguage suffix or a$POS tag on an interposed word was dropped. Found on real Spanish transcripts, where the two parsers had disagreed for as long as the rule existed. -
Under
--parser=re2c, an unrecognised annotation killed the utterance.[@ xyz]reported E321 (“unparsable utterance”) rather than E207 (“unknown annotation”): every specific bracket form had a lexer rule and anything else fell to a bare[no parser rule could use. E321 is a statement about the parser where E207 is a statement about the file. -
Under
--parser=re2c,un++doreported a misplaced linker. The word body consumed+only when an atom followed, so the word ended atunand++matched the linker rule: E766, “a linker placed after utterance content”, on a construct containing no linker. It reports E233, “empty part in compound word”, as the specification says it should.
0.15.0 - 2026-08-25
Changed
-
ExtractedWord.langbecomesExtractedWord.language: ExtractedLanguage, and the word gains aspan. The old field carried only a word’s OWN@smarker, so a word inside a<...> [@s:hin]span came out of extraction indistinguishable from an unmarked word in a plain English utterance. The extractor WALKS the tree and therefore knew about the span; it discarded what it had computed, and every NLP consumer downstream was left unable to recover it. Batchalign’s morphotag read the old field, so span-governed words fell out of second-language dispatch and were tagged against the tier language.ExtractedLanguageisUtterance | Own(marker) | Span(span). It is not anOption, because “the utterance governs” is a real answer rather than a missing one, and treating no-mark as no-language is the mistake the type exists to prevent.ExtractedLanguage::resolve(span, tier, declared)gives the resolved language directly. -
GoverningMarker::resolve_at(span, ..)resolves without aWord. The resolver only ever used the word forword.span, to place diagnostics, so a consumer holding an already-extracted word had to fabricate aWord::new_uncheckedpurely to satisfy the signature. Batchalign was doing exactly that.resolve_word_language_with_markeris deleted; one span-based core serves both paths.
0.14.0 - 2026-08-25
Removed
-
resolve_word_languageis gone from the public API. It answered “what language is this word?” without saying what SCOPE it was asking under, and silently assumed a word’s own marker was the whole story. Once<...> [@s]spans existed that was false, and the function had nowhere to put the span. UseGoverningMarker::of(word, enclosing_span)followed by.resolve(..): the constructor takes the scope, so a caller with none writesNoneexplicitly.resolve_word_language_with_markeris no longer public either; it is a primitive of that operation. -
FileStem::from_stris renamedfrom_stem. It can never implementstd::str::FromStr, whose signature has no input lifetime, while this type borrows its stem; the old name invited callers to expect a trait they could use generically.
Added
-
Multi-word code-switch spans:
<word word> [@s]and<word word> [@s:code]. Every word in the scope takes the switched language, exactly as if each carried the@s/@s:codesuffix, so a switched stretch no longer has to be annotated word by word. Bare[@s]resolves the way a bareword@sdoes. As with any scoped annotation, a single content item needs no angle brackets:hallo [@s]is well-formed.A word inside the span may carry its own marker, and the word wins. This is attested usage rather than a case to reject: transcripts mark a switched stretch with the span and individual borrowed words inside it with the donor language. Resolution is innermost-first (word, then span, then utterance) and each layer records its own provenance, so
language_metadata[].sourcegainsspan_shortcutandspan_explicitalongside the existingword_*values. A span and a suffix can resolve to the same CODE, and that field is the only way to tell which mark decided it.Consumers reading a word’s
langfield alone will under-report switches: a span is an annotation on the group, and the words inside keeplang: nullunless suffixed.language_metadatacarries the resolved answer for every word regardless of which mark produced it.
Fixed
-
E220andE763gated on the wrong language inside a code-switch span. Span resolution reached metadata but not word validation, so a word’s recorded language and the language it was checked against could disagree:<ha# kelev> [@s:heb]in an English-headed file was reportedE763as English while its own metadata said Hebrew. Both paths now share one precedence decision. -
A
[@s:code]span could not be serialized to JSON at all. The enum was internally tagged, which serde cannot use for a variant carrying a string, sochatter to-jsonfailed at runtime on any transcript containing one while the bare[@s]form worked. The committed JSON Schema described a shape the serializer could never emit; it is regenerated and smaller. -
Every generated error fixture parsed with a spurious
MISSING newlinerecovery node, because the generator stripped the trailing newline the grammar requires. Invisible for as long as each fixture also emitted a real diagnostic to hide it behind. Restoring it changed the diagnostics of zero of the 335 existing examples. -
chatter debug fix-sno longer rewrites an utterance containing a code-switch span. It strips each word’s@ssuffix after writing the[- LANG]precode, so for a word inside a span the span would then govern it and its language would silently change:<how@s:fra to@s:fra> [@s:eng] .became[- fra] <how to> [@s:eng] ., whose words resolve to eng. It now refuses such utterances, which is lossless where rewriting was not. No released version could do this, because[@s:eng]did not parse before this release; it was found and fixed within the same cycle. -
E220no longer fires on a word whose language is unresolved. It treated an empty candidate set as “no language permits digits”, so an unresolvable@sproduced “illegal digits in language X” with no X.E763already skipped in that case and the two are documented as agreeing. The visible effect:[- zho] ni3hao3@s .now reportsE248alone, which names the actual defect, rather thanE248plus a consequence of it. -
Language-gated word rules run whenever the language is KNOWN, not only when the file declares one. The gate asked whether
@Languagesexists, which is a different question:<...> [@s:eng]names a word’s language with no header present, and nothing was checking those words. -
A word carrying its own
@sinside a span, and a span on a replaced word or a retrace, are now validated under the span like any other span-governed word. Only annotated GROUPS were handled, sohallo [@s]was recorded as switched in metadata while being validated against the tier language.
Changed
-
“CA” named three different things, and two names picked the wrong one. The symbol registry’s exported arrays are computed from
parse_roleand were calledca_element_symbolsandca_delimiter_symbols, naming PROVENANCE on a value holding a PARSE ROLE. They are nowword_attached_symbolsandpaired_stretch_symbols, and their union, which builds the set forbidden inside aword_segment, isALL_MARKER_SYMBOLS. Library consumers readingtalkbank_model::generated::symbol_setssee the renamed constants.The registry already carried both facts:
notation_familysays where a symbol’s notation came from, and 2 of the 25 aredisfluencyrather than Conversation Analysis. Those two are≠(blocking) and↫(segment repetition), which every fluency corpus depends on, so the old name was false for exactly its load-bearing members. Theca_elementandca_delimiterNODE names are unchanged, andparser.c,grammar.jsonandnode-types.jsonregenerate byte-identical. -
ChatOptionFlag::enables_ca_modeis replaced byhas_effect(CaOptionEffect). The old name asserted that@Options: CAturns Conversation Analysis parsing on. It does not: CA-originated markup needs no option at all, and the flag’s scope is material judged specifically weird CA. A predicate that reads “is this file CA” is what invites gating SYMBOL ADMISSIBILITY on the option, which would be wrong for every symbol, including the genuinely CA-originated ones.The two effects are named separately because they are not the same kind of thing:
TerminatorRequirementWaivedwaives a requirement, whileParentheticalIsCaOmissionchanges what a construct MEANS. Calling both leniency would be a quieter version of the same conflation. The match on the effect is exhaustive, so a third effect, or a second flag granting one, breaks compilation rather than silently inheritingCA’s answer.Validation verdicts: UNCHANGED. Renames and one predicate; every call site computes what it computed before.
-
Tree-sitter 0.26.13 across the workspace, the grammar crate, the spec workspace and the
tree-sitter-clidevDependency, plus desktop npm devDependency bumps and jsonschema 0.51.Validation verdicts: UNCHANGED over the sampled corpus, and that needed measuring rather than assuming. 0.26.13 avoids wide error nodes on unparseable input, which is a change to RECOVERY, so it can move what
validatereports on malformed CHAT without changing one byte of the regeneratedparser.c(which is in fact byte-identical here) and without any fixture in the suites noticing, because they all parse. The corpus differential is what can see it: over 2,147 files at stride 50, stratified per repo, against the v0.13.0 released build, there were no new error codes, no per-code count increases, no newly failing roundtrips and no new cross-backend disagreements. That is a statement about the sample, not about the whole corpus; at this stride a defect in a few dozen of ~106,000 files could still hide.The generated typed CST traversal is byte-identical apart from its generator provenance stamp, and
just spec-genmoved no artifact.
0.13.0 - 2026-08-21
Validation verdicts: UNCHANGED. Nothing here moves what validate reports
on a CHAT file. Every prior entry states this either way, and the published
promise is that an entry without the note did not move its verdicts, so an
entry that omits it cannot be told from one nobody filled in.
Removed
-
talkbank_transform::capitalizeis GONE. The English capitalization transform announced in 0.7.0 (capitalize_english,capitalized_pronoun_i,is_capitalizable_initial,capitalize_first) is deleted. It is the only change here that affects a library consumer.Why: chatter is the CHAT-format authority, and English orthography is a convention of one language rather than a fact about CHAT. Nothing inside chatter ever called it; its two users were downstream generators, which wanted different policies. One of them had already written its own version of
is_capitalizable_initialand documented that chatter’s answered a different question. The module also had no stopping rule: pronoun “I” and sentence capitals today, then contractions and proper nouns on request.If you used it: copy it into your own generator, where the policy belongs. It is built entirely on public API (
walk_words_mut,Word::category,Word::untranscribed), so nothing about the move needs chatter internals. Note that the version shipped here had three defects inis_capitalizable_initial, all from deciding a structural question fromcleaned_text(), which strips the very prefixes the question needs: a non-letter-initial word did not consume the utterance-initial slot, so the capital landed on the following word; an apostrophe-initial word received no capital at all; and the&-fragment guard could never fire, so a filler took the sentence capital. Ask the typed model instead.num_wordsis unaffected and stays: it serves E220, a rule chatter enforces.
Changed
-
chatter new-filebuilds its template through the typed model. The emitted skeleton is produced by parsing and serializing a typedChatFilerather than formatting text, so it is roundtrip-proven by construction; the default output is unchanged. -
docs/errors/index.mdis one table, sorted by code. It emitted one##section per spec, 236 of them with 31 exact duplicates, each over a single-row table, and reprinted every description. It is now a flat table with Code, Name, Category, Kind, Level and Status columns: 234 lines where it was 2,247. Anything that scraped the old section structure will need updating; anything that followed theE###.mdlinks is unaffected. -
The error-spec format is TOML frontmatter with a required CLAIM per example. Landed in stages within this unreleased window, superseding earlier entries’ details: metadata moved from
## Metadatabullets to+++frontmatter (an unknown or missing field is a load error); the authoredLayerfield was then DELETED (which pipeline stage catches a rule is recorded per example in the generatedspec/observations/snapshot, and every example is a fixture in the validation corpus); andExpected Error Codeswas replaced byclaim = 'violates' | 'legal' | { subsumed_by = ... }, whose negative halves (a code that must NOT fire) are enforced.levelmoved from the spec file to the example, where it is required: a code can be violated at one level in one example and another in the next (E519 has header-level and utterance-level violations), so the fault site is a fact about the example; a code’s page renders the distinct set. A non-emptyDescriptionremains required. This matters only if you author specs againstspec/errors/. -
Corpus tests require
TALKBANK_DATA. The re2c integration tests defaulted to a hard-coded directory under$HOME, which could only ever be right on one machine and silently sent everyone else to a path that does not exist. The variable is now required and its absence fails loudly.just corpus-testsneeds it set; the default test suite is unaffected, since those tests are#[ignore]d.
Internal
Not part of any published API, listed because the commits are marked breaking:
spec/errors/*.md now has ONE parser in the spec workspace rather than two (ErrorCorpusSpec and
its types are deleted), and the spec format’s vocabulary moved to a new
dependency-light talkbank-spec-vocabulary crate that both cargo workspaces
share. The generators and talkbank-parser-tests crates are publish = false.
0.12.0 - 2026-08-16
Validation verdicts: CHANGED. Four rules report where they were silent:
E241 on illegal untranscribed spellings, E756 on any empty dependent tier,
the participants check on files declaring an empty set, and the re2c backend
on empty tiers it used to paper over. If you gate a pipeline on validate,
diff your own corpus before upgrading; see
What a Version Bump Promises.
Adjudicated against real corpus data before shipping, per the standing
grammar-change gate. The full-stride differential against the shipped 0.11.0
build covers all 106,507 corpus files and reports EVERY error code unchanged
except E241, whose 661 new instances are every one an illegal short or miscased
spelling of an untranscribed marker: 624 ww, 18 Www, 10 XX, 6 Ww, 2
Xxx, 1 Xx. All adjudicated INTENDED, the rule correctly flagging invalid
data, and the 194 affected files join the cleanup queue. No new cross-backend
disagreements and no newly-failing roundtrips.
Added
-
E241 rejects the illegal untranscribed spellings. The corpus authority ruled that
wwis not legal CHAT andwwwis canonical, addingyyagainstyyyunprompted. Which spellings are wrong is now DERIVED from the canonical set rather than listed, sowwcannot be missed whilexxandyyare caught, which is what happened before. Eight instances in the differential sample, every one adjudicated as the rule correctly flagging invalid data. -
E756 covers every dependent tier, not only
%x*. A tier line whose payload is absent or whitespace-only declares nothing. The rule always said that; only its name was%x-specific, and it could not be applied to a standard tier until the model could represent an empty one. Before this, an empty%eng:was read as VALID by the re2c backend and rejected by tree-sitter through an undescribed code, so the two backends disagreed about a file neither could explain. Zero instances in the differential sample: the construct is invalid CHAT and correspondingly rare.The rule now reaches EVERY tier whose grammar body is free text, which is every dependent tier except the structured ones (
%mor,%gra,%pho,%mod,%sin,%wor), whose bodies are not free text and whose empty case fails earlier and more specifically. That boundary is a grammar fact, not a list: a tier qualifies exactly when its rule marks its bodyoptional(...).%timgained anEmptystate to make this expressible, since both of its content variants hold a non-empty string; the Phon tiers (%xmodsyl,%xphosyl,%xphoaln,%xphoint) answer from the word or group count they already reported.
Fixed
- An empty dependent tier is no longer papered over. The re2c backend met
%eng:with no content and substituted a single space, which made the tier look well formed and the whole FILE read as valid where tree-sitter reported errors. The model can now say that a tier declares nothing, so the parser reports what the file contains and E756 judges it. - An empty
%xtier survives a roundtrip, andnormalizeno longer swallows the file.%xtst:with no content reported E756 from the PARSE path and returned without adding the tier to the model, so the line vanished on roundtrip while an empty%eng:was preserved. Worse, because the report came from parsing rather than validation,chatter normalizetreated the whole file as unparseable and wrote NOTHING. The parser now says what the file contains and the validator judges it, as it does for every other tier kind. - An empty
%tim:is a%timtier. The re2c backend lowered it to an unsupported DEPENDENT TIER and reported E605, “unsupported dependent tier ‘%tim’”, about a tier name that is perfectly supported; a whitespace-only body additionally drew E603 (“Invalid %tim tier format: ‘’”) alongside E756, two codes for one fact and the more specific of them false. Same for an empty%xphoaln:and%xphoint:, which conflated an absent body with a malformed one. All four now report E756 on both backends. - The participants check reads the declaration. An empty participant set used to disable the check rather than fail it, so the files least likely to be well formed were the ones exempted from the rule.
- An annotation’s separator is not part of its text.
[=! contacts], written with two spaces, parsed as" contacts"in one backend and"contacts"in the other. That was the last content-level disagreement between the two parser backends across all 107,403 corpus files. chatter validate --format jsonno longer writes cache housekeeping to stderr. Two facts leaked there:Cleared N cache entrieson every--forcerun, andnote: pruned N unreachable cache row(s)...whenever a prune fired. Both broke the documented promise that JSON mode’s stderr is empty, and the test suite contained two tests with contradictory expectations about it, one requiring stderr empty and one asserting it contained the cleared count. The first only failed when a prune happened to fire, which is why both shipped green through four releases.
Changed
-
Breaking (library):
TimTiergained a third variant.TimTier::Empty { span }represents a%tim:line that declares nothing, which neitherParsednorUnsupportedcould hold: both carry aNonEmptyString. Code matching onTimTierexhaustively must add an arm.TimTier::empty()constructs one,declared_content()returnsNonefor it (as_str()still flattens to""forDisplayand serialization), and the serde form is unchanged apart from""now deserializing toEmptyinstead of erroring. -
Breaking (library): the
test-utilsfeature is REMOVED, and with itChatCleanedText::test_uncheckedandChatRawText::test_unchecked. Not renamed: gone. This is the breaking change that bites FIRST, because cargo refuses to resolve a graph that asks for a feature which no longer exists, so it fails before anything compiles and is invisible to a “what will fail to build” scan. A consumer sees:package `X` depends on `talkbank-model` with feature `test-utils` but `talkbank-model` does not have that featureBuild fixtures through the parser instead:
TreeSitterParser::parse_wordfollowed byChatCleanedText::from_word. The hatch was removed because a type whose existence proves “this text came from a parsed AST” is only as strong as its weakest constructor, and one any dev-dependency could switch on was that constructor. Downstream adoption on the day of release found three fixtures that had been asserting on a shape production cannot emit (a terminator in awordslist), passing only because the hatch let them fabricate it. -
Breaking (library):
BulletContent::empty(). A named constructor for a payload that carries nothing, distinct fromfrom_text(""), which fabricates an empty text segment that is not in the file. Additive; no existing call site changes. -
NDJSON surface: a new record type. Those facts now arrive on stdout as
{"type":"cache","action":"clear"|"prune"|"warning",...}, emitted only when cache maintenance did something. Silencing them under--format jsonwas considered and rejected: they are results a caller can act on. A consumer that ignores unknowntypevalues needs no change; one that errors on an unrecognisedtypewill see these. Thetypefield’s documented value set is now"file","summary","cache", and the contract page says to treat unknown values as ignorable. See Diagnostic contract.
0.11.0 - 2026-08-13
Validation verdicts: CHANGED, in BOTH directions. Six error codes that had
silently degraded to E316 “unparsable content” now report themselves again
(E202, E307, E311, E314, E370, E375), and two false positives are gone. If you
gate a pipeline on validate, diff your own corpus before upgrading; see
What a Version Bump Promises.
Adjudicated against real corpus data before shipping: the operator’s corpus differential over a 2158-file stratified sample reports byte-identical per-code counts and no newly-failing roundtrips against v0.10.0. The changes below are all on malformed input, which a curated corpus contains almost none of.
Library APIs: BREAKING. This release changes the public API in several ways. The list below is from a mechanical diff of the public surface between the two tags, made after the notes first shipped saying “additive” and then being corrected twice as a downstream consumer hit one break after another. A release note written from memory of a 415-file change is a guess; this one is a measurement.
Removed items (6):
FormType::A. The@amarker was retired by the corpus authority in 2024 and is absent from the form-marker registry that now generates every site of that closed set; the variant survived only because sixteen hand-written copies of the list disagreed. No replacement: the construct is not CHAT.ALL_MARKERS,all_markers_string. Superseded by the same registry.collect_bracketed_content,collect_bracketed_item. Superseded by the typed traversal.counts_for_tier_in_context. Usecounts_for_tier, now re-exported attalkbank_model::alignment.iso.
Added, and breaking for an exhaustive match:
FormType::Undeclared(String)carries the raw text of a marker naming no declared form, soword@zzroundtrips instead of being silently rewritten toword@z:zz.
Changed signatures:
ChatFile::validateandvalidate_intotakeTranscriptName<'_>rather thanOption<&str>.NonebecomesTranscriptName::Anonymous; a real name becomesTranscriptName::Named. TheOptioncould not say which of “no name” and “a name we failed to read” it meant, and both reached the same branch.
Library APIs: additive. talkbank_model::alignment re-exports
walk_words, walk_words_mut and counts_for_tier, which previously required
naming the helpers module.
Fixed
-
A recovery node could displace an entire
tier_body. An utterance ending in “ .“ was told its terminator was missing (E305), and a retrace or bracket at utterance start took the rest of the line with it. The parser was reading its own recovery artefact as evidence about the user’s file. The generated typed CST traversal is regenerated from a generator that no longer absorbs an ERROR child at whatever position its cursor had reached. -
Six codes degraded to the E316 catch-all. Which classifier a recovery node reached was decided by WHERE tree-sitter had put it, so the same construct was named precisely at utterance start and generically after spoken material.
MainTierRegionis now stated by the caller that knows it, and every main-tier Unexpected sink routes through one owner. -
E246 blamed a lengthening marker for a stray tab. The classifier saw a
:before the recovery node, and that:was the SPEAKER’s. A tab inside the main tier now reports the tab. -
E758 pointed at whitespace nowhere near a tab. It claims “extra whitespace between the tab and the tier content”; filling that slot never established the adjacency the sentence asserts, so ordinary space between two words was reported as a leading-space violation. The span is now built only when it starts at the tab’s end byte.
-
E754 retired.
@land@lsno longer require a single character. -
Windows: the content-catch-all gate reported every exempted file as new, because repo-relative paths were compared with the host separator against a forward-slashed list.
Changed
-
Two wall-clock test assertions became hang detectors with order-of-magnitude ceilings; both were tuned to one machine and one of them turned the Windows matrix red doing correct work.
-
chatter-desktopandtalkbank-llmsetdoctest = false. Neither has doc examples, and each was paying a full rustdoc compile to run zero doctests.
0.10.0 - 2026-08-07
Validation verdicts: CHANGED, in the stricter direction. Two new error codes reject retrace constructions that previously passed silently, and two existing rules stopped being suppressed by the shape of the content they were asked about. Adjudicated over all 107,376 corpus files: E377 fires 53 times in 42 files and E378 15 times in 12 files, all real transcription defects and queued for data cleanup; restoring E372 and E704 costs zero new instances, so those two are pure correctness.
Library APIs: BREAKING. Retrace loses its annotations field, both
content enums gain an AnnotatedRetrace variant, and per-word language records
lose word_index. Pre-1.0, so this is a minor bump.
Every fix below except the merge one is the same defect: a traversal carrying
its own private list of which content variants contain other content, plus a
catch-all arm for everything the list forgot. Five such traversals existed.
Added
- E377
RetraceWithNoMaterial. A retracing marker whose material is another marker, so it retraces nothing of its own. One rule covers the unbracketedна [//] [/] наand the bracketed<<a> [/]> [//], because the lowering folds both into the same tree; naming it for the shape rather than for a spelling is what makes that possible. Deliberately narrow: 11,163 retraces in the corpora sit inside another retrace and only 4 wrap a lone marker, so a “no retrace inside a retrace” rule would have rejected ordinary stutter chains (<the [/] the piece> [//] the people) in exactly the aphasia and fluency corpora that study them. - E378
RetraceWithoutWords. The retraced material must contain a word at some depth. Phrased against absent WORDS rather than a present event, so<the floor on the &=laughs water> [//]stays legal while<&=sigh> [/]does not. The boundary falls out of what the model already calls a word:0det [/] 0det dogis legal (an omitted determiner is lexical content),<xxx> [/] xxxis legal (untranscribed speech is speech), and0 [=! snuffles] [/] okis not.
Fixed
- A retracing marker’s position among its annotations was discarded.
dog [* p:w] [/] doganddog [/] [* p:w] dogare different claims: the first codes the error on the abandoned attempt, the second on the retrace. chatter built the identical model for both and wrote the first back as the second. 12,226 places in the corpora put an annotation immediately before a retrace marker. A second adjacent marker had nowhere to go at all, soна [//] [/] наround-tripped asна [/] на, losing a marker outright in 105 places across 46 files, 31 of them bilingual or language-impairment corpora where disfluency is the research variable. - E704 (overlapping bullets) was silently disabled on any utterance whose content held a retrace or a group. The predicate for “does this utterance say anything timeable” recursed into neither, so two speakers’ bullets could overlap by a full second and report nothing whenever either line contained a retrace.
- E372 (nested quotation) was invisible below every container except an
annotated group.
“a <“b”> [/] c”is a quotation inside a retrace inside a quotation, and reported nothing. - Per-word language metadata skipped every word inside a quotation,
phonological group, sign group or retrace, in a tool whose per-word
language resolution is the point.
hao3 “ni3” <ma> [/] maproduced records for two words out of four. - The re2c backend diverged from tree-sitter on marker runs. It split a run where tree-sitter folds it, dropped retrace markers on events, and never raised E377 at all, despite three doc comments saying it did. The parser-equivalence gate now covers all three.
chatter mergedropped donor@Commentrows when the reference file had none of its own.- Five missing CHANGELOG link references. Every release from v0.6.0 to
v0.9.1 shipped a
## [X.Y.Z]heading with no matching[X.Y.Z]:definition. It renders as literal bracketed text rather than as a broken link, so the book’s link check reports zero errors and cannot see it. The version gate now requires both halves of the entry.
Changed
Retrace::annotationsis gone. Annotated retraces areAnnotatedRetrace(Box<Annotated<Retrace>>)in both content enums, parallel to the existingGroup/AnnotatedGrouppair. Parser lowering is now a left fold over the marker run, one wrapper per marker, which absorbed three hand-rolled copies of the same tail.- Adjacency is validated, not refused at parse time. Folding the offending input faithfully means it still round-trips, so a file that trips E377 stays recoverable rather than being partly discarded during recovery.
- Per-word language records no longer carry
word_index, andget_word_languageis removed. The records are a#[serde(transparent)]list, so a consumer reads position withenumerate(). The stored index was a second representation of that position whose documentation claimed it matched the tier-alignment domains; it cannot, because%morexcludes retraces and%phocounts them, so no single integer indexes both. --parser re2cis documented as reporting unreliable diagnostic positions. The lexer emits spans and the parser discards them; until that is plumbed through, the flag is for cross-checking verdicts, not for locating them.
Internal
- Design rule 3 (no
_ =>catch-all over the content enums) is now enforced by the compiler, through#![deny(clippy::wildcard_enum_match_arm)]added per file as each is cleaned, seven so far. A reintroduced catch-all is a compile error at the exact line, which no scalar count could be.cargo run -p talkbank-parser-tests --bin audit_content_catch_allsinventories the 24 modules still to clean. ContentStructureis the single owner of which content contains what, carryingWordRefandGroupRefpayloads so a caller can ask not merely whether something is a container but which one. That set had been encoded independently in 18 files, and two copies disagreeing about phonological and sign groups is what let E377 escape from inside‹...›.
0.9.1 - 2026-08-05
Validation verdicts: UNCHANGED. No rule was added, removed or altered, and no file changes its valid/invalid verdict. This release completes the library API that v0.9.0 closed the fields on, and every entry below was found by compiling a real downstream consumer against v0.9.0, which is a gate this project did not previously have.
Added
into_vec(),take()andretain()on every collection newtype. v0.9.0 made these types’ inner fields private but shipped only the READING half of the resulting API. With no consuming accessor there was no way to move the items out, so a consumer rebuilding a content list or resegmenting a file could only clone throughas_slice().to_vec(), on paths that run per utterance; and with noretain, every caller wrote take-edit-rebuild by hand, which hands a closure a&mut Vec<_>and isDerefMutunder another name. Downstream, these three methods delete three helper functions and sixteen copies of one incantation.- One owner for that API.
collection_newtype_ops!now emits the accessor set for all seventeenVec-backed newtypes. They had drifted while hand-written:into_vecwas on 6 of 17,as_sliceon 11,as_mut_sliceon 6, so what a consumer could do depended on which type it happened to hold. TierContentItemsandBracketedItemsare re-exported frommodel. Both werepubbut reachable only through a glob, so a consumer could not name the type to reconstruct one after its field closed.
Fixed
- A doc-comment claim that was not true. Several comments said
reconstruction “goes through
new, where a future invariant would be enforced”. Every one of these seventeen types also hasimpl From<Vec<T>>andimpl Deref<Target = Vec<T>>, sonewis not the only door and no invariant is enforceable on them today. Closing the fields prevents literal construction and destructuring, and nothing more. The docs now say that, and name the open question (whetherFromshould becomeTryFrom) rather than implying it is already answered. - A test whose “unique” temporary directory repeated 98% of the time. The
name was the pid plus
SystemTime::now(), but the pid is constant across a test binary and 19,584 of 20,000 consecutiveSystemTimesamples measured identical, so parallel tests shared a directory and one’s cleanup deleted the other’s file. Now a process-wide counter.
0.9.0 - 2026-08-05
Validation verdicts: CHANGED, in the permissive direction. Files that
earlier versions wrongly REJECTED now parse: an unquoted @Media filename may
contain dots, parentheses, interior spaces and non-ASCII characters. Nothing
that used to pass now fails. A comparison over a 2,136-file
stratified sample of the reference corpora reports no new error code and no
count increase on any code, and no newly-failing roundtrip file.
Three rules changed with no effect on any known transcript. E767 (new) reports
whitespace before the @Media comma; those files were already invalid, and what
changes is the diagnostic. E768 (new) cannot be reached from a .cha file at
all. E602 became E756 on empty user-defined tiers, a construct that occurs zero
times in the wild corpus.
This release closes the library’s newtype surface ahead of 1.0, so it carries a lot of breaking API change and very little behaviour change.
Fixed
@Mediarejected legal media filenames.media_filenamewas an ASCII allowlist ([a-zA-Z0-9_-]+), so a dot, a space, a parenthesis or any non-ASCII character made the header fail to match. The failure surfaced as E330 “Missing media_type node” on a line that visibly ended in, audio, and as E525 about a header chatter had recognised perfectly well. A filename is now defined the way the format defines it, as everything up to the comma that introduces the media type, with the quoted form still available for URLs (which may contain commas). This was costing real transcription runs: a media file named in Chinese, or containing a space, could not be referenced at all.- A
%mortier could be silently dropped on an empty user-defined tier.UserDefinedDependentTier::contentwas aNonEmptyString, so the model could not represent a%xtier with empty content and the two parsers disagreed about what to do with one. The state is now representable and rejected by a validation rule (E756) rather than being unrepresentable and handled twice. - The validation cache could panic on drop inside an async runtime. This was the second half of the nesting bug fixed in 0.8.0: the first half covered the call, this one covers teardown.
- A
%morclone was a no-op, cloning a reference rather than the owning vector it was meant to copy. - A declared speaker with no
@IDwas reported as undeclared. The “Speaker *X not declared in @Participants” check read the@Participants-to-@IDjoin rather than the@Participantsheader, so for a speaker declared without an@IDit asserted the opposite of the file. The missing@IDis a real fault and E522 already reported it correctly. The neighbouring “@Participants header missing or has no participants” check had the same confusion: an empty join and an absent header are different facts. - E767 never fired in the editor. It was implemented as a file-level sweep,
and the LSP calls
validate_headers_only, which does not run those. Both@Mediapayload rules now live on the per-header dispatcher that every entry point calls, so the CLI and the editor report the same thing. The LSP’s per-speaker code lens had a quieter version of the roster bug: a speaker without an@IDgot no lens while speaking. - A spec file had been failing to load silently.
E502_wor_cascade_regression.mdcarried a malformed title, and the loader downgraded every load failure to a warning on stderr, so it simply left the corpus unnoticed. The loader now fails closed on a spec it cannot parse, and distinguishes a spec from the prose that shares its directory.
Added
ChatFile::declared_speakers()returns every speaker declared in@Participants, in declaration order, each enriched with its@IDmetadata when present.participantsis populated from the@Participants-to-@IDjoin, so a speaker declared without an@IDraised E522 and was then absent from the map: consumers saw fewer speakers than the file declares. Prefer this for “who is in this transcript”;all_participants()remains the@IDjoin.ChatFile::participant_entries(), the named@Participantsextraction, alongside the existingid_headers().MorWord::analysis()borrows the analysis half of a%moritem (lemma[-Feature]*) so a consumer whose token model keeps the tag and the analysis in separate fields need not serialize the whole item and strip thePOS|prefix back off.MorWord::write_chatnow delegates to it, so the two renderings cannot drift.DependentTierEntry::kind(),span()andcontent_span(), the last giving the byte range of a tier’s content without its label or terminator.MediaFilename::parse(),unquoted(), andMediaFilenameProblem.- E767: whitespace between the
@Mediafilename and its comma. Reported from the validation layer so both parser front ends raise it from one implementation. - E768: an
@Mediafilename that cannot be written to a header and read back unchanged. Unreachable from CHAT by construction; it guards the JSON ingress, where a document can carry a value no transcript could express. string_newtype_read_impls!, the read and render surface shared by every string newtype, so a newtype WITH an invariant can share it instead of copying it.Status: unreachable_from_chatfor error specs: a rule that IS implemented but that no CHAT input can trigger, so it carries no corpus fixture and owes a named out-of-corpus test instead. This closes a hole in the gate meant to stop an implemented rule shipping untested: a spec with no example used to fail to parse, and the loader turned that into a warning, so the gate never saw the one case it names. Both directions are now checked, a spec marked unreachable that carries an example is also an error.
Changed
- BREAKING: no model newtype exposes its inner field. Every newtype in the
model, including every one generated by
string_newtype!, now has a private field. Code reading.0usesas_str()/as_slice()/raw(); code mutating through it uses the named accessors. - BREAKING:
DerefMutis gone from the collection newtypes. While it existed, a private field bought nothing: any caller could still push, clear or replace the contents.as_mut_slice()allows element mutation without allowing the collection to be resized. - BREAKING:
MediaHeader::newtakes aMediaFilename, notimpl Into<MediaFilename>, andMediaFilenamehas nonew, noFrom<&str>and noFrom<String>.parseis the only way in. An@Mediafilename containing the delimiter was constructible, andbuild_chatbuilt one. - BREAKING:
build_header_linesandbuild_media_headerare fallible, andBuildChatErrorgains aMediaFilenamevariant, because a caller-supplied media name is external input that@Mediacannot always represent. - BREAKING:
UserDefinedDependentTier::contentis no longer aNonEmptyString. - BREAKING: the crates are edition 2024.
- Deserialization of
MediaFilenameis lenient, like every other checked newtype in the model: the serde boundary reconstructs what the document held and validation reports the violation with a code and a span.
0.8.0 - 2026-08-03
Validation verdicts: UNCHANGED. No rule was added, removed or altered, and
no file changes its valid/invalid verdict in this release. What changes is that
the desktop app can run at all, that runs differing only in --suppress share a
cache again, and the library API named under “Changed” below.
Fixed
- Chatter Desktop can validate again. Since v0.6.0 the desktop app could not
start a run at all: it stopped on “Starting…” forever, on every machine and
every folder. Tauri drives a command on its async runtime, and the validation
cache bridges its synchronous API to an async database by owning a runtime and
blocking on it; nesting runtimes panics, the panic unwound out of the command,
and the IPC call then never resolved OR rejected, so the window had nothing to
report and no error to show. The cache now runs such a call on a thread with no
ambient runtime, so nesting cannot arise, and the desktop
validatecommand always produces an outcome, reporting a panic as a failed run rather than as silence. The CLI was never affected. Introduced 2026-07-07; shipped in v0.6.0 and v0.7.0.
Changed
-
Desktop commands return typed errors instead of
String. Each command now names the failures it actually has (TargetError,ValidationStartError,ClanError,InstallCliError,RevealError,ExportError,OpenExternalError), so a failure can be matched on and carries its source error. Errors still cross the IPC boundary as the same display text, so nothing the user sees changes. -
--suppressno longer throws the validation cache away. Suppression is a presentation preference: it changes which diagnostics are printed, never which ones the validator computes. v0.6.0 folded the suppression set into the cache key, so every distinct--suppresslist got its own private cache andchatter validate ~/corpusfollowed bychatter validate --suppress xphon ~/corpusre-validated all ~106,000 files from cold instead of hitting the cache. Runs that differ only in--suppressnow share one cache;--strict-linkers, which genuinely turns extra checks on, still validates afresh. Suppression behaviour itself is unchanged: a suppressed code is not reported, and a file with other diagnostics still counts invalid. -
The cache no longer grows without bound across releases. Every read binds the current rules version, so rows written under a superseded one can never be matched again, yet nothing deleted them: only a 30-day age cutoff existed, which answers a different question. Each release therefore stranded a complete copy of the corpus in the database, which had reached 464,773 rows across 88 versions (about 190 MB of a 243 MB file) for a corpus of ~106,000 files. Opening the cache now deletes rows outside a two-generation window (the current version plus the most recently written previous one, so a rollback or a bisect is not cold), rewrites the file so the space actually returns to the filesystem, and reports what it reclaimed.
-
Rule selection and presentation policy are now separate types.
ValidationConfigheld both “which rules run” and “how diagnostics are shown”, and the validation cache key was derived from the whole thing, which is what let a display preference partition the cache. It is replaced bytalkbank_model::RuleSelection(what is computed; the only input to the cache key) andtalkbank_transform::PresentationPolicy(what is shown). The cache crate cannot name the second, since the crate that owns it depends on the cache, so folding a display preference into the key is now a compile error rather than a judgement call.Library callers:
ChatFile::validate_with_configandvalidate_with_alignment_and_configare nowvalidate_with_rulesandvalidate_with_alignment_and_rules, taking aRuleSelection, and they report the complete diagnostic set with nothing filtered.ConfigurableErrorSinkmoved totalkbank_transformand takes aPresentationPolicy. The validation runner’s config fieldmodel_configis now the pairrulesandpresentation.
0.7.0 - 2026-08-03
Validation verdicts: UNCHANGED. No rule was added, removed or altered, and no file changes its valid/invalid verdict in this release. What changes is what a run REPORTS about itself when it does not complete normally.
Changed
-
A validation run now always terminates its event stream, and says how. Previously a run whose thread died emitted no terminal event at all, and the three surfaces each guessed differently: the CLI exited non-zero, the desktop app waited forever showing “Discovering files”, and the TUI marked the run COMPLETE, so a dead run was presented as a finished one. Terminality now belongs to the runner, which guarantees a terminal event on every exit path, so all three surfaces inherit the same guarantee instead of reconstructing it.
-
A run that lost files can no longer report success. A panicking worker was caught, logged where no graphical user could see it, and then ignored: the run reported
Finishedwith partial statistics, so a 500 file corpus could validate 480 and be presented as “all valid”.Finishednow means every discovered file was accounted for and is the only basis for a claim about the whole input; a run that covered less reportsFinishedIncompletewith the number of files lost.chatter validateexits non-zero in that case, where it previously exited 0. -
Cancelling a run now cancels it. The cancel request was a single token on a channel that the dispatch loop, every worker, and the end of run check each consumed destructively, so exactly one of them observed it: cancelling stopped one worker while the rest drained the queue, and the run’s own statistics usually recorded that it had not been cancelled.
-
ValidationEventgainsAbortedandFinishedIncomplete(BREAKING for library consumers). The enum is deliberately NOT#[non_exhaustive]: a consumer that upgrades gets a compile error and has to decide what a dead or partial run means for its own interface, rather than silently inheriting “pretend it finished”, which is the defect these variants exist to fix.
Fixed
-
Desktop: a run that never started looked identical to one in progress. The app showed “Discovering files” from the moment it sent the request, and the backend’s own discovery event set the same state, so a backend that never answered was indistinguishable from one still working. The two are now separate states: the app shows “Starting” until the validator actually responds, and says so if that takes more than a few seconds. A run stuck there is a start-up fault rather than anything about your files, which is worth quoting in a bug report.
-
Desktop: an aborted run is no longer a dead end. It reports why it stopped and offers Re-validate, instead of leaving the window with no way forward.
0.6.0 - 2026-07-31
Validation verdicts: CHANGED. Files that earlier versions accepted may
now be rejected, and files they rejected may now be accepted. Both
directions occur in this release: the Phon %x fixes below remove false
rejects, while the removed error codes and the suppression fix change what
validate reports and what exit code it returns. Pin with ~ if you depend
on a fixed rule set.
Added
chatter fix, built on the span-splicing engine. It supersedes the deletedchatter lint: the oldlint --fixis nowfix --apply.fixcovers the full fix catalog rather than three codes, applies fixes at exact byte spans validated against the source text, and repairs a clean utterance in a file whose other regions did not parse (the utterance containing an edit must have parsed clean, or the edit is refused and reported, never silently dropped). Every catalog entry carries a batch-safety tier and a bare--applywrites only the mechanical ones; a semantic fix is written only when its code is named with--code; an ambiguous fix is only ever reported, never written by this command.
Removed
-
chatter lint(the--fixauto-fixer) is deleted. It was a live span-driven byte writer built before the splice engine’s safety guarantees existed: it readerror.location.spanwith no dummy-span guard (Span::DUMMYis{0,0}, a real file offset), calledString::replace_rangewith nois_char_boundarycheck (a panic on a non-character-boundary span), inserted its E301 terminator fix at zero width with no dummy-span guard either (corrupting the@UTF8header had one ever fired at offset 0), and detected no overlap between fixes. An audit found zero production callers (notalkbank-toolsreference, no workspace script, no downstream pipeline usage; only its own tests and the book mentioned it).chatter fixis its successor; see Added above. -
Five
ErrorCodevariants, each unreachable or redundant:LongFeatureLabelMismatch,NonvocalLabelMismatch,UnexpectedTierNode,UnexpectedMorphologyNode, andLegacyWarning(with its generated spec entry). Consumers matching onErrorCodewill see these gone. -
Twelve of the language server’s twenty-one quick fixes. Each was attached to a code it did not repair: the action offered for a duplicate header inserted a missing one, and eleven others were similarly mismatched. The nine that remain repair the diagnostic they are attached to. Quick-fix matching now goes through parsed error codes rather than string literals, so a renamed code is a compile error instead of a silently dead action.
Fixed
-
Phon
%xdependent tiers, reconciled against the upstream spec. Three fixes, two of which were false rejects on valid Phon output:%xphoalnno longer requires its word count to equal%mod/%phoexactly. The spec allows a pause present on only one of the two tiers to consume a word slot only on the tier that contains it, so the counts legitimately differ by one.- Numeric inter-word pauses (
(1.5),(1:05.2)) are accepted on the syllabification tiers, alongside the three untimed forms. They were rejected on the grounds of being unattested in available corpora, which is not a basis for refusing a construct the spec declares legal. - Intra-word pauses (
^, U+005E) are tokenized rather than absorbed into the neighbouring phone. A word-final^previously produced a spurious error, and a mid-word^silently became part of the following phone. Reconstruction preserves the pause in place, per the spec’s rule that stripping each unit’s:CODEand concatenating must reproduce the source word exactly.
-
validate --suppressno longer zeroes the invalid count and the exit code. Suppressing a code removed it from the report AND from the tallies, so a file with genuine OTHER errors could be counted valid and the command could exit 0. A file that still has unsuppressed diagnostics now counts invalid and the command exits non-zero, as it should. A file whose every diagnostic was suppressed does count valid: that is what asking for those codes to be suppressed means. -
The validation cache key covers every dimension of the verdict, including parse behaviour and the active rule set, as a required parameter rather than a hand-picked subset. A cached verdict from one configuration could previously be served for another.
-
Strict parsing no longer discards the model it built on failure, so a caller can inspect what parsed alongside the diagnostics.
Changed
-
Diagnostic classification happens once, from the active rule set, rather than being recomputed at three call sites that could disagree.
-
The diagnostic kind is generated from the spec instead of mirrored in a hand-maintained match, and a divergence between the spec and the
ErrorCodeenum now fails the build in both directions rather than falling through to a default.
0.5.1 - 2026-07-30
Fixed
validate --forcewas unusable at corpus scale (v0.5.0 DOA): the cache refresh calledclear_prefixonce per resolved FILE, and each call scanned everyfile_pathin the cache, so a corpus-sized invocation did quadratic work (on a 136k-file cache, effectively forever) at 100% CPU behind a blank screen before the progress display started. The refresh is now one batchedDELETE ... IN (...)pass over the resolved file list, andclear_prefixitself became a single range-predicate statement instead of a scan-and-loop. Pinned by a real-CLI regression test that warms a 6,000-file cache and bounds the forced pass (old: 34s at that size; new: seconds, dominated by validation itself).
0.5.0 - 2026-07-30
Removed
- Two ungrounded CA-mode validation exemptions.
@Options: CAno longer disables the E241 illegal-untranscribed checks, nor the E701/E704 temporal checks (which it had skipped wholesale via an early return, while E362 bullet monotonicity kept running on the same files). Neither skip had a CLAN CHECK counterpart: CHECK’s whole CA behavior is three suppressed errors (21 terminator, 155 parenthesized word, 123 leading space), and chatter keeps exactly those three. Measured before removal on ALL 994 kept CA-declared files: both gates protected zero occurrences. The temporal skip’s recorded rationale (leniency-policy Decision 6, false positives on CA reference files) no longer reproduces, since the temporal rules gained the 500 ms tolerance and per-speaker semantics; its Revisit line anticipated this removal. CA files with genuine timing defects or illegal untranscribed markers are now diagnosed like any other file.
Fixed
-
E326 now says when the skipped line looks like a CHAT line pushed off column 1. An indented dependent tier (
%mor: ...) was reported as “Unsupported line skipped”, accurate but useless: the reader hunts for junk when the fix is deleting one space. The message now names the shape (“looks like a dependent tier line pushed off column 1; it must begin at column 1”) for tier-, main-tier-, and header-shaped lines, with a suggestion to remove the leading whitespace. -
An annotated word’s wrapper span was never set, left
Span::DUMMYat construction while the annotated event, action, and group paths all set a real one. Two consequences: any diagnostic located on an annotated word pointed at byte zero, and E757 could not see a bracketed code glued to the following word (hello [!]there) at all, because its detection is span adjacency. The wrapper now spans the word through the enclosing node’s end, covering its trailing[...]codes, exactly as the retrace paths do.
Changed
- E757 now covers every bracketed code, not only retraces.
hello [/]xwas rejected;hello [!]xandbobo [= toy]xwere silently accepted, though they are the same defect and the code’s own description already said “bracketed code”. Juxtaposition-matrix cell 8, ruled REJECT 2026-07-18. Mirrored in the re2c front end, where a bare closing bracket joins the retrace tokens. The 2026-07-18 matrix scan found][letterunattested corpus-wide and the differential confirms no new instances, so no kept file is affected.
Added
-
talkbank-transformgained a default-onvalidation-runnerfeature. The corpus-scale validation runner is the crate’s only SQL consumer (sqlx, viatalkbank-cache), so it, that dependency, and the runner-onlycrossbeam-channel/num_cpusnow sit behind the feature. Default builds are unchanged; a consumer that wants the transform surface without a SQL stack opts out withdefault-features = false. The path predicateis_chat_transcript_pathmoved to the feature-independenttalkbank_transform::paths(still re-exported fromvalidation_runner), since the corpus walk and CLI walks need it on every build. -
E766, a linker placed after utterance content (
the dog ran +" away .). Linkers connect an utterance to the previous one, so they are utterance-initial by definition; a misplaced one used to surface as generic unparsable content (E316), which gave the transcriber nothing to act on. The grammar now parses the misplaced linker into the CST (the same strict+catch-all pattern as the curly-quote rule) so the diagnostic names the construct at the exact token, in both parser front ends. One deliberate carve-out: a++glued to words on both sides (un++do) is a word run with an empty compound part and keeps its E233 diagnosis. A side effect of the grammar change is finer error recovery on several unparsable-content shapes: diagnostics that used to blame a whole line now land on the exact offending region (e.g. an unmatched<now yields E316 on<wordwith the rest of the utterance parsed normally). -
E765, a free-standing
:or;separator, or a pause, glued to the item after it (:and,;;,(.)dog). Same family and same span-adjacency mechanism as E764; the preceding side stays valid, sinceword↘anddog,are documented convention anddog:fuses into the word.Juxtaposition-matrix cell 7 was ruled REJECT for the whole separator class, against an estimate of roughly six affected files. A real-corpus comparison measured that reading at 270 new instances on a 2%, 2,134-file sample (about 13,500 corpus-wide), every inspected one legitimate CA notation rather than a missing space:
≡is latching and is written glued on both sides, and the intonation arrows attach to the material they mark, including directly before an overlap close. Adjudicated UNINTENDED, so the rule ships narrowed to the plain punctuation separators and pauses, where the differential is clean. Whether any CA mark should forbid trailing glue is left open, with receipts in the spec. -
E764, a
&-prefixed form glued to the preceding word (dog&-um,dog&~gaga,dog&+fr). The shape parses as two words, because&cannot continue a word, so a missing space silently manufactures a word boundary that the transcriber did not write and nothing reported. Style rule in the E749/E751/E757 family, detected by span adjacency, mirrored in the re2c front end as a token scan. Glued omission (dog0is) is not this code: it yields one malformed word and E220 already rejects it.Juxtaposition-matrix cell 6, ruled REJECT 2026-07-18; zero main-tier attestations in the kept corpus at adoption, so no existing file is affected. Validator-only: grammar, model shape, and roundtrip behavior are untouched.
0.4.1 - 2026-07-27
Fixed
-
talkbank_transform::dependent_tiers::replace_or_add_tiercould not be called on an utterance. It still tookSmallVec<[DependentTier; 3]>afterDependentTierEntrywas introduced andUtterance::dependent_tiersbecameSmallVec<[DependentTierEntry; 3]>, so the one thing the helper exists to do no longer type-checked. It shipped in this state in 0.3.6 and 0.4.0.It compiled because it was internally consistent, and no test caught it because this workspace has no callers of it: the helper is public API for downstream consumers, and the only one was pinned to an older release. The regression guard added with the fix is a compile-time function taking
&mut Utterance, so any future drift between the signature and the field fails the build rather than passing silently.On replace, the existing entry’s
TierSeparatoris preserved (only the payload is regenerated, and the separator is the provenance E758 is detected from); on append the new entry isCLEAN. Serialization canonicalizes to a single tab either way, so this affects diagnostics, not output.
0.4.0 - 2026-07-27
Added
-
Three validation rules that catch real, previously-invisible defects in transcript data. All three live entirely in the validator: the grammar, the model’s serialized shape, and roundtrip behavior are untouched.
-
E761,
%grarelation head is not a Universal Dependencies relation. A%gralabel isHEADorHEAD-SUBTYPE; UD fixes the head set at 37 universal relations and leaves subtypes open and language-specific, so the head is checked against that closed set and the subtype is never checked. Nothing validated relation labels before, in chatter or in CLAN CHECK, so a corrupted label rode silently into every downstream analysis that reads the dependency graph. Grounded in a survey of the entire corpus (138,565,864 relation instances across 106,158 files): all 37 universal heads are attested, 150 distinct labels occur, and exactly three heads fall outside the set, all of them defects (IOBforIOBJ,PAD,PUNCTTforPUNCT). -
E762, the prefix marker
#stands alone as a word or opens one. The marker attaches to the END of the prefix it marks, and the prefix is a word of its own (Hebrewha# kelev), so neither shape can be that construct in any language. Language-independent, and zero-attested corpus-wide. -
E763, prefix marker in a language that does not use it. Gated on the WORD’s resolved language rather than the file’s
@Languagesheader, exactly as the digits rule (E220) is, so a code-switched word brings its own rules with it. Languages that write the marker:heb,ara. Word-internal markers stay legal wherever the language allows the marker at all.
-
-
TreeSitterParsernow implements the sharedChatParsertrait (talkbank_model::ChatParser), making the two parser backends interchangeable behind one generic bound at every granularity (file, header, utterance, tier, word, relation). Previously onlyRe2cParserimplemented the trait, so consumers selecting a backend per target (tree-sitter natively, pure-Rust re2c on wasm) had to hand-roll a cfg-gated facade. Every trait method delegates to the matching inherentparse_*_fragmentmethod, so trait-path and inherent-path behavior are identical; conformance is pinned bytalkbank-parser/tests/chat_parser_trait.rs. -
Dedicated error codes for two malformations that previously fell through to the generic E316 unparsable-content catch-all, from the CHECK-parity adjudication of CLAN CHECK errors 52 and 11: E759 (an utterance beginning with a postfix annotation such as
[/],[<], or[: text], which has no preceding material to scope over) and E760 (a%moritem with an empty part-of-speech field,|we). Both are recognized by the tree-sitter front end’s error analysis and mirrored in the re2c oracle’s front end; both files were already rejected, so no validity verdict changes, only the diagnosis.
Removed
-
nextest. CI, the cross-platform workflows and the documented local commands run plain
cargo test, with job time limits in place of nextest’s hang guard. -
TalkBank XML support, in full. The
to-xmlcommand, thetalkbank_transform::xmlemitter, thecorpus/reference-xml/golden corpus, thexml_goldenandxml_schema_validatesuites, the bundledtalkbank.xsd/xml.xsdschemas, the XML Emitter book chapter, and thequick-xmldependency.TalkBank stopped generating TalkBank XML on 2025-10-29, when its last consumer said he no longer used it, and the published
data-xml/distribution has been offline since. Phon moved off the format some time ago. Nothing produced by this emitter had a consumer.Breaking:
chatter to-xmlno longer exists and there is no replacement. Usechatter to-json, which is the format the toolchain actually maintains.talkbank_transform::xml::XmlWriteErroris gone from the public API surface.
Changed
- The
%gradocumentation, examples, reference corpus, and error-spec fixtures no longer use retired TalkBank relation labels (SUBJ,JCT,POBJ,COM,VOC,MOD,NEG,PRED,COMP,ADV,INCROOT,QUANT,LINK), which E761 now rejects. None of them occurs anywhere in the real corpora; they were fixture inventions that would have taught readers of the API docs a vocabulary the validator rejects. Replaced throughout by the UD relations the corpora actually use (NSUBJ,OBL,CASE,DISCOURSE,VOCATIVE,AMOD,ADVMOD-NEG,CCOMP,EXPL,DET,DEP). Thegra_incrootgrammar construct deliberately keepsINCROOT: it pins the property that relation labels are open text at the grammar layer, which is why the vocabulary is a validation policy and not a syntax.
Fixed
-
Word validation now reaches words nested inside groups. Main-tier validation iterated content items flatly and matched only
Word,AnnotatedWordandReplacedWord, with a catch-all that silently discarded every container, so a word inside a retrace, a reformulation, an angle group or a quotation was never word-validated at all.The symptom: the identical token was rejected outside a group and accepted inside one. In English
hello3 dog .was invalid (E220) whilehello3 [/] hello dog .was valid, on every release up to this one. Every word-level rule inherited the hole, so E220 has carried it for as long as the rule has existed; the newer prefix-marker rules inherited it on arrival.Corpus impact, measured over all 106,158 files: 341 to 348 invalid files, 8 new error instances across 7 files (E241 x2, E252 x4, E248 x1, E763 x1), each a pre-existing data defect that had been hiding inside a group rather than any change in what counts as valid CHAT.
-
ErrorCollector::is_empty()violated the standard Rust contractlen() == 0 <=> is_empty(): it answered “is the internal buffer unallocated?”, so a collector created withwith_capacity(which pre-allocates) reported non-empty while holding zero errors. Found by the 1.0 contract-set API audit; now implemented aslen() == 0with a regression test. -
TreeSitterParser::parse_gra_relation_fragment(and the trait’sparse_gra_relation) rejected EVERY bare%grarelation and leaked a spurious E709 diagnostic into the caller’s sink, because the wrapper appended a scaffold terminator with the never-valid index 0 (0|0|PUNCT) and the tier wrapper rejects on any internal diagnostic. The scaffold is now valid CHAT (2|1|PUNCT), and a scaffold-region filter guarantees diagnostics against wrapper scaffolding can never reach the caller. The re2c backend was unaffected (it parses the relation directly); the fix restores backend agreement. Caught by the newChatParsertrait conformance test. -
Validation cache: initialization is now concurrency-safe across processes, not just threads. Every opener takes an exclusive advisory file lock (
talkbank-cache.init.lock, beside the database) around first-time create + migrate, so parallelchatterruns (or parallel test processes) sharing one cache directory can no longer race sqlx’s SQLite migration (UNIQUE constraint failed: _sqlx_migrations.version, the 2026-07-13 flake) or collide on first-connection WAL setup. Lock acquisition is bounded: on timeout, opening fails with the new typedCacheError::InitLockTimeoutand the CLI degrades to running uncached instead of blocking. The 2026-07-13 bounded retry is retained as a backstop for older builds that share the cache directory without honoring the lock protocol. Regression coverage: a cross-process stress test races 8 processes over a fresh cache directory for 4 rounds under a hard deadline, so both failure modes (constraint error and hang) fail the suite instead of flaking or wedging it. -
Desktop release: the macOS updater bundle is now uploaded under a per-arch asset name (
Chatter-<target>.app.tar.gz). Previously both the aarch64 and x86_64 macOS jobs uploaded the arch-independentChatter.app.tar.gz, which raced on the shared release asset (the v0.3.6release-desktopupload failure) and pointed both darwin entries inlatest.jsonat a single URL holding one arch’s binary. Fresh.dmgdownloads were unaffected; the desktop auto-updater is the surface this corrects. (Ships with the next release.) -
Public API: a downstream crate that depends only on
talkbank-parsercan now name the error type of its parse methods. The sixTreeSitterParser::parse_*methods returnParseResult<T> = Result<T, ParseErrors>, butParseErrors/ParseResultwere not reachable from thetalkbank-parsercrate root (only via apub(crate)module), forcing consumers to add a separatetalkbank-modeldependency or stringify at the boundary; both are now re-exported. Also re-exportedtalkbank_model::SylWordError(the error ofclassify_syl_word/tokenize_syl_word), which was omitted from the model root while its sibling phon parse-error types were present. Completes the BUG-3 audit: a compile-test now names every public fallible constructor’s error type so this class cannot regress.
0.3.6 - 2026-07-17
Fixed
- The Phon
%x-tier content checks (introduced with the %x fold-in) no longer mass-flag valid Phon exports. Two wild-corpus conventions the original specification never confronted are now accepted: (1) pause fillers ((.),(..),(...)) mirrored at the same word position on%mod/%pho/%xmodsyl/%xphosyl(and as pause pairs on%xphoaln) to keep word-aligned tiers in index lockstep, which E735 previously rejected as malformedphone:CODEunits (roughly 13,000 spurious errors across the PhonBank corpora); and (2)^and IPA.syllable-boundary notation in%mod/%phowords, which the segment-level%xphoalnreconstruction comparison now ignores exactly as it already ignored stress markers (roughly 770 spurious E740/E741). Genuine misalignments (index-shift chains, pause fillers standing in for real words) are still reported. Users who adopted--suppress xphonto silence the storm can remove it and regain the genuine%x-tier checks. - Generated error-documentation pages (
docs/errors/) no longer fuse words across wrapped spec lines or drop backticked text: the spec text extractor now renders soft line breaks as spaces and includes inline code spans.
Added
-
New validation rule E752: timing bullets without an
@Mediaheader. A transcript carrying timing evidence (utterance bullets or%worword timing) must declare the media those timestamps index; completes the media-consistency family (E544: declared linkage without timing; E552: declaredunlinkedcontradicted by timing). Mirrors CLAN CHECK error 112. -
New validation rule E753: a word consisting only of a repetition segment (fully
↫...↫-wrapped, no stem outside the delimiters) is rejected; word-category prefixes (&-filler,&~nonword,0omission) count as a stem. Adopted from GUI CLAN CHECK error 151 as a chatter-authority rule (the unix CHECK build never enforced it). -
New validation rule E754: the
@lletter form must carry exactly one letter of stem (b@l); multi-letter content belongs under@k/@ls. Repeated-segment material (↫b^↫b@l) does not count toward the stem, matching real CLAN CHECK behavior. Mirrors CLAN CHECK error 76. -
New validation rule E755: a
[- CODE]utterance-level language must be declared in@Languages(utterance-level presence is substantial). Mirrors CLAN CHECK error 152. -
Word-level explicit language codes (
word@s:CODE) are now validated against the ISO 639-3 registry (E519), the same rule that guards@Languagesand@ID; declaration in@Languagesremains not required. -
@L1 ofvalues are now typed ISO 639-3 language codes and validated against the registry (E519), completing registry validation at every position language codes appear. Wild usage was already uniformly codes; generation viabuild_chatnow takes aLanguageCodefor the participant first language. -
E756 (empty user-defined
%xtier) replaces W601: the rejection is unchanged; the old code fired as a hard error despite its warning prefix, so the number was the bug. The diagnostic message also no longer double-prefixes the tier name (%xfoo, not%xxfoo).
Removed
W210andW211are retired, and their numbers are not reused. No production code path emitted either; CLAN CHECK accepts W210’s glued-terminator construct, and W211’s shape is valid overlap-hugging CA notation. The JSON schema no longer lists them.- The E254 warning (word-level
@s:CODEnot listed in@Languages) is retired: an explicit word-level language code is self-contained and deliberately carries no declaration requirement.@Languagesdeclares the transcript’s substantial languages; a one-word insertion is not substantial presence. (This matches CLAN CHECK, which dropped its own@sdeclaration requirement in 2019.)
0.3.5 - 2026-07-15
Emergency release restoring corpus-correct word parsing. Versions 0.3.3 and 0.3.4 have been YANKED (releases and tags removed).
Fixed
- Reverted the whitespace-boundary overlap-custody grammar introduced
in 0.3.3. Its GLR-arbitrated word readings fragmented words carrying
four or more glued markers (for example multi-syllable-pause chains
like
or^ga^ni^zi^ra), causing spurious E252/E331/E600/E705 validation errors across real corpora and, worse, a serialization mutation (a space inserted into such words on rewrite). Word parsing is restored to the 0.3.2 grammar, verified by an error-code differential and a roundtrip comparison against the 0.3.2 binary over a corpus sample: identical profiles. - A regression test pins that multi-marker words parse as one word and validate cleanly.
Retained from the yanked releases
- Typed
@uphonetic word forms (UNIBET). - The
build_chatheader emitters and @ID demographics fix. - The shared English capitalization transform.
- The long-tier stack-overflow fix and its regression test.
- The SQLite cache concurrency-safety fix; CI runs under nextest.
0.3.4 - 2026-07-15 [YANKED]
Added
-
@uphonetic forms are now typed phonetic content. A@uword (a UNIBET/IPA phonetic transcription standing in a word slot, e.g. the spoken side of an aphasia[: target]replacement) now models its content as a dedicatedWordContent::Phonetic(WordPhonetic)node instead of orthographic text, in both parsers. Orthographic word-hygiene rules structurally cannot apply to phonetic content; the phonetic string itself stays deliberately lenient (IPA, ASCII UNIBET, X-SAMPA), matching the%photier’s stance.to-jsonemits{"type": "phonetic", ...}for these nodes (schema updated);cleaned_textremains the phonetic string verbatim; the sanitizer redacts phonetic forms like spoken text. Scope is@uonly; sibling special forms remain orthographic words. -
build_chatnow emits the full standard header set. The general CHAT-generation schema (TranscriptDescription/ParticipantDesc) gained typed optional fields for@Date,@Situation,@Options,@Transcriber,@Comment, per-speaker@L1 of, and@PID(preserved from a source, never minted), each emitted in canonical header order.@IDdemographics (age, sex, group, SES, education, custom) are now carried throughParticipantDescinstead of being silently dropped, fixing empty demographic slots in generated@IDheaders. -
Shared English capitalization transform (
talkbank_transform::capitalize): capitalizes the pronoun “I” family and the first real word of each utterance on the typed model, for generators whose sources are all-lowercase (improves downstream%moraccuracy). Token-level helpers are public for generators that capitalize their own word representation.
Fixed
chatter validateno longer headlines a warnings-only file as an error. A file whose findings are all warnings (which is valid CHAT, and was already counted valid in the summary) now prints⚠ Warnings in <file>instead of the contradictory✗ Errors found in <file>, and the “fix structural errors first” hint fires only on hard errors. Presentation only; validation logic unchanged.- The validation cache no longer fails to initialize when opened
concurrently. Two
chatterruns sharing a cache directory (or a multi-threaded consumer) could race the one-time SQLite setup and hitUNIQUE constraint failed: _sqlx_migrations.versionor a WAL init collision, silently disabling caching for that run. Concurrent opens on a fresh cache directory now retry the transient init race and all succeed.
0.3.3 - 2026-07-13 [YANKED]
Added
- Desktop app: a “Check for Updates…” menu item and a periodic background update check. The app previously checked for a new release only at launch, so an app that was rarely relaunched could sit far behind. It now also checks every six hours in the background, and the app menu has a manual “Check for Updates…” item that reports when you are already up to date.
- Desktop app: a real “About Chatter” panel with the version, a short description, and clickable links to the TalkBank site and the source repository, replacing the bare version-only default.
talkbank_transform::build_chat: assemble a validated CHAT file from a typed transcript description. Given participants, optional media, and utterances as pre-formatted CHAT main-tier text (TranscriptDescription), it synthesizes the header block, parses each utterance through the tree-sitter parser, and returns aChatFile. The description carries amedia_status, so a transcript that names its media but has no timing bullets yet (pre-forced-alignment) can emit@Media: <id>, audio, unlinkedand stay valid instead of falsely claiming linkage (E544).talkbank_transform::num_words::expand_number: spell digit tokens as language-appropriate number words (13 lookup-table languages, CJK, and English ordinals/decades), so generated CHAT satisfies E220 (numeric digits are not allowed in words for languages that do not permit them).
Changed
- Overlap custody now follows whitespace boundaries, with canonical overlap serialization. Overlap markers bind to the token on the correct side of a whitespace boundary, and serialization emits a single canonical form.
- tree-sitter updated to 0.26.11 across the workspace (CLI, grammar bindings, and the generated parser).
Fixed
- Long dependent-tier reconstruction is now linear-time. A quadratic blowup on very long utterance tiers is eliminated; pathological inputs that previously stalled the parser now reconstruct in linear time.
- Desktop app: the validation settings popover no longer opens hidden behind the results panel. It was rendered below the panels in the stacking order; it now sits above them.
- Desktop app: the “up to date” dialog now dismisses on the first OK. A listener leak (an async menu subscription whose cleanup could run before it resolved) let duplicate listeners accumulate, so one menu click stacked several identical dialogs.
0.3.2 - 2026-07-10
Added
chatter rediarize: repair speaker attribution from external diarization turns. Takes a transcript whose utterance timing is trusted but whose speaker labels are not, plus a speaker-turns JSON file ({"source": ..., "turns": [{"track", "start_ms", "end_ms"}]}) from an external diarizer, and re-attributes each timed utterance to the dominant overlapping turn. Utterances with no turn coverage are flagged, never guessed. Reconciled@IDrows are inserted in the header block.--summary-jsonemits a machine-readable outcome summary (per-utterance reattributions and flag reasons) for downstream tooling.- Four validation rules for constructs that do not make sense, each adjudicated against real CLAN CHECK behavior and the wild corpus: E748 leading-zero media-bullet times; E749 comma glued to the following word; E750 whitespace inside angle-group delimiters; E751 pause marker glued to a word.
Fixed
- The re2c oracle lexer now tokenizes short-form parenthesized material the same way the canonical parser does (its catch-all previously swallowed a trailing delimiter), keeping the two independent parsers in cross-check agreement on the new spacing rules.
Changed
- Rust toolchain pin bumped to 1.97.0 (CI workflow pins synced); workspace and spec lockfiles refreshed; desktop dependency bumps (jsonschema 0.47, TypeScript 7).
- Documentation: an architecture page on overlap-marker binding (why
edge-adjacent overlap markers bind into words, the ideal top-level
model, and the conversion-layer path); the grammar’s empty-
extras(all-whitespace-explicit) design rationale is now recorded at the declaration site.
0.3.1 - 2026-07-08
Fixed
- Every public fallible constructor’s error type is now publicly
nameable.
LanguageCodeError(fromLanguageCode::new),XphointParseError, andPhoalnParseErrorwere not re-exported, so downstream crates could not store them in typed#[source]fields and had to stringify at the boundary; found by the first real downstream consumption of the 0.3.0 API. A new API-surface guard test pins the contract so a constructor error type can never silently become unnameable again.
0.3.0 - 2026-07-07
Added
--llm-cache <file>(envCHATTER_LLM_CACHE) for holistic speaker-id judgment. A persistent, write-through JSON response cache forspeaker-id/pipeline/batch --judgment holistic: an identical request (same endpoint, model, and rendered prompt) is served from the cache instead of making another LLM call, so re-running a batch after a crash or an unrelated code change does not re-pay completed sessions. Absent flag and env variable means uncached, unchanged from before.
Fixed
chatter batchno longer reports holistic suggestions as merges. In holistic-judgment mode the per-session pipeline exits 0 after writing a suggestion to the pending file without merging (the operator adjudicates first); the batch summary counted those as “merged” and reported zero pending work. Outcomes are now classified by whether the merged output actually exists, and the summary separately counts merges, suggestions awaiting adjudication, and low-confidence refusals awaiting adjudication.- E552 (
@Mediasaysunlinkedbut timing exists) now says where the timing was found and how to fix it. When the only timing evidence is word-level bullets inside a%wortier (invisible in normal display), the message names the%wortier and offers both remedies (the media is in fact aligned: removeunlinked; or the%wortier is stale: remove it) instead of asserting the media is linked and pointing at bullets the user cannot see. The main-tier-bullet case keeps its direct advice. - Chatter Desktop’s single-file validation now shares the CLI’s validation
engine. Previously, validating a single
.chafile in the desktop app (as opposed to its parent folder) bypassed the on-disk cache entirely, skipped the@Media-filename check (E531), and could not honor--roundtrip/--parser/--strict-linkers. All of these now work identically tochatter validateand to the desktop’s own folder validation, and a new Settings panel exposes the equivalent options. - Chatter Desktop no longer shows “N files, all valid” before a run has actually finished. The file tree previously derived this message from the partial, still-streaming result set, so it could flash “all valid” mid-run whenever no error had streamed in yet.
0.2.1 - 2026-06-24
Added
- The
talkbank-lsplanguage server now ships as a standalone release artifact. Prebuilt, code-signedtalkbank-lspbinaries for macOS (Apple Silicon and Intel), Linux (x86_64 and aarch64, static musl), and Windows are attached to the GitHub Release, each with its owntalkbank-lsp-installer.sh/talkbank-lsp-installer.ps1. Any LSP-aware editor can now install the server without building it from source; it is a first-class artifact in its own right, not only the binary the VS Code extension bundles per platform.
0.2.0 - 2026-06-23
Added
- More of CLAN CHECK’s invalidity is now enforced. A batch of CHECK-parity
rules was implemented so
chatter validaterejects more invalid CHAT:E514: an@IDline’s corpus field is required (CHECK 63).E547: a constant participant header must follow the@IDblock.E548: closes the case CHECK 126 covers.E549: a speaker may not be declared twice (CHECK 13).- Duplicate
@IDlines and out-of-order@Optionsfields (CHECK 13, 125). - A dependent tier used without being declared (CHECK 17).
- An out-of-range
@Time Duration(CHECK 35). - An
@Mediaheader marked unlinked while the transcript still carries timing bullets (CHECK 124), and an@Mediafilename that does not match the data file (CHECK 157). - A replacement
[: ...]now requires a preceding space (CHECK 161). - Tree-sitter recovery nodes are surfaced as invalidity rather than silently
repaired: a surviving
ERRORnode maps toE316and aMISSINGnode toE342(with the re2c oracle mirroring it), covering a group with no annotation and swallowed recovery nodes inside comma-list headers (CHECK 5/6/106/108).
- Phon:
U(unknown) is accepted as a legal syllable-constituent code on the%xmodsyland%xphosyltiers. - A formal behavioral CHECK-validity parity test suite that runs real CLAN CHECK and chatter on the same fixtures and fails if either side drifts.
Changed
-
chatter updatenow self-updates in process. It embeds the axoupdater self-updater as a library, reads the cargo-dist install receipt (keyed by the package name), and replaces the running binary from GitHub Releases. This removes the package-name coupling that previously madechatter updatereport “not installed” on a correctly installed binary. -
The CLI package is renamed
talkbank-clitochatter(the crate now lives atcrates/chatter/). The generated install scripts are thereforechatter-installer.shandchatter-installer.ps1(previouslytalkbank-cli-installer.*); update any pinned install URL accordingly. The binary is stillchatter, and the library/API crates keep theirtalkbank-*names. -
Validation is stricter. Because of the new CHECK-parity rules above, some files that passed
chatter validateunder 0.1.1 may now report errors. This is intended: chatter is the CHAT-validity authority and is at least as strict as CLAN CHECK. -
Word-level explicit language codes (
word@s:CODE) are now validated against the ISO 639-3 registry (E519), the same rule that guards@Languagesand@ID; declaration in@Languagesremains not required.
Removed
- The standalone self-updater binary (cargo-dist
install-updater = false). Thechatter updatesubcommand is unchanged for users; it now updates in process instead of shelling out to a separate program.
Fixed
- The recovery-node invalidity backstop is scoped to localized errors so it does
not over-flag, and several malformed
@IDtest fixtures were corrected. - Hardened the CHECK-parity audit and corrected a CHECK 126 verdict it had falsely certified; the curated CHECK error-code map is restored in place of a brittle keyword heuristic.
0.1.1 - 2026-06-22
Fixed
- Validation cache could serve a stale verdict across rule-set changes.
chatter validatekeyed its result cache on the cache crate’s package version, which does not change when validation rules change, so a “Valid” result cached before a new rule (such as a retrace-marker check) existed kept being served, while a fresh conversion of the same bytes correctly rejected them. The cache key now folds in a fingerprint over every error-code rule, so adding, removing, or renaming any rule invalidates stale entries; the cache is kept and still functions, only keyed correctly. - CLI usage lines pin the binary name to
chatterregardless of the invoked path (clapbin_name). - The book renders Mermaid diagrams again (restored mdbook-mermaid assets).
- Desktop app version is now locked to the release version. The desktop
bundle (
.dmg/.exe/.deb) and the Tauri auto-updater manifest now report the same version as the CLI. A version-sync gate (scripts/sync-app-version.py, enforced in CI and at release time) keepstauri.conf.json,package.json, the workspace version, and this changelog from drifting, so the updater can never again advertise a version the installed bundle does not match.
Changed
- CI book toolchain bumped to mdBook 0.5.3 and mdbook-mermaid 0.17.0.
- Build: force
serialize-javascript >= 7.0.5to clear advisories, and bumprandin the spec crate. - Docs: the book intro is de-staged for the public release (download-first).
0.1.0 - 2026-06-15
First public release.
Added
- CHAT-format core. A strict, incremental tree-sitter parser
(
talkbank-parser) with an independent re2c oracle parser (talkbank-parser-re2c) that cross-checks it on every file; a typed CHAT data model with structured validation, error codes, and tier alignment (talkbank-model); and CHAT-to-JSON / JSON-to-CHAT / XML conversion, normalization, transcript-merge, and redaction pipelines (talkbank-transform). - Phon extension tiers. The four Phon
%xdependent tiers (%xmodsyl,%xphosyl,%xphoaln,%xphoint) are parsed and validated as first-class CHAT tiers, on by default (pass--suppress xphonto opt out): syllabification constituent codes and phone-vs-source reconstruction, model-to-actual phone alignment, and per-phone time intervals, with dedicated error codes. chatterCLI.validate,normalize,to-json/from-json/to-xml,merge,speaker-id,batch,pipeline,adjudicate,sanity-scan,lint,clean,watch,new-file,show-alignment,validate-utseg,schema,update, and a content cache.- Language server (
talkbank-lsp): real-time validation, hover, go-to-definition, and cross-tier alignment for any LSP-aware editor. - Desktop app (
Chatter): a Tauri-based CHAT validation app, shipping in the coordinated release alongside the CLI. - Auto-update. The
chatterCLI self-updates withchatter update(the bundled cargo-dist / axoupdater self-updater), and the desktop app checks for and installs new releases on launch (Tauri updater). Both pull from GitHub Releases. The CLI self-updater is experimental. - Prebuilt binaries for macOS (Apple Silicon and Intel), Linux, and
Windows, plus desktop installers, attached to the GitHub Release. The
macOS desktop
.dmgis signed and notarized.
Known limitations
- The merge and adjudication surface is experimental.
merge,adjudicate,speaker-id, andsanity-scanwork, but their interfaces and heuristics may change before 1.0. - Windows binaries are not code-signed yet, so Windows SmartScreen warns on first run (choose “More info” then “Run anyway”). macOS CLI binaries are codesigned but not notarized; install via the release installer script to avoid the Gatekeeper quarantine prompt.
- Not on crates.io yet. crates.io publication is deferred.
Earlier changes (moved from the book)
The book states the current design only. These entries record fixes and behaviour changes that its pages used to narrate. They are not attributed to a release because the book did not record one; dates are those the book gave.
Spec system
- The machine-written
_autoerror specs are gone.corpus_to_specsandenhance_specswere deleted (spec-system redesign, R5), andspec/errors/carries a.human-authoredmarker that every generator refuses to write into. In August 2026, 152 of 238 error spec files were*_auto.md; they recorded what chatter did rather than what it should do, because the tool wrote the spec’s own filename code whenever its (never existing)expectations.jsongave no codes. By 2026-09-03 one_autofile was left, and it has since been merged intoE519.md. - Duplicate spec files for one code were reconciled on 2026-09-03: the
residue pairs (E202, E241, E604) were deleted,
E243_auto.md’s example was re-filed under E202, and E316, E342, E375, E522, E360 and E502 were each merged into oneE###.md. That changed the keys the re2c parity baseline uses, soKNOWN_DIVERGENCESwas regenerated. Categorywas removed from specs on 2026-08-19 (a free-text grouping nothing read);Levelmoved from the file onto each example (Phase 2, 2026-08-21); all error specs moved to+++TOML frontmatter (Phase 1b, 2026-08-21).- Every example now carries a required typed
claim(R2, 2026-08-21), which deletedSpecSelfDemonstrationGateand its 36-entry baseline. The authoredlayerfield was deleted (R4, 2026-08-21), together with the string-based error tests: the authored field disagreed with the observation snapshot on 17 examples.kindandstatusmoved from each spec file to the code registry (R1, 2026-08-26), which removed thespec_statusgate, thespec/errors <-> ErrorCodedivergence check and the per-codekindagreement loop.statusno longer defaults toimplementedwhen absent (this had been true of 104 of 238 specs on 2026-08-11). A retired code number reused is now a registry load error, replacing a comment in the enum. - The
docs/errors/*.mdpages are a registry artifact written byjust spec-gen; the standalonegen_error_docsbinary was deleted. The hand-written artifact table in the spec chapter was replaced by one generated from the registry.grammar/test/corpus/generated/andmanual/are separate trees; they had been one tree, which destroyed 1,468 lines of hand-mined corpus tests twice in three days. - The backend-parity harness now carries each example’s declared source path, so contextual rules such as E531 run for both parsers; dropping it had made both backends appear to miss E531.
Reference corpus overhaul (Phases 0-6)
- The reference corpus grew from 345 English-only files to 374 files in 20
languages, then was reorganised into nine topical subdirectories under
corpus/reference/. At the end of the overhaul: concrete grammar node coverage went from 316/334 (94.6%) to 334/334 (100%); error specs from 177/181 to 181/181 (169 with CHAT examples, 12 documented stubs); the golden artifacts were regenerated andreference_corpus.rsrebuilt with 374 cases. - Phase 0 built
corpus_node_coverage(it confirmed 18 uncovered node types); Phase 1 builtextract_corpus_candidatesand selected 25 files across 20 languages (eng, zho, fra, deu, spa, jpn, nld, heb, por, ell, tur, hrv, pol, ita, hun, rus, est, dan, ara, isl); Phase 2 added four handcrafted files inconstructs/(rare-terminators.cha,uptake.cha,best-guess.cha,unsupported.cha) for the 18 gaps; Phase 3 ran batchalign3 morphotag over the language files; Phase 4 created error specs E707, E711 and E717, corrected E376’s recorded code, filled 17 triggerable stub specs, documented 12 untriggerable ones (E001, E002, E211, E317, E318, E340, E374, E377, E378, E380, E385, E386), corrected 5 misclassified specs (E319-E322, E376) and builtperturb_corpuswith 11 mutation strategies. - Mining the MacWhinney subcorpus (407 files) found zero tree-sitter parse errors, and mining all of Eng-NA took over four minutes; that is why perturbation is the systematic route.
- The former Chumsky direct parser could not handle
unsupported_linenodes (373 of 374 roundtrips passed under it); it has been removed and tree-sitter is the sole canonical parser.
Chatter 1.0 readiness work log (2026-09-05 onward)
- Baseline on 2026-09-05: 223 error specs (179 implemented, 37 not implemented, five unreachable from CHAT, two deprecated); of 418 examples, 368 satisfied their claims, 50 were deferred and none failed. The CHECK mapping audit stopped inferring parity from code names and reading an obsolete source path; it reads the compiled registry. Real CHECK grounding passed (12.59 s and 12.45 s on two runs) over the committed fixture set.
- E246 was marked implemented (its reachable
(he):example emits E246 and E209;hel:oand(he)l:oare legal controls). E212 gained a whole-file violation and a CA-mode legal control (an explicit category prefix prevents the CA normalizer from rewriting0(the); the standalone shortening reports E209 and E212), and its originalhello world .example now declareslegal. E251’s malformed word sample moved to E342 as an active missing-element example (tree-sitter reports E255/E342; re2c reports E209/E253/E255; neither emits E251), raising verified examples to 365 and lowering deferred to 51. Wordprosodic checks use the privateProsodicWordmeasurement and two linear passes with constant-time neighbor queries (each marker used to search a prefix or suffix again); E244-E252 behaviour and diagnostic order are unchanged.JsonSchemaPolicyselects serialization after the shared named CHAT parse. A single-file conversion with schema validation skipped used to succeed on a mismatched media name; both policies now reject mismatches. Directory conversion exposes failing diagnostics and returns exit 1.just regennow refreshes the JSON schema (model documentation is embedded in it). The schema writer preserves unchanged files; generation is the explicitjust schema-gen. The generator no longer rewrites$refsiblings intoallOf(Draft 2020-12 allows siblings; the transform rewrote literal data as though it were a schema and its removal dropped 74 wrappers and five shape-only tests).just traversal-genand the other generated-Rust recipes stage output throughscripts/generate_if_changed.py;tree-sitter generateruns into a staging directory with--outputand--checkis read-only.- The spec artifact writer used to delete every current named output before
writing;
Ownership::NamedFilesnow deletes only retired names.GeneratedDirreplacedclear_ownedand proves ownership before pruning. A bounded nextest trial did not improve the warm generator suite. - The foundation publication check derives the held-back set from Cargo
metadata and requires
publish = false; it rejectedtalkbank-llm, which was publishable by default and now holds publication back explicitly. - Diagnostic indexes: editing a string in place reused stale line positions,
and another source’s line map panicked on a multibyte character.
SourceIndexreplaced the publicenhance_errors_with_line_map(a breaking library API change) and the hidden thread-local cache. - The re2c parser separated source lifetime from temporary token-storage
lifetime, removing every production
Box::leak;SinToken::new_uncheckedwas removed andSinTier::from_tokensreturnsResult. - The hygiene scanner mistook a nested
fn report(...)declaration for a call and missed a call with whitespace before(; both are fixed with a regression.
Word grammar
standalone_wordhad at one point been coarsened into one opaque DFA token, with a Chumsky direct parser re-parsing it intoWordContent. That cost two parsers with independent bugs, validation that could not find markers without re-parsing, one opaque editor node, and acleaned_text()that scanned for marker characters. When the Chumsky parser was eliminated the structured word grammar was restored (markers re-excluded fromword_segmentvia the symbol registry, one CST child per marker,WordContentaligned 1:1 with grammar nodes, the purity invariant made a gate).- In commit
fdceeac2the consolidatedword_segment_purity.txt(8 named tests) was replaced by per-construct test files generated from the specs.
Parser
- The editor uses the
ParsedRevisioncache (parse_chat_file_revision); the LSP no longer owns its own edit calculator or a separately replaceable source/tree pair. The raw CST, strict-model and streaming incremental methods remain as compatibility entry points. Tree-sitter incremental parsing no longer clones the whole-document tree. - E311 (parser-only unclosed-replacement diagnostic), the
UnclosedDelimiterwrapper and its E312/E313 text classifiers, and theBracketRecoveryclassifier (which inferred annotations from text prefixes) were removed; the malformed inputs remain rejected by grammar recovery, with generic E316 where no structural evidence supports a narrower fault. The raw-argument error collector and its recursive wrapper were removed. - A dummy replacement span used to conceal missing separators at either closing bracket; replacement producers now retain the whole CST wrapper’s span.
- A missing
@UTF8anchor is an optional grammar slot: the canonical parser used to discard the whole document and report the present headers as missing; it now keeps them and shared validation rejects the file with E503. - Header lowering used a start-only check that accepted the first of two
headers and discarded the second;
HeaderFragmentadmission now requires the node to account for all caller text. Routing only@PIDthrough the shared pre-@Begindecoder had let@Window,@Color wordsand@Fontfall through to successfulUnknownvalues. - Fragment adapters no longer use the legacy error sink’s length heuristic.
Symbols and CA terminators
- A Chumsky-based direct parser provided combinator-based fragment parsing until it was removed in March 2026; tree-sitter is the sole parser.
- The derived symbol arrays were named
ca_delimiter_symbolsandca_element_symbolsuntil 2026-08-25; they arepaired_stretch_symbolsandword_attached_symbols. The two arrays used to be hand-written and need a disjointness check, which was deleted when they became derived from oneparse_rolefield. Before 2026-08-12 no gate compared the symbol outputs against the registry;generated_symbol_sets_are_currentwas added then, and on 2026-08-20 its hand-written generator list became a glob ofspec/symbols/generate_*.js. Its first run found two rustfmt-wrapped Rust outputs that the generator unwrapped. - The parser/model used to promote trailing CA markers into utterance
terminators through a post-hoc
resolve_ca_terminator()pass; the pass was removed and CA arrows and≈/≋staySeparatorcontent items. - The validation page no longer describes one downstream consumer’s server behaviour, PyO3 boundary types and report directory, and no longer calls the reference corpus “the sacred semantic target”.
re2c backend
- As of 0.19.0 re2c-parsed values borrow the caller’s source (the earlier
Box::leakstrategy is gone). The re2c newline token used to fuse consecutive breaks and lose blank-line structure. The main-tier-only whitespace scan and its separate CA probe were removed in favour of the shared file validator.WordLengthening::countand the re2c AST moved fromu8toNonZeroUsize(source runs over 255 colons could overflow or silently wrap, and zero counts were repaired withmax(1)on serialization); JSON now accepts longer runs and rejects zero. - Postcode, glued-replacement and separator silence recorded in older parity reports are fixed. The “both parsers correct” claim for CI validation is not currently true.
- Stored benchmark timings (tree-sitter vs re2c): small file (13 lines) 44 us vs 9.6 us (4.6x); medium file with dependent tiers 69 us vs 9.4 us (7.3x); large file 7,734 us vs 970 us (8.0x); batch of 35 files 21.7 ms vs 3.0 ms (7.2x). Older wild-corpus percentages and a 140-case diagnostic table were dropped as not describing the current spec suite.
Correctness architecture work log (2026-09-08)
- The first session of work against the correctness plan: the workspace guard
that permitted one named test at a time was removed (the whole suite ran in
31 seconds for 3,048 tests); 368 dead snapshots then the last 56 were
resolved and
snapshot-hygienejoinedgate(the 56 were 47 stale.chastems from a reorganised corpus layout, 4 from asnapshot_testsmodule rewritten to plain assertions, 4 superseded copies and 1 naming a missing file); undemonstrated error codes were corrected from fifty-three to ten (219 codes carried a spec, 166 demonstrated, 44 excused by registry status, five of the ten closed that night); 107 fabricated-AST constructions becameWord::simple(66 others pass two different strings); the grammar corpus was found to be 211 generated cases (233 recorded, later re-derived as 139 specs) expecting what the parser produced, with 137 of 138 construct specs carrying acstblock nothing asserts (41 naming nonexistent node types, 43 containing...). - Review of the first probe mechanism found
Outcome::Clean(String, Examined)forgeable, a suite of one control probe satisfying every check, a tier axis with one reachable value (everyPreconditiondeclaredPrePush), an absent directory minting the same witness as an empty one, and a second-tree hole closed bygate_discipline. Both Python ratchets (fabricated-AST and demonstration) became gates and six script files were deleted. A final review found the re2c fragment entry points passing the caller’s raw sink into%gralowering (a head overflow reported at byte 2 instead of the caller’s offset) andjust gaterunning at the inner-loop tier because the tier was exported fromtest; the tier became the_testrecipe argument. - The probe run was 5.3 s of a 13.7 s
just test; threads made the standalone binary 4.5x faster (1.2 s) butjust testslower (16.6 s) and were reverted; a shared read cache shipped (loop 12.0 s). The ratchet that counted comments scored prose as the hazard (505 and 601, of which 54 were prose) and was corrected on 2026-09-08. Coverage measured 2026-09-08 over the whole suite: validation 89.2% reported vs 47.7% parse-backed; model 83.8% vs 43.0%; parser 69.3% vs 68.8%; transform 89.5% vs 89.5%. E370 located its marker through a second main-tier serializer; the parser started recordingRetrace::marker_spanthat day and the renderer was deleted. - The
uncovered_branchesJSON field was renameduncovered_region_starts. The repository-root error-corpus generator resolved one parent too many and wrote outside the repository (66 files found beside it); it was fixed to refuse a root without the manifest directory.
Contributor workflow documentation
just pushonce ran four fast checks (and at another time no tests at all) under a comment claiming to be the full CI gate; a greenjust testwas read as a green gate and CI went red on a doctest. The gate is nowjust gate. The book once told contributors to runcargo checkbeforecargo test(which recompiles the dependency graph twice), listed eightjustrecipes when there were thirty-one, described amake verifytarget as “not yet ported” (there is no Makefile), gave per-file--test <name>targets that had not existed for some time, and listed fewer CI jobs than exist.build.rsonce claimed a CI job verified the vendored re2c lexer; there has never been one. A specific, plausible-looking error message was once defended as a loss when the corruption producing it was fixed. Per-push CI no longer runs clippy or the feature-off build (just release-lint).
Annotations
- Until 2026-08-26
UtteranceContenthad no bareAction, so the parser wrapped every unannotated action in anAnnotatedwith an empty list; across a 106,000-file corpus that was 20,184,072 values claiming to be annotated while carrying nothing (almost all a bare0marking silence in daylong recordings).BracketedItemhad no bareGroup, so an unannotated nested group became anAnnotatedGroupwith an empty list. Two error codes meant to catch the empty case could not: one was disabled because bare[*]is valid CHAT, its number was reused for a rule that was unreachable. Both bare variants were added and the empty state became unconstructible.
Form markers
- The form-marker meanings were corrected wholesale on 2026-08-11 against the
CHAT manual’s “Special Form Markers” table: six had been glossed with
plausible expansions of the letters (
@kas “kinship”,@p“proper name”,@sl“slang”,@sas“second attempt success”,@g“gemination”,@ls“letter sequence”).@awas removed from chatter the same day: the corpus authority had eliminated it from every file on 2024-09-03 together with@eand@lp; the other two were dropped from chatter then and@awas overlooked.
%mor
- Earlier documentation described a “comma-stripping” convention where
PronType=Int,Relbecame-IntRel; the grammar and parser preserve the comma. The UD MOR redesign (2026) removedMorSuffix,MorCompound,MorPrefix,MorSubcategory,AnnotatedChunkandChunkfrom the data model, taking it from about 12 types to 4.
Phon tiers
- CLAN’s dependent-tier definitions added
%phointon September 25, 2026 (clan-infocommitf062b58), closing the earlier missing-declaration issue for the unprefixed tier. Current Phon exports no longer use the leadingxon tier names. The Phon%xchecks were once opt-in via--check-xphon; they are on by default and the flag is a deprecated no-op.
JSON output
- A
word_indexfield in the per-word language output existed until 2026-08-07 and was removed as derivable and misleading. The independent raw/cleanedWordconstructor arguments andset_raw_textwere removed;raw_text()returns an owned string and no raw-text cache can go stale.
Validation rules
- E756 (empty dependent tier) is formerly W601, renumbered because it always
was a hard error. It read only user-defined
%xtiers until 2026-08-15 because the model could not represent an empty standard tier, so an empty%eng:had nowhere to be recorded and the two parser backends disagreed about it.%comand%addwere exempt from its whitespace-only check by accident until 2026-09-08. - E757 once caught only
hello [/]there; it now applies to any item ending in a bracketed code (hello [!]there,bobo [= toy]there). The cause was a parser omission: an annotated word’s wrapper span was left DUMMY at construction, so the glue was invisible to a span-adjacency check and any diagnostic reported on an annotated word pointed at byte zero. - E767: an
@Medialine with a space before the comma used to fail to match, so the header fell back toUnknownand reported E525 alongside E330; the grammar now parses it so the rule can name the space (a change of diagnostic, not of verdict). - E764: nothing reported
dog&-um(two words) before the rule existed. E243 now reports a bare or embedded pipe in a word (CLAN CHECK error 48). The%grarelation-head check was added because neither chatter nor CLAN CHECK validated relation labels, so a typo likePUNCTTrode silently into analyses.
Test infrastructure measurements
- With unpacked debug artifacts disabled, a clean
spec/targetmeasured 586 deps entries, no.rcgu.ofiles and 1.3 GB; the full spec suite took 20.14 s from an empty target and 1.64 s warm. A bounded nextest trial on the generators library (51 tests, one binary, four workers, warm) took about 0.6 s against 0.4 s for Cargo, so Cargo remains the runner. - A measured no-op
just regenpreserved bytes and nanosecond modification times of all 3,815 tracked files, took 8.177 s and compiled nothing; the nextjust testtook 10.625 s with no compilation (2,985 passed, 61 ignored, across 34 test harnesses).
Documentation and CHECK assessment
- The CHECK assessment manifest’s notes had accumulated fix narratives
(“GAP CLOSED 2026-07-09”, “ROOT CAUSE was…”, “E316 until 2026-09-08”,
“previously missed”, “wrongly recorded as not-firing”, and similar); they now
state the current verdict and its grounds. Rulings and their dates are kept.
Facts the old notes carried: the streaming lowering used to drop
@Begin:and@BeginsERROR nodes (the whole-tree recovery backstop now reports E316); duplicate@IDlines were missed until E549;@Time Durationrange was unchecked until E540; E552 was added as the inverse of E544;@Mediafilename E531 was dead through the CLI until the file stem was threaded throughvalidate_single_file_streaming;@zXXXwithout a colon fell through to a bare user-defined label until the@z:colon was required; E242 named only the close-quote case until 2026-09-01;%mormalformed words reported E316 beside E702 until 2026-09-08; CHECK 152 was wrongly in the dead list until 2026-07-15. - The spec-tooling page once described a bootstrap-era pipeline: it referred to
make test-gen(there is no Makefile), listed as open a concern aboutspec/toolscarrying parser/model dependencies (resolved by thespec/runtime-toolssplit), prescribed per-spec metadata that no loader read (ownership,draft/accepted/deprecated), and proposed aninput/ir/emit/validate/syncmodule split and aspec lintbinary that were never built. - The branch-protection required-check list named only four jobs until 2026-07-26, having been written before the wasm, app-version-sync and shellcheck jobs existed.
ParseError::build(...).finish()is infallible:try_finish()andParseErrorBuilderErrorwere removed, and a hand-picked subset of the form markers that once sat in the word-syntax page glossed@sias “signed word” (it is singing;@slis signed language).- A consumer’s ledger cited E754 (
LetterFormMultipleLetters, retired 2026-08-11) in August 2026 for a repair that is still correct.
Compile-time investigation (2026-03, pre-fold)
- The compile-times investigation found that a global sccache
rustc-wrapperwas disabling incremental compilation (2.7% Rust cache hit rate; 36 of 37 compilations non-cacheable), that full DWARF debug info inflated link times, and that third-party crates at-O0ran serde, regex and tree-sitter paths about 10x slower than necessary. Pre-fold measurements on the original ten-crate workspace: clean build about 3-5 min (estimated) to about 39 s; incremental rebuild after touchingtalkbank-modelabout 60-90 s to about 4 s. The 2026-04-28 batchalign3 fold roughly tripled the third-party dependency surface, which made[profile.dev.package."*"] opt-level = 1(and theprofile.testequivalent) prohibitive; both were removed.
API changes recorded in the book
-
ChatDate::Validheld{ day, month, year, raw }public fields; it is nowValid(CheckedChatDate)with private components and the accessorsday(),month(),year()andas_str(). -
The
chatter merge,chatter pipelineandchatter batchcommands were removed from the CLI; the structural library intalkbank_transform::transcript_mergeremains. -
The development-loop section of the CI and release page was written on 2026-08-27 after a single parser fix cost a day to the process around it rather than to the fix. The parity baseline stopped listing E550 and E747 once file and fragment participant recovery agreed and both lexers preserved single logical line breaks.
This page last changed: 2026-06-24 (commit 34abe802). The whole book last changed: 2026-10-07 (commit 5e895791).
Installation
Status: Current Last modified: 2026-08-30 14:11 EDT
chatter targets Windows, macOS, and Linux. There are two ways to
install it: the prebuilt binaries (recommended for most people,
including clinicians and researchers) and a from-source build (for
contributors or unsupported platforms).
Prebuilt binaries (recommended)
Every GitHub Release attaches prebuilt binaries for macOS (Apple Silicon and Intel), Linux (x86_64 and ARM64), and Windows (x64), plus desktop-app installers.
chatter CLI
One-line installers (they download the binary for your platform and place it on your PATH):
-
macOS and Linux:
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/TalkBank/chatter/releases/latest/download/chatter-installer.sh | sh -
Windows (PowerShell):
powershell -ExecutionPolicy Bypass -c "irm https://github.com/TalkBank/chatter/releases/latest/download/chatter-installer.ps1 | iex"
On Windows the binary is not yet code-signed, so SmartScreen may warn on first run: choose More info, then Run anyway. The macOS binaries are codesigned, and the installer above does not set the quarantine attribute, so Gatekeeper does not prompt.
Prefer a manual download? Grab the archive for your platform from the
latest release
and extract chatter onto your PATH. (On macOS, a browser-downloaded
archive is quarantined; right-click the binary and choose Open once,
or run xattr -d com.apple.quarantine ./chatter.)
Verify:
chatter --version
chatter --help
chatter desktop app
The desktop app (“Chatter”) is for people who prefer a window to a terminal. Download the installer for your platform from the latest release:
- macOS: the
.dmgis signed and notarized; open it and drag the app to Applications. No Gatekeeper override is required. - Windows: the installer is not yet signed (same SmartScreen note as above: More info then Run anyway).
- Linux: an AppImage and a
.debare provided.
Updating chatter
chatter keeps itself current so you do not have to track releases by
hand.
-
CLI: run
chatter updateThis self-update runs in-process:
chatter updateembedsaxoupdateras a library and downloads the newest release from GitHub Releases directly, without a separate bundledchatter-updateprogram. (The self-update facility is experimental and works the same way regardless of how you installed the CLI.) -
Desktop app: the app checks for updates on launch and offers to install a new version when one is available.
From source
Building from source needs only a stable Rust toolchain (install via
rustup, which supports Windows, macOS, and Linux).
Node.js and the Tree-sitter CLI are needed only when working on the grammar
or generated artifacts. Use the Node version in grammar/.nvmrc, then run
npm ci in grammar/; its lockfile pins the CLI to 0.27.0 so regeneration
uses the same toolchain as CI.
Clone and install the CLI:
git clone https://github.com/TalkBank/chatter.git
cd chatter
cargo install --path crates/chatter --locked
This installs the chatter binary to ~/.cargo/bin/ (macOS/Linux) or
%USERPROFILE%\.cargo\bin\ (Windows). To update a source install, pull
and re-run the cargo install command above (chatter update is only
for installer-based installs).
Building the libraries
If you are developing with the Rust crates directly, from your chatter checkout root:
cargo build --workspace --all-targets --locked
cargo test --workspace --locked
cargo clippy --all-targets -- -D warnings
See the contributor setup for additional commands.
Directory layout
Everything lives in a single repository:
<your-chatter-checkout>/
├── grammar/ # Tree-sitter grammar
├── crates/ # All Rust crates (talkbank-* + the chatter binary)
├── spec/ # CHAT specification
├── apps/ # Tauri desktop app (chatter-desktop)
└── book/ # Chatter mdBook (this book)
The CLI, grammar, crates, and the LSP/desktop integrations all live in this single repository.
This page last changed: 2026-08-30 (commit 733da964). The whole book last changed: 2026-10-07 (commit 5e895791).
Quick Start
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
This page gets you from zero to productive with chatter in five minutes.
Install chatter first if you haven’t already.
Validate a CHAT file
Check a single transcript for errors:
chatter validate transcript.cha
If the file is valid you get a summary (a cache-statistics block follows it;
use --quiet to suppress all output and rely on the exit code):
=== Summary ===
Total files: 1
Valid: 1
Invalid: 0
If there are problems, you’ll see rich diagnostics with the exact location
and a stable error code. For example, a *CHI: line missing its terminator:
✗ Errors found in transcript.cha
E305 (https://talkbank.org/errors/E305)
× error[E305]: Expected terminator not found (line 6, column 1)
╭─[input:6:1]
6 │ *CHI: hello world
· ─────────┬─────────
· ╰── here
╰────
help: Add a terminator at the end: Standard (. ? !), Interruption
(+... +/. ...), or CA intonation (⇗ ↗ → ↘ ⇘ ...)
Every error code (E305, E705, etc.) is documented with fix guidance in the
validation error reference.
Not every diagnostic is an error. Some codes are warnings: the file is valid
CHAT, but something is worth flagging (for example W110, an @Media name
that differs from the transcript’s own name only in letter case). A file
whose only diagnostics are warnings is reported as valid, and its heading
reflects that. For a Session.cha declaring @Media: session, audio:
⚠ Warnings in ./Session.cha
W110 (https://talkbank.org/errors/W110)
⚠ warning[W110]: Media filename 'session' differs from file name 'Session'
│ only in letter case. The names must match exactly: a case-insensitive
│ filesystem finds the recording either way, a case-sensitive one does not.
│ (line 6, column 1, bytes 101..124)
╭─[input:6:1]
6 │ @Media: session, audio
· ───────────┬──────────
· ╰── here
╰────
help: Make the names identical: update @Media to "@Media: Session, audio" or
rename the transcript, and give the recording the same spelling.
The summary still counts this file under Valid, and the exit code stays 0.
Validate an entire corpus
Point chatter at a directory, it walks recursively, validates in parallel,
and caches results:
chatter validate corpus/
The interactive TUI shows progress and lets you browse errors per file.
Use --format json for machine-readable output, or --quiet for CI
(exit code 1 on errors).
Convert to JSON
Get a structured representation of any CHAT file:
chatter to-json transcript.cha
The output conforms to the TalkBank CHAT JSON Schema.
Convert back with chatter from-json.
Watch for changes
Edit a file and get live validation feedback:
chatter watch transcript.cha
Every time you save, chatter re-validates and shows updated diagnostics.
What next?
- CLI Reference: all commands, flags, and output formats
- Validation Errors: every error code, with examples and fix guidance
- Batch Workflows: corpus-scale validation and analysis
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CLI Reference
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
The chatter CLI is the primary command-line surface for the TalkBank CHAT toolchain.
The following diagram shows the command dispatch structure. Each top-level command dispatches to a handler in the corresponding crate.
flowchart TD
chatter(["chatter"])
chatter --> validate["validate\n(chatter)"]
chatter --> normalize["normalize\n(chatter)"]
chatter --> tojson["to-json\n(talkbank-transform)"]
chatter --> fromjson["from-json\n(talkbank-transform)"]
chatter --> showalign["show-alignment\n(chatter)"]
chatter --> watch["watch\n(chatter)"]
chatter --> fix["fix\n(talkbank-transform splice)"]
chatter --> clean["clean\n(chatter)"]
chatter --> newfile["new-file\n(chatter)"]
chatter --> cache["cache\n(stats, clear)"]
chatter --> schema["schema\n(JSON Schema output)"]
chatter --> debug["debug\n(overlap-audit, linker-audit,\nfind, sanitize, fix-s)"]
chatter --> update["update\n(self-update, experimental)"]
chatter --> speakerid["speaker-id\n(experimental)"]
chatter --> rediarize["rediarize\n(experimental)"]
chatter --> adjudicate["adjudicate\n(experimental)"]
chatter --> sanityscan["sanity-scan\n(experimental)"]
Top-Level Commands
chatter validate PATH...
chatter normalize INPUT
chatter to-json INPUT
chatter from-json INPUT
chatter show-alignment INPUT
chatter watch PATH
chatter fix PATH... --apply
chatter clean PATH
chatter new-file
chatter cache stats
chatter cache clear --prefix PATH
chatter schema
chatter debug ...
chatter update # experimental: self-update to the latest release
chatter speaker-id INPUT # experimental
chatter rediarize INPUT --turns T # experimental
chatter adjudicate ... # experimental
chatter sanity-scan ... # experimental
Use chatter --help or chatter <command> --help for the exact live surface.
validate
Validate CHAT file(s) or directory tree(s). Accepts multiple paths.
Usage: chatter validate [OPTIONS] <PATH>...
chatter validate file.cha # single file
chatter validate file1.cha file2.cha file3.cha # multiple files
chatter validate corpus/ # directory (recursive, parallel)
chatter validate file.cha corpus/ other.cha # mix of files and directories
chatter validate corpus/ -f json # structured JSON output
chatter validate corpus/ --force # ignore cache, revalidate everything
chatter validate corpus/ --audit out.jsonl # bulk audit to JSONL file
chatter validate corpus/ --suppress xphon # suppress named error group
chatter validate corpus/ --suppress E726,E727 # suppress specific error codes
chatter validate corpus/ -j 8 # use 8 parallel workers
chatter validate corpus/ --max-errors 50 # stop after 50 errors
Options:
| Flag | Description |
|---|---|
-f, --format text|json | Output format (default: text) |
--list-checks | Print every validation check with Active/Planned status, then exit. It reads no file, so a <PATH> beside it is a usage error |
--skip-alignment | Skip dependent-tier alignment checks |
--force | Ignore cache, revalidate all files |
-j, --jobs N | Parallel workers for directory mode (default: CPU count). N must be at least 1; --jobs 0 is a usage error (exit 2) |
--quiet | Only emit errors, suppress success messages (text output only: with --format json it is a usage error) |
--max-errors N | Stop once N errors (never warnings) have been found across all files. Files already being validated finish; if any were left, the run says Stopped after reaching --max-errors N; M file(s) were not validated. (stderr in text mode, a stop record in JSON mode) and exits 1. N must be at least 1; --max-errors 0 is a usage error (exit 2) |
--roundtrip | Test serialization idempotency (developer tool) |
--parser tree-sitter|re2c | Parser backend (default: tree-sitter; re2c is opt-in for faster batch validation). Diagnostic line and column numbers are not reliable under re2c, see the note below |
--strict-linkers | Enable the opt-in cross-utterance checks of quotations (the +"/. and +". terminators and the +" linker) and of the completion linkers +, and ++; off by default. --list-checks marks each code this turns on as [Opt-in] |
--suppress xphon | Silence the Phon %x dependent-tier checks (E725-E728, E735-E746), which run by default |
--audit FILE | Stream errors to JSONL file (bulk audit mode). Its own output: a usage error with --format or --quiet. A file that cannot be written whole fails the run. An audit only reads an existing cache (no create, migrate, clear, prune or write), so --force with it is a usage error too |
--suppress CODES | Suppress error codes or groups (comma-separated). A value that names no known code or group is a usage error (exit 2) |
Every path argument is expanded the same way: a file is validated as given,
and a directory contributes the .cha files under it (links followed, each
directory once). If an argument or any entry under a directory cannot be
read, validate reports each one and exits 1 before validating anything.
If the validation cache fails during a run (for example, a locked
database), the files are validated without it and the run says so
(Warning: N cache read(s) or write(s) failed on stderr in text mode,
cache_errors in the JSON summary).
--parser re2creports unreliable diagnostic positions.The re2c lexer DOES produce a source span for every token; the parser discards it (
parser/mod.rs,lexer.map(|(tok, _span)| tok)), so the converter assigns every model node a dummy span and diagnostics that compute a position from those spans point somewhere arbitrary. The same file validated both ways:tree-sitter error[E370] ... (line 7, column 13) <- the offending tier re2c error[E370] ... (line 2, column 7) <- points at @BeginThe VERDICT is trustworthy on both backends and the two are held to structural equivalence by the parity oracle; only the reported location is not. A wrong position that looks plausible is worse than none, so treat
--parser re2cas suitable for batch pass/fail and use the default backend when you need to find the error in the file.Restoring the positions means carrying the lexer’s spans through the token slice rather than re-deriving them, which is bounded work rather than a redesign.
Suppress groups: xphon expands to the whole Phon %x
dependent-tier validation surface (%xmodsyl/%xphosyl/%xphoaln/%xphoint,
codes E725-E728 and E735-E746). These checks run by default; pass
--suppress xphon to silence the group. (--check-xphon is a deprecated
no-op, accepted so existing scripts do not break.) The
--suppress flag can mix groups and codes: --suppress xphon,E316.
Suppression does not cost you the cache. It changes what is printed, not
what is validated, so runs that differ only in --suppress share cached
results: chatter validate corpus/ followed by chatter validate corpus/ --suppress xphon reuses the first run’s work. --strict-linkers is the other
kind of flag, since it turns extra checks on, so it validates afresh.
normalize
Serialize a CHAT file into canonical formatting.
chatter normalize input.cha
chatter normalize input.cha -o normalized.cha
chatter normalize input.cha --validate
chatter normalize input.cha --validate --skip-alignment
Flags:
-o, --output <PATH>: write to a file instead of stdout.--validate: validate (including alignment by default) before writing the normalized output.--skip-alignment: with--validate, skip the dependent-tier alignment checks (the rest is still validated). It requires--validate; alone it is a usage error (exit 2).
normalize writes to stdout unless you pass -o/--output. There is no --in-place flag.
JSON Conversion
# Single file
chatter to-json input.cha # pretty-printed JSON to stdout
chatter to-json input.cha --compact # minified JSON to stdout
chatter to-json input.cha -o output.json # JSON to file
# Directory (recursive, preserves structure)
chatter to-json corpus/ --output-dir json/ # incremental by default (mtime check)
chatter to-json corpus/ --output-dir json/ --compact # minified output (saves disk)
chatter to-json corpus/ --output-dir json/ --force # full rebuild
chatter to-json corpus/ --output-dir json/ --prune # remove orphaned .json files
chatter to-json corpus/ --output-dir json/ --jobs 4 # parallel workers
# Reverse and schema
chatter from-json input.json -o output.cha
chatter schema
chatter schema --url
Single-file mode: to-json validates by default. Use --skip-validation,
--skip-alignment, or --skip-schema-validation to bypass checks.
--skip-validation already skips alignment, so passing it with
--skip-alignment is a usage error (exit 2).
Failures are reported the same way in both modes: a line
ERROR: <path>: <failure>, then the rendered diagnostics of a parse failure,
a validation failure, an incomplete validation or an internal failure, and
the command exits 1.
Directory mode: Walks recursively, converting each .cha to .json under --output-dir
with the same relative path. Incremental by default: skips files whose JSON is
already newer than the source. Use --force to rebuild all. Use --prune to remove
.json files with no matching .cha (handles renames/deletions). --prune never
follows a symbolic link in the output tree: a linked file or directory is left
alone, so nothing outside --output-dir can be deleted. Use --jobs N for
parallel conversion (defaults to number of CPUs; N must be at least 1, and
--jobs 0 is a usage error). If any directory or entry under the input cannot
be read, to-json reports each one and exits 1 before converting anything, as
fix and the debug commands do with their path arguments. (validate
instead validates the rest and reports each unreadable path as a read error
in its results, which fails the run.) An input directory with no .cha file
(an empty tree, or a mount point with nothing mounted) is refused the same
way, ERROR: no .cha files found in DIR, exit 1, before anything is
converted or pruned: --prune over it would delete every .json under
--output-dir.
Each mode takes only its own options. --output-dir, --force,
--prune and --jobs apply only to a directory input, and -o/--output
only to a file; giving one for the other kind of input is a usage error
(exit 2), as is a directory input without --output-dir.
Editing and Inspection Commands
show-alignment
Print the dependent-tier alignment for a CHAT file (debugging aid).
chatter show-alignment file.cha
chatter show-alignment file.cha -t mor # one tier type
chatter show-alignment file.cha -t gra -c # compact one-line-per-alignment output
Flags: -t/--tier <mor|gra|pho|sin> (omit to show all available
tiers); -c/--compact (one line per alignment).
watch
Watch a CHAT file or directory and re-validate on every save.
chatter watch file.cha
chatter watch corpus/
chatter watch corpus/ --skip-alignment --clear
Each changed file is validated as chatter validate --quiet would validate
it, with the default rules and the shared cache, and a file that passes says
so. Flags: --skip-alignment (faster reruns); -c/--clear (clear the
terminal between runs).
fix
Apply catalog fixes to CHAT file(s) at exact byte spans. Every file is parsed and validated, each diagnostic is resolved against a per-code fix catalog, and the resulting edits are admitted only into utterances that parsed clean (a broken region elsewhere in the file never blocks a fix, and is never itself rewritten) before being spliced in.
chatter fix file.cha # report only, writes nothing
chatter fix corpus/ --apply # write the mechanical fixes
chatter fix file.cha --apply --code E259 # opt a semantic fix into writing
Every catalog entry carries a batch-safety tier, and this command enforces it rather than trusting the caller:
- Mechanical (one right answer, no semantic judgment): written by a
bare
--apply. - Semantic (deterministic, but changes meaning enough to need a human
naming it): written only when its code is named with
--code. - Ambiguous (several valid answers, no evidence in the file picks
one): never written by this command, regardless of
--code; only reported.
A bare fix reports what it would do and writes nothing; --apply writes.
A bare fix is the dry run, so there is no --dry-run flag: passing one is
an unknown-argument usage error (exit 2). A file that cannot be read, or a
fix that --apply cannot write, is reported on stderr, is not counted as
applied, and makes fix exit 1.
Flags: --apply (write; without it, fix only reports what it would do);
--code <CODE> (repeatable;
narrows the diagnostics considered to exactly the named codes, and is how
a semantic-tier code opts into being written; an unknown code is a usage
error, exit 2); --skip-alignment. An argument list that names no
transcript, or a path that cannot be read, stops fix and the debug
tools before anything is processed (exit 1).
Missing facts are not guessed. E308, E504 and E507 do not offer participant,
role or language placeholders, even with --code. Supply the actual facts;
naming a diagnostic does not authorize inventing them.
E604 removal selects the complete typed %gra tier in its owning utterance,
including continuation lines. Other dependent tiers may intervene and remain
unchanged. This is a semantic deletion requiring --code E604; multiple target
tiers refuse selection rather than choosing one arbitrarily.
Header admission is narrow. General header edits have no enclosing utterance and are reported as skipped. W109 has a separate capability admitting only a clean, typed media-filename token; it cannot rewrite other header data or rename the transcript. Recovery-aware utterance fixes retain their own admission and changed-output verification.
clean
Show the cleaned text for each word (a debugging aid for the text-normalization pipeline).
chatter clean file.cha
chatter clean file.cha --diff-only # only words where raw differs from cleaned
chatter clean file.cha --format json
Flags: --diff-only; --format text|json.
new-file
Create a new minimal valid CHAT file from defaults.
chatter new-file
chatter new-file -o starter.cha --speaker CHI --language eng
chatter new-file -o adult.cha -s MOT -l eng -r Mother
chatter new-file -c brown -u "hello world ."
Flags:
-o, --output <PATH>: stdout if omitted-s, --speaker <CODE>: defaultCHI-l, --language <ISO 639-3>: defaulteng-r, --role <ROLE>: defaultTarget_Child-c, --corpus <CORPUS>: corpus identifier in the@IDheader (defaultcorpus)-u, --utterance <TEXT>: optional initial main-tier utterance content
Cache Commands
chatter cache stats
chatter cache stats --format json
chatter cache clear --prefix /path/to/corpus
chatter cache clear --all --dry-run
The validation cache lives under the platform cache directory and stores per-file validation results. validate --force refreshes cache state for the specified path.
cache clear needs exactly one of --all and --prefix PATH (otherwise it is
a usage error, exit 2). --prefix selects the entries for that path and
everything under it, by whole path components; a relative prefix is taken
from the current directory, as the cache stores absolute paths.
--dry-run says what the clear would do and writes nothing: how many entries
it would clear, or, for a cache an older build left, that it would first
migrate the database to this build’s schema (a migration can remove
duplicate entries, so the count is known only afterwards; the clear itself
migrates, then clears). With no cache database, both say there is none,
create nothing and exit 0. cache stats only reads, too: with no cache it
says No cache database at PATH and exits 0, and it reports an older
schema without migrating it.
validate --force clears exactly the rows of the files it validates,
however their paths were typed (a.cha, ./a.cha or an absolute path).
What the cache does and does not speed up
Files that passed are remembered; files with errors are re-checked every time. This is deliberate, and it is worth knowing because it decides how fast a re-run feels.
A file that validated cleanly is skipped entirely on the next run, as long as its contents have not changed. A file that had errors is validated again from scratch, because the cache remembers only THAT a file had errors, never what they were: the codes, the line numbers, the quoted source and the suggestions have to be produced by actually reading the file. Storing them instead would mean showing you an older release’s wording for an error that has since been improved, which is worse than waiting.
In practice this costs nothing on a corpus in good shape. A full run over the ~106,000 kept TalkBank transcripts takes about 6 seconds when cached, because only ~141 files have errors to re-check. It is noticeable in the opposite situation, part way through cleaning up a corpus where most files still fail, or just after a new release tightens a rule. Two things help there: narrow the target to the directory you are working in rather than the whole corpus, and fix files as you go, since each one that passes joins the fast path permanently.
Two other things reset the cache, both expected:
- Editing a file. The cache follows file contents, so a changed file is always re-validated, and reverting a change restores the earlier result.
- Upgrading Chatter. A new release can change what counts as valid, so every cached result from an older version is retired and the first run after an upgrade is a full one. Later runs are fast again. The previous version’s results are kept, so downgrading does not force another full run.
--suppress does not reset anything: it changes what is printed, not what is
checked, so runs differing only in --suppress share the same cached results.
debug
Developer / debugging subcommands for CHAT analysis. Not intended
for routine end-user workflows; surface and behavior may change
between releases. Run chatter debug --help for the live list. Current
subcommands include:
-
overlap-audit: analyze CA overlap markers (⌈⌉⌊⌋): pairing, temporal consistency, orphans. -
linker-audit: audit linker / special-terminator usage across a corpus (cross-utterance pairing for+<,++,+^,+",+,,+≋,+≈, plus+...,+/.,+//.,+"/.etc.). -
find: filter CHAT files by@Languagesand body content (token / substring counts) across a corpus tree; emits paths, JSONL, or CSV. -
sanitize: strip contributor lexical content while preserving structure, for protected-corpus debugging. See the Sanitize user-guide page for the full workflow. -
fix-s: normalize whole-utterance same-language@sruns into a[- lang]precode, clear the per-word@smarkers (including those on fillers and nonwords), and append any missing explicit@s:LANGcodes to@Languages. Trigger conditions and safety rules:- Every word-bearing item in the utterance, including fillers
(
&~,&-,&+), nonwords, and retraced material, must carry an explicit language marker AND every marker must resolve to the same target language. If a single filler such as&~dang3lacks a marker, the utterance is left untouched (the predicate cannot prove it is monolingual). - Bare
@sshortcuts on fillers must be cleared when the rewrite fires. A bare@sresolves relative to the surrounding tier language, so adding a[- LANG]precode without clearing the shortcut would flip the filler’s language to the precode target.fix-sclears the shortcut to keep the original meaning intact. - The pre-validation rule that catches the unrewritten pattern is
E255 (whole-utterance same-language
@srun);fix-sis the canonical repair.fix-salso appends to@Languagesany@s:LANGcode the file uses but does not declare (no rule requires the declaration: an explicit word-level@s:CODEis valid undeclared). - True no-op on already-correct files: a file is rewritten only when
a
[- lang]conversion or@Languagesrepair can be proved necessary.
- Every word-bearing item in the utterance, including fillers
(
-
join-retrace: auto-repair dangling-retrace (E370) utterances. An utterance whose last main-tier content is a retrace marker with nothing after it is joined with the next same-speaker utterance. The--scopeflag (value-enum, defaultrepetition) selects which retrace kinds qualify:--scope repetition(default, Wave 1): only[/]partial-repetition retraces qualify, and only when the successor’s leading words repeat the retraced material. This is the conservative, OBVIOUS-only repair suitable for most automated use.--scope corrections(Wave 3a, opt-in): also joins correction retraces:[//](Full),[///](Multiple), and[/-](Reformulation). Corrections replace rather than repeat the retraced material, so the leading-words prefix check is skipped; same-speaker presence alone is the gate. Use--dry-runfirst to review every proposed correction-join before writing.--scope all(Wave 3b, broadest, opt-in): joins ANY dangling retrace kind, including[/]Partial where the successor does NOT repeat the retraced material. This covers genuine child-language disfluencies: false starts, partial words, disfluent repetitions, expansions, and fillers where the transcriber correctly coded a[/]but the successor cannot repeat the abandoned material. Same-speaker presence alone is the gate. Always use--dry-runfirst when running this scope on new data.
Shared behavior for all joined pairs:
- The join produces one utterance: the first utterance’s content (keeping the trailing retrace marker) followed by the successor’s content, terminated by the successor’s terminator. Main-tier time bullets are unioned (start from the first, end from the successor).
- Dependent tiers are dropped. If either side carried
%mor,%gra, or any other dependent tier, the joined utterance drops all of them (a naive%gramerge would yield two ROOT relations, whichchatter validaterejects as E723). Such joins are reported as “needs re-morphotag” so the file can be re-run through morphotagging afterwards; the main tier alone remains valid CHAT. --dry-runreports what would be joined without modifying files.
Speaker and Review Commands (experimental)
These commands inspect, review, and relabel CHAT transcripts of the
same recording, in the tradition of CLAN’s reliability and comparison
tools (rely, trnfix). They are experimental and in active
development: flags and behavior may change, and several modes are not
yet complete. Work on copies and validate the output.
| Command | What it does |
|---|---|
speaker-id | Assign CHAT-conformant speaker codes to an anonymously-labeled file, from an explicit mapping or by text similarity against a reference transcript. |
rediarize | Re-attribute utterance speakers from an external diarizer’s timestamped turns (JSON), keeping the words: repairs transcripts whose ASR under-counted or mixed speakers. |
adjudicate | Resolve pending decisions (currently speaker-id) interactively or from a scripted decision file, writing results to an override file. |
sanity-scan | Post-merge QA: flag sessions whose automatic decisions look suspicious by an out-of-band heuristic, for operator review via adjudicate. |
Full guides: Speaker ID, Rediarize, and
Review Tools. The holistic mode of speaker-id can call
an LLM provider when configured; deterministic modes need no network access.
There are no merge, pipeline or batch commands (why),
and no drop-in CLI for fuzzy event matching.
Exit Codes
| Code | Meaning |
|---|---|
0 | Success: all files valid, or the command completed without errors |
1 | Failure: validation errors found, parse errors, or the command failed |
2 | Usage error: invalid arguments or missing required options (from clap) |
chatter validate exits 0 only when the run covered every file it was
given and none was invalid, unreadable or a tool failure (warnings do not
fail a run). It exits 1 for any such file, for a run that stopped or lost
files, and for an input that named no transcript, on every surface: text,
JSON, an audit file and the TUI alike (a TUI closed before its run ended
exits 1). This makes it safe to use in scripts and CI pipelines:
chatter validate corpus/ --quiet --tui-mode disable || echo "Validation failed"
Use --quiet to print only problems while still relying on the exit code.
Use --format json for machine-readable structured output (JSON objects go
to stdout; the exit code is the same).
A consumer that stops reading early (chatter ... | head) closes standard
output; every command then stops and exits 1, whatever it was printing,
since the output it promised is incomplete.
The output flags are decided together, once, into one surface: plain text,
quiet text, JSON, an audit file, or the interactive TUI. The TUI is chosen
automatically only for plain text with stdout a terminal (--tui-mode auto,
the default); --format json, --quiet and --audit never open it, and
--tui-mode force beside any of them is a usage error (exit 2), as are
--audit with --format or --quiet, and --format json with --quiet.
--max-errors applies to every surface, the TUI included.
Output Contracts
- Text output is intended for humans.
- JSON output is intended for automation and downstream tools.
- Error codes and the JSON Schema are documented public contracts; see the Integrating section of this book.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Validation Errors
Status: Current Last modified: 2026-10-07 (commit 5e895791)
The CHAT validator produces diagnostics at two severity levels: errors (must fix) and warnings (should fix). Each diagnostic has an error code that maps back to a documented spec and validator rule.
chatter validate is the binding judgment on whether a byte sequence is valid CHAT. When it reports an error, the file is invalid CHAT: clean the data rather than working around the check.
But that is a conclusion, not a reflex, and it cuts both ways. chatter is
under active development and its current behaviour is not sacred. If a
diagnostic does not make sense against your file, the fault may be ours, and we
would rather hear about it than have you edit a transcript to silence it. Tells
that a diagnostic is a chatter bug: the message names a construct your line does
not contain, the code is a generic “unparsable content” rather than a specific
rule, or no documented rule justifies it. Report it with the smallest file that
reproduces it. Editing data to quiet a wrong check destroys the evidence and
leaves the bug in place for the next person. A warning flags a questionable but parseable construct you should review. Where chatter and an older tool such as CLAN’s check disagree on whether a file is valid, chatter validate is authoritative (see CHECK Parity Audit for how the two are reconciled).
Reading Error Output
The validator emits rich diagnostics that include the error code, a source-pointed snippet, and a suggested fix:
× error[E304]: Missing speaker in main tier (line 15, column 3)
15 │ * hello world .
· ╰── here
╰────
help: Add a speaker code between * and : (e.g., *CHI:)
Each diagnostic contains:
- File path and location (line:column)
- Severity:
errororwarning - Error code:
Eprefix for errors,Wprefix for warnings, with a URL pointing at the per-code documentation page - Message: human-readable description
- Suggestion: actionable fix guidance where available
Error Code Ranges
| Range | Category | Examples |
|---|---|---|
| E1xx | UTF-8 and encoding | E101: Invalid line format |
| E2xx | Word-level content | E202: Missing form type after @, E203: Invalid form type marker, E207: Unknown annotation |
| E3xx | Main tier (speakers, terminators, content) | E301: Empty/missing main tier, E304: Missing speaker, E305: Missing terminator, E306: Empty utterance, E307: Invalid speaker, E308: Undeclared speaker |
| E4xx | Dependent tier structure | E401: Duplicate dependent tier |
| E5xx | Headers | E501: Duplicate header, E504: Missing @Participants, E505: Invalid @ID format |
| E6xx | Dependent tier validation | E601: Invalid dependent tier, E604: %gra without %mor |
| E7xx | Alignment, Phon tiers, structure | E705: Main/%mor count mismatch, E721: %gra index error, E747: Blank line, E748: Leading zero in bullet time, E749: Comma glued to next word, E750: Space inside angle group, E751: Pause glued to word, E752: Timing bullets without @Media, E753: Word only repetition segments, E755: Undeclared utterance language, E756: Empty dependent tier, E757: Bracketed code glued to following word, E758: Leading space on tier (non-CA), E759: Annotation at utterance start, E760: %mor item with empty POS, E761: %gra relation head not a UD relation, E762: Prefix marker # standalone or word-initial, E763: Prefix marker # in a language that does not use it, E764: Prefixed form glued to the preceding word, E765: Separator glued to following content, E766: Linker not utterance-initial, E767: Whitespace before the @Media comma, E768: @Media filename not representable |
| W1xx-W6xx | Warnings | W108: Speaker not found in @Participants (non-fatal contexts) |
Common Errors and Fixes
E256: Curly single quote used as a word character
A curly single quotation mark (U+2018 or U+2019), commonly introduced by
autocorrect or speech-to-text, is not a legal CHAT word character. CHAT words
use the ASCII apostrophe (U+0027, the plain '). For example, a contraction
typed as don + U+2019 + t is rejected; write don't with the ASCII
apostrophe instead. chatter flags the curly form wherever it appears in word
content and points the diagnostic at the exact character. This mirrors CLAN
CHECK errors 138 and 139.
E243: Private-use characters or noncharacters in a word
Chatter rejects Unicode private-use scalars and noncharacters in lexical words, including supplementary planes. This is a transcription-interchange policy, not a claim that those scalars are invalid UTF-8. The diagnostic identifies the scalar, category and containing word. Use the intended standard transcription character; do not silently delete text or assume changing the encoding repairs it.
Ordinary high-BMP and supplementary characters are not rejected by this rule, and CLAN’s private-use markup has no exemption. Existing control-character and CHAT punctuation checks still apply. U+FFFD is not a noncharacter: this rule does not reject it, but its presence can indicate earlier loss during decoding. See the Unicode FAQ for the categories and the CHECK assessment for the scope of the comparison.
E304: Missing speaker code
A main tier line must have a speaker code after the *:
*CHI: hello world .
An empty speaker code (*: hello .) triggers E304.
E308: Undeclared speaker
Every *SPEAKER: code must be listed in @Participants. Add the missing speaker to the header:
@Participants: CHI Target_Child, MOT Mother
E370: Retrace marker with nothing to retrace
A retrace or repetition marker ([/], [//], [///]) must be followed by the
repeated or corrected material; per the CHAT manual the marker always refers to
the text that follows it. A marker followed only by a terminator has nothing to
retrace:
*CHI: <the> [/] . ← invalid: [/] is not followed by repeated material
*CHI: <the> [/] the cat . ← valid: the repeated material follows the marker
This mirrors CLAN CHECK error 119 (and the related retrace checks 151 and 159).
E505: Invalid @ID format
Check that pipe-separated fields are correct and the speaker code matches @Participants:
@ID: eng|corpus|CHI|2;6.||||Target_Child|||
E705: Main/%mor alignment mismatch
The number of %mor items must match the number of alignable words on the main tier. Retraces, pauses, and events are not counted. The validator shows a columnar diff:
Main tier %mor tier
────────────── ──────────────
I pro|I
want v|want
to inf|to
go v|go
home, ⊖
E714 / E715 and E733 / E734: phonology count mismatch
E714 and E715 report too few or too many %pho tokens. E733 and E734 report
the corresponding %mod mismatch. The tier names use distinct codes because
actual and model phonology are different evidence layers.
%wor does not use any of these codes. It is a timing sidecar, and a count
mismatch does not make legacy CHAT invalid. Timing consumers handle the state
explicitly:
Missing: no%wortier;Drifted: current main-tier and%worslot counts differ;CountMatched: counts match, but no timing is exposed yet;Uncorroborated: canonical display tokens differ;Corroborated: positional slots are available, with lexical identity from the main tier and timing from%worword bullets.
The current FilteredLexicalV1 policy includes regular words, fillers,
retraced regular words, and the original spoken word of a replacement. It
excludes fragments, nonwords, xxx/yyy/www, omissions, pauses, and actions.
Changing that policy requires a new named policy and evaluation; it must not
silently reinterpret existing %wor data.
And this is valid too:
*EXP: what's is dis [: this] ?
%wor: what's •37050_37471• is •37491_37631• dis •37631_38131• ?
E721: %gra sequential index error
%gra entries must have sequential 1-based indices: 1|...|... 2|...|... 3|...|...
E748: Leading zero in bullet timestamp
A media bullet time component is written with a leading zero before
another digit, for example \u{15}012_200\u{15}. Bullet times are
plain millisecond integers; write 12, not 012. A bare 0 (as in
0_200) is legal. This mirrors CLAN CHECK error 90 (“Illegal time
representation inside a bullet.”). The bullet’s numeric value still
parses, so downstream tooling sees the intended times; the diagnostic
alone makes the file invalid.
E749: Comma glued to the following word
A comma on a speaker tier must be followed by a space or end-of-line:
write hey , you, not hey ,you. Mirrors CLAN CHECK error 92. The
check looks at the word immediately after the comma in document order
(including inside <...> groups); constructs that place their own
character after the comma (group and overlap marks, CA symbols) are
not flagged.
E750: Space inside angle-bracket group delimiters
Group delimiters hug their content: write <dog> [/], never < dog>
or <dog >. Mirrors CLAN CHECK error 160. Each offending space gets
its own diagnostic; the group still parses, so downstream tooling sees
the intended structure. chatter fix --apply --code E750 removes only the
offending delimiter-adjacent space. This mechanical catalog repair carries a
distinct recovery-safe edit state, so it may repair the parser recovery that
tainted its own utterance while ordinary edits remain barred from recovered
content; the standard post-splice reparse still has to prove the result.
E751: Pause glued to the preceding word
A pause marker must be space-delimited from the word before it: write
hello (.) there, not hello(.) there. Mirrors CLAN CHECK error 57.
E531, W109, W110: transcript and @Media names
The @Media name must equal the transcript’s own file name, spelled exactly.
E531 (error) means the two are different names. W109 (warning) means they are
the same name but not both in Unicode NFC, the standard composed spelling of
accented letters. W110 (warning) means they differ only in letter case. W109
and W110 do not stop validation, but a name that differs in either way is
found on macOS and not on Linux. See
File Names and @Media for the rule, why it
matters and how to fix each.
E752: Timing bullets without an @Media header
The transcript carries timing evidence (an utterance-final bullet, a
bullet inside an utterance, or %wor word timing) but no @Media header
declares the recording those timestamps index. Add an @Media header
naming the media file (or remove the timing bullets if the transcript is
genuinely unlinked). Completes the
media-consistency family: E544 covers declared linkage without timing,
E552 covers a declared unlinked contradicted by timing. Mirrors CLAN
CHECK error 112.
E753: Word consisting only of a repetition segment
A word whose entire material sits inside segment-repetition delimiters
(↫hi↫ with nothing outside the arrows) marks the repetition of a word
that is not there; attach the repeated segment to its host word
(↫p↫parents) or transcribe a stand-alone fragment as a filler or
nonword form. Filler and other word-category prefixes (&-, &~, 0)
count as material outside the arrows. Adopted from GUI CLAN CHECK error
151 as a chatter rule.
E519 at word level: language codes must be real everywhere
The ISO 639-3 registry check that guards @Languages and @ID also
applies to explicit word-level switch codes (word@s:CODE, including
+/& multi-code forms) and to @L1 of values: the code needs no
declaration, but it must name a real language. Utterance-level [- CODE] precodes are covered
by E755 plus the header check.
E755: Utterance language not declared in @Languages
A [- CODE] precode marks a whole utterance as being in another
language, which is substantial presence: declare that language in
@Languages. Deliberate contrast: a word-level @s:CODE insertion
needs NO declaration (ok@s:eng in a Cantonese transcript is valid
as-is), because @Languages lists the transcript’s substantial
languages, not every language that appears. Mirrors CLAN CHECK error
152.
E756: Empty dependent tier
A dependent tier with empty or whitespace-only content declares an
annotation that is not there; add the content or remove the line.
Whitespace-only counts as nothing on every free-text tier, %com and
%add included; CLAN
CHECK 31 rejects the same lines.
This covers every tier whose body is free text, which is every
dependent tier except the structured ones (%mor, %gra, %pho,
%mod, %sin, %wor). An empty structured tier fails earlier and
more specifically, because its body is not free text and there is no
“you declared nothing” to report.
The rule covers standard tiers such as %eng: as well as user-defined
%x tiers, and is a hard error.
E757: Bracketed code glued to the following word
A bracketed code’s closing ] must be space-delimited from what
follows: write hello [/] x, not hello [/]x. The parse is
unambiguous either way, which is exactly why this is a style rule: the
corpus stays canonically spaced. Mirrors CLAN CHECK error 19.
E758: Leading space before tier content in a non-CA file
A space between the tier’s tab delimiter and the first content item
(*CHI:<tab><space>dog .) is invalid unless the file declares
@Options: CA; CA transcripts legitimately column-align content with
spaces after the tab. Mirrors CLAN CHECK error 123.
E759: Annotation at utterance start
Postfix annotations (retraces [/] [//], overlap markers [<]
[>], replacements [: text], the quotation code ["]) scope over
the material BEFORE them; an utterance whose content begins with one
has nothing for the code to attach to. Mirrors CLAN CHECK error 52.
E760: %mor item with an empty part-of-speech field
A %mor item beginning with the | separator (|we) declares no
part of speech; every item is pos|stem with a non-empty POS. The
modern reading of CLAN CHECK error 11 (the depfile mechanism is
legacy; the non-empty-symbol invariant is real).
E761: %gra relation head is not a Universal Dependencies relation
A %gra label is HEAD or HEAD-SUBTYPE. UD fixes the head set at 37
universal relations and defines subtypes as language-specific and
open-ended, so only the head is checked; a subtype such as NMOD-POSS
or ACL-RELCL passes untouched. CLAN CHECK does not validate relation
labels, so without this rule a typo like PUNCTT for PUNCT would ride
silently into every analysis that reads the dependency graph. Common causes: truncation (IOB for IOBJ), typos, and the
retired TalkBank labels (SUBJ, JCT, POBJ, INCROOT), none of
which occurs in the corpora.
E762: prefix marker # stands alone or opens a word
The prefix marker separates a bound prefix from its stem and attaches
to the END of the prefix, which is a word of its own (Hebrew
ha# kelev, “the dog”). So a word that is nothing but #, or one that
opens with it (#dog), cannot be that construct in any language. This
covers CLAN CHECK 71 and the #-undeclared facet of CHECK 11.
E763: prefix marker # in a language that does not use it
Languages that write the marker are heb and ara; anywhere else it
is a stray character, usually a typo or a conversion artifact. The gate
reads the WORD’s resolved language, not the file’s @Languages header,
exactly as the digits rule (E220) does, so a code-switched word marked
@s:heb inside an English file is accepted. Word-internal markers
(mi#ha#shuk) stay legal wherever the language allows the marker at
all.
E765: separator glued to the following content
A free-standing : or ;, or a pause, must have a space after it:
:and, ;;, (.)dog are invalid. The preceding side is untouched:
word↘ and dog, are documented CHAT convention, and dog: is not two
items at all (the colon fuses as lengthening).
Every CA mark is out of scope, on corpus evidence. Implementing the
whole separator class flagged 270 instances in a 2% corpus sample (about
13,500 corpus-wide), all legitimate notation: ≡ is latching and is
written glued on both sides (y≡I≡) because that is what it encodes, and
the intonation arrows attach to the material they mark, including
directly before an overlap close (⌊I don't know⇗⌋). Whether any CA mark
should forbid trailing glue is unresolved.
E766: linker not utterance-initial
Linkers (+", ++, +<, +^, +,, +≈, +≋) tie an utterance to the
PREVIOUS one, so they may only open the utterance. One placed after content
(the dog ran +" away .) is meaningless and is named here, at the exact
token, instead of surfacing as generic unparsable content (E316).
One deliberate carve-out: a ++ glued to words on both sides (un++do) is
not a linker but a word run with an empty compound part, and keeps its E233
diagnosis.
E767: whitespace before the @Media comma
In @Media the comma separates the filename from the media type, so the
filename ends where the comma begins and a space between them belongs to
neither. Real CLAN rejects it too (CHECK 148), and an unambiguous style
violation that CLAN rejects is an error here.
The construct is unambiguous, so the grammar deliberately PARSES it rather
than failing, which is the only way to name the rule and point at the exact
space. Were the line to fail to match, the whole header would fall back to
Unknown and report E525 about a header chatter had recognised perfectly
well, alongside E330 “Missing media_type node” on a line visibly ending in
, audio. Deleting the one space validates clean.
E768: @Media filename cannot be written and read back
The filename is delimited by the comma that introduces the media type, so a few strings cannot survive a round trip through the header: an unquoted comma, surrounding whitespace, a line break, a stray double quote, or an empty name. A quoted remote URL may contain a comma.
You will not see this one from a transcript. Both parsers end the filename at
the comma, so no .cha file can express a violating value; the rule guards a
ChatFile that arrived as JSON, where deserialization is deliberately lenient
and validation is what reports the violation.
E757 scope: every bracketed code, not only retraces
hello [/]there, hello [!]there and bobo [= toy]there are all caught: the
rule applies to any item that ends in a bracketed code.
An annotated word’s wrapper carries a real span (as its annotated event, action, and group siblings do), so the glue is visible to a span-adjacency check, and a diagnostic reported on an annotated word points at the word.
E764: prefixed form glued to the preceding word
dog&-um parses as TWO words, because & cannot continue a word. So a
single missing space silently manufactures a word boundary the
transcriber never wrote, and nothing else reports it. Applies
to the three & prefixes (&- filler, &~ nonword, &+ fragment).
Glued omission (dog0is) is a different shape: 0 is ordinary word
text, so it yields one malformed word and is already rejected by E220.
E243 and the pipe character
| is the %mor tier’s delimiter and has no meaning in main-tier word
text; a bare or embedded pipe in a word reports E243
(IllegalCharactersInWord). Covers the grounded shape of CLAN CHECK
error 48.
Generated Error Documentation
The source of truth for error-code details is spec/errors/. Maintainers can
The browsable error catalog under docs/errors/ is generated from those specs
and committed, so it is regenerated with every other artifact:
just spec-gen
That generated reference includes the error description, example inputs, suggested fixes, and the layer that catches the diagnostic.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Chatter Desktop
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
Chatter Desktop is a native graphical validation app for CHAT files, released
alongside the chatter CLI. Prefer the chatter CLI for scripted or batch
validation; use the desktop app when you want a standalone graphical validation
experience without a terminal.
When to use Chatter Desktop
Chatter Desktop (apps/chatter-desktop/) is the right tool when you want to:
- Validate CHAT files through a graphical interface, no terminal required
- Drag and drop a file or folder and read errors with source snippets
- Work on the desktop without setting up a terminal workflow
Related surfaces:
- Validate CHAT from the command line: use
chatter validate
This page documents the desktop surface:
- Chatter Desktop (
apps/chatter-desktop/), the CHAT validation GUI
Current status
- Release contract: released alongside the CLI in the public chatter release
- Distribution: ships in the coordinated chatter release alongside the CLI; also buildable from source (below)
- Platforms: macOS, Windows, and Linux
Staying up to date
Chatter Desktop keeps itself current. When you launch it, it quietly checks for a newer release; if one is available it asks whether to update, and on your confirmation it downloads, installs, and restarts into the new version. The app also checks every six hours while running. Use Check for Updates… in the application menu to check immediately; a manual check reports when you are current or when the check fails. Overlapping checks share one operation, so repeated menu clicks do not start duplicate prompts or installations.
If a background check cannot reach the network, the app keeps working on the installed version. A failed check is not evidence that the installed version is current. Retry the menu command later or download an installer from the release page. Export any results you want to retain before accepting an update: installation relaunches the app, and the current results are not a saved session. About Chatter shows the installed version and links to the project.
Getting Started
Install the released application
Open the Chatter release page and choose the desktop installer for your platform, rather than a CLI archive:
| Platform | Desktop download | Installation |
|---|---|---|
| macOS, Apple silicon | Chatter-macos-apple-silicon.dmg | Open the disk image and copy Chatter to Applications. |
| macOS, Intel | Chatter-macos-intel.dmg | Open the disk image and copy Chatter to Applications. |
| Windows | Chatter-windows-setup.exe | Run the installer. |
| Linux | Chatter-linux-x86_64.deb or .AppImage | Use the Debian package manager, or make the AppImage executable and launch it. |
The macOS application and disk image are signed and notarized by the release pipeline. Windows installers currently lack Authenticode signing and may show a SmartScreen warning. Verify that you downloaded the intended project release; do not disable operating-system security globally to install it. Updater signatures are separate from operating-system signing.
No Rust toolchain, Node.js, terminal, or separate CLI installation is needed to use a released desktop app. CLAN is optional and needed only for Open in CLAN.
A first validation session
- Choose one
.chafile or a folder. Folder validation includes subfolders. - Leave Tree-sitter selected. Keep optional roundtrip and strict-linker checks off unless you need them for this session.
- Wait for completion; a blank problem list during discovery or processing does not certify the files.
- Select a file with a diagnostic or processing failure, read its message, and use Copy, Reveal, or Open in CLAN as appropriate.
- Edit the original transcript in your editor, save it, and Re-validate. Export results if you need a record before starting another target.
Validation does not edit or rename your transcripts. The desktop app is not a CHAT editor or an automatic repair interface.
Build from source
cd apps/chatter-desktop
npm ci
cargo tauri dev # launches the app with hot reload
cargo tauri build # produces a distributable app bundle
Use the repository’s pinned Rust toolchain and the Node version used by its desktop workflow. Linux also needs Tauri’s native system dependencies. A distributable updater-enabled build needs signing configuration; the two build commands above are not a substitute for the coordinated release pipeline. See Desktop App Testing and CI and Release.
Using the App
Opening files
Chatter validates one target at a time: a single .cha file or one folder.
Three ways to start validating:
- Choose File: opens a file picker filtered to
.chafiles - Choose Folder: opens a folder picker; validates all
.chafiles recursively - Drag and drop: drag one
.chafile or one folder onto the app window
When idle, if you’ve previously validated a target, the drop zone shows “Last: corpus/reference/, Re-validate?” as a clickable shortcut.
Reading results
The main window has three areas:
┌──────────────────────────────────────────────────────────────┐
│ [Choose File] [Choose Folder] or drag here [System|Light|Dark] │
├──────────────────┬───────────────────────────────────────────┤
│ 3 FILES WITH │ Filter by code… [All|Errors|Warnings] │
│ ERRORS / 120 │ │
│ │ ▾ [E302] Missing @End header │
│ 📁 corpus/ │ ┌───────────────────────┐ │
│ ✗ file1 (3) │ │ 41 │ *CHI: hello . │ │
│ ✗ file3 (1) │ │ 42 │ │ │
│ │ │ │ ^ │ │
│ │ └───────────────────────┘ │
│ │ 💡 Add @End on the last line │
│ │ [Copy] [Open in CLAN] │
├──────────────────┴───────────────────────────────────────────┤
│ Progress: 45/120 │ 4 errors │ ~2m 30s remaining │ [Cancel] │
└──────────────────────────────────────────────────────────────┘
-
File tree (left), collapsible directory tree showing files with diagnostics (including warnings) or read, parse and roundtrip failures. Valid files without diagnostics are hidden to reduce clutter. A header shows “N files with errors / M total”. Files are sorted alphabetically.
-
Error panel (right), for the selected file, shows each error with its code in
[E001]format, severity color, message, source snippet with caret underlines, and multi-span labels for complex errors (e.g., alignment mismatches across tiers). CHAT-specific formatting is handled: tabs expanded to 8-column boundaries,\x15bullets rendered as•, underline markers shown as styled underlined text. Suggestions prefixed with 💡. -
Status bar (bottom), streaming progress during validation, ETA after 5+ files, total error count, and action buttons.
Filtering errors
A compact filter bar appears above the error cards when a file has diagnostics:
- Code filter: type “E7” to show only alignment errors, “W” for warnings, etc.
- Severity toggle: switch between All / Errors / Warnings
The file header updates to show filtered vs. total count (e.g., “3 errors (7 total)”).
Collapsible error cards
Each error card has a clickable header that toggles between expanded and collapsed view. Collapsed cards show only the error code and first line of the message. When a file has 5 or more errors, an Expand All / Collapse All button appears.
Validation settings
A ⚙ Settings popover next to the file picker exposes the same knobs the CLI’s flags do, since both surfaces build the same underlying validation config:
| Setting | Equivalent CLI flag | Default |
|---|---|---|
| Roundtrip check | --roundtrip | Off |
| Parser | --parser tree-sitter|re2c | Tree-sitter |
| Strict cross-utterance linkers | --strict-linkers | Off |
| Parallel jobs (at least 1) | --jobs N | All CPUs |
Settings are disabled while a validation run is in progress and apply to the next run (including Re-validate).
Parallel jobs takes a whole number of at least 1, or empty for all CPUs.
Anything else (0, 1.5, a negative number) is not taken: the field says so,
and that the next run uses the last valid value. The backend refuses 0 as
well.
When a run finishes, its summary also says if the validation cache failed (for example, a locked database), or would not open at all, with the reason: those files were validated without the cache, so the results stand, but the next run will not be faster for them.
Re2c is experimental and incomplete. It is not a second validity authority; use Tree-sitter for ordinary work and report disagreements with a minimal CHAT example. Strict-linker checks enforce additional quotation/completion conventions that are not enabled for ordinary validation. Roundtrip checking compares the parsed model with its serialized and reparsed form; it is not an audio check or a guarantee that every byte keeps its original formatting.
The app shares the CLI’s validation cache and rule-aware engine. Changed files are revalidated; a cached validation result is not itself proof that an optional roundtrip check ran. Selected settings apply to the next run, not the run already in progress. Settings currently reset to their defaults when the app restarts.
Failures, warnings and incomplete results
Read and parse failures may have no CHAT error card because the validator could not obtain a usable transcript. They still appear in the file tree, and the detail panel shows the reason. A roundtrip failure is likewise a failed check, not “No errors.” Report unexpected roundtrip failures rather than editing the data merely to silence an internal mismatch.
“All files valid” requires a completed, non-cancelled, nonempty run with every file accounted for as valid and no visible diagnostics or roundtrip failures. Cancelling, losing files during a run, or failing to read a file cannot earn that summary. An empty folder means No CHAT files found, not a successful validation population. The title, notification and status bar share the same completion summary.
Warning-only files remain visible even though a warning is not a hard error.
For example, W109 identifies a nonstandard Unicode spelling in the @Media
name, the stored transcript filename, or both. Typing a different Unicode form
of the same filesystem path must not change that diagnosis. The optional CLI
repair chatter fix --code W109 --apply <file> normalizes only the media-name
token; it never renames the file. A file-only warning can remain afterward.
See Headers for normalization details.
Dark mode
Chatter follows your system appearance by default. A System / Light / Dark toggle in the drop zone area lets you override. Your preference is remembered across sessions.
The dark palette uses muted Apple-style colors, readable miette error highlighting on dark backgrounds.
Clickable file paths
Click the file name in the error panel heading to reveal the file in Finder (macOS), Explorer (Windows), or the default file manager (Linux).
Copy errors
Each error card has a Copy button that copies the full miette-rendered error text (plain text, not HTML) to your clipboard for pasting into issue reports or messages.
Actions
| Action | Where | What it does |
|---|---|---|
| Re-validate | Status bar / last-target hint | Re-run validation on the same target (picks up edits) |
| Cancel | Status bar (during validation) | Stop the current run |
| Export | Status bar | Save results as JSON or plain text via a save dialog |
| Open in CLAN | Per-error button | Opens the file at the error location in the CLAN editor |
| Copy | Per-error button | Copies the plain-text error to clipboard |
| Reveal in file manager | File name heading | Opens the file’s parent directory |
“Open in CLAN” only appears when the CLAN application is detected on your
system (macOS and Windows only). It adjusts line numbers to account for headers
that CLAN hides (@UTF8, @PID, @Font, @ColorWords, @Window).
Exporting results
After a run ends, Export opens a save dialog. JSON preserves per-file diagnostics and status; plain text includes each file’s status and its rendered diagnostics. Read, parse and roundtrip failure reasons are included even when there is no CHAT diagnostic card. Copying one card exports only that diagnostic, not the outcome of the whole file or folder.
A cancelled run contains only the results obtained before cancellation. An export is a record of those results, not proof that every requested file was checked. Retain the run’s completion/cancellation context with the report; the current per-file export does not include a full session manifest with settings, application version and run coverage. Revalidate to obtain a complete population before making a whole-folder validity claim.
Keyboard shortcuts
| Shortcut | Action |
|---|---|
| Ctrl+R / Cmd+R | Re-validate |
| Escape | Cancel running validation |
All other navigation is mouse-driven (click files, scroll errors).
Window title
The window title updates to reflect the current state:
- Idle: “Chatter”
- Starting: “Chatter, Starting…”
- Discovering: “Chatter, Discovering files…”
- Running: “Chatter, Validating (45/120)”
- Finished: a diagnostic/failure summary, or “Chatter, All 74 files valid” (with “; 3 warnings” when some files have warnings only)
- Cancelled: “Chatter, Cancelled (3 files not checked)”
- Empty: “Chatter, No CHAT files found”
- Incomplete: “Chatter, Incomplete (2 files not checked)”
- Stopped: “Chatter, Run stopped unexpectedly”
A cancelled run says how many files it never reached, and its status bar leads with that number. It is its own state, never a finished run with a flag, so it cannot show “All N files valid”.
The last two are failures, and they never claim anything about your whole folder. Incomplete means the validator finished but some files were never opened, so the counts it shows describe only the rest; you will see how many were missed, and re-validating is the right response. Stopped means the run died without producing results at all. Neither one can show “All N files valid”, because that sentence is a claim about every file and neither run examined every file.
“Starting” and “Discovering” are different states, and the difference is worth knowing if you ever need to report a problem. Starting means the app has asked the validator to begin and has not heard back; nothing has been scanned yet, so a run stuck there is a fault in start-up rather than anything about your files. Discovering means the validator is walking the folder, which legitimately takes time on a large one. If the app sits on “Starting” for more than a few seconds it says so in the status bar, and that message is worth quoting in a bug report.
ETA
After 5 or more files have been processed, the status bar shows an estimated time remaining (e.g., “~2m 30s remaining”). The estimate updates every second.
Notifications
When validation finishes while the app is not focused, a system notification shows the summary (“Validation complete, 14 errors in 3 files”).
First launch
On first launch, an onboarding overlay explains the four main interactions: drag files, error panel, keyboard shortcuts, and export. Dismiss with “Got it”, it won’t appear again.
CLI Bundling
Install the CLI separately from the same release when you need scripted validation or repairs. The current desktop bundle configuration does not ship a CLI resource or an Install CLI Command menu item. A backend installation command is not evidence that a released bundle contains a CLI. Bundling remains future work, not a prerequisite for using desktop validation.
Troubleshooting and reporting a problem
| Symptom | What to check |
|---|---|
| Stuck on Starting | Quote the startup message; this is before file discovery. |
| Discovering takes time | Large or network folders can be slow. Try one local file to isolate the issue. |
| Read error | Confirm the file still exists and the app can read it; check volume availability and permissions. |
| No CHAT files found | Select the intended folder and check that transcripts have .cha extensions. |
| Open in CLAN unavailable | Confirm a supported CLAN installation; validation itself does not need CLAN. |
| Update failed | Keep using the installed app, retry later, or use the official installer. |
| Roundtrip failed | Retain the transcript and reported reason; this can indicate a parser/serializer defect. |
Include the installed version from About Chatter, operating system, selected parser/settings, whether the target was one file or a folder, and copied diagnostics. Share a minimal sanitized example when possible. Review exported results before sharing: they can contain file paths and transcript snippets. Do not put private participant data in a public issue.
Current limitations
- One validation target per run; no in-app editing or automatic fix command.
- No persistent validation-result session; export results before closing or updating.
- Re2c remains experimental; Tree-sitter is the default.
- Desktop navigation is primarily mouse-driven; it is not the terminal UI’s key map.
- CLAN integration is platform-dependent. Native end-to-end automation is available on Linux/Windows; macOS requires a real-app smoke review as well as seam tests.
Architecture
The desktop app lives in apps/chatter-desktop/:
apps/chatter-desktop/
src-tauri/ Rust backend (Tauri v2)
src/
main.rs Bin entry, calls chatter_desktop_lib::run()
lib.rs Tauri app setup (Builder + module wiring)
protocol.rs Shared command/event names + request types
commands.rs validate, cancel, open_in_clan, export, reveal, install_cli
events.rs ValidationEvent → frontend event bridge
validation.rs Desktop validation orchestration for one target
src/ React + TypeScript frontend
components/ DropZone, FileTree, ErrorPanel, ProgressBar, OnboardingOverlay
hooks/ useValidation, validationState, useTheme
protocol/ Command/event names + TypeScript transport mirrors
runtime/ Tauri transport + capability-focused runtime seam
The Rust backend calls validate_directory_streaming() and
validate_files_streaming() from talkbank-transform directly (folder vs.
single-file targets respectively), the same streaming validation pipeline and
on-disk cache used by the CLI and TUI. Events flow over crossbeam channels to
the Rust side, then are serialized to JSON and emitted to the frontend via
Tauri’s event bridge.
Cancellation uses ArcSwapOption for lock-free atomic swap of the cancel
sender, no mutex.
The frontend keeps Tauri-specific code confined to src/runtime/tauriTransport.ts.
React components and hooks consume narrower capabilities (validationRunner,
validationTarget, clan, exports) instead of reaching for one broad
desktop service object.
Comparison with TUI
| Feature | TUI (chatter validate) | Desktop app |
|---|---|---|
| File selection | CLI arguments | Drag-and-drop, file picker |
| Navigation | Keyboard (Tab, arrows) | Mouse click |
| Error display | Two-pane terminal UI | Scrollable panels with source snippets |
| Error filtering | , | Code filter + severity toggle |
| Copy error | , | Copy button per error |
| Open in CLAN | c key | Button per error |
| Export | --format json or --audit FILE | Save dialog (JSON or text) |
| Streaming progress | Progress bar | Progress bar + ETA |
| Dark mode | Terminal theme | System/Light/Dark toggle |
| Caching | Same engine | Same engine |
| Who it’s for | Power users, CI | Researchers, linguists |
Both use the identical validation engine and produce the same error codes.
When to Use Which Tool
The TalkBank toolchain offers validation through three interfaces. Each serves a different workflow:
| Tool | Audience | Use when |
|---|---|---|
| Chatter Desktop | Researchers, linguists | You want a graphical, drag-and-drop CHAT validation app without using a terminal. |
chatter validate (TUI) | Power users | You’re comfortable in a terminal and want keyboard-driven navigation. |
chatter validate (CLI) | CI, scripts | You need machine-readable output (--format json) or batch audits (--audit). |
Chatter Desktop focuses on validation only.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CLAN Line Numbering
Status: Current Last modified: 2026-05-29 17:31 EDT
When you click “Open in CLAN” in the desktop app or press Enter in the TUI, chatter sends the error location to the CLAN editor. CLAN opens the file and places the cursor at the error. This usually works seamlessly, but there is one caveat: CLAN and chatter count lines differently.
Hidden Headers
CLAN hides five header types from its editor display:
| Header | Purpose |
|---|---|
@UTF8 | Character encoding declaration |
@PID | Persistent identifier |
@Font | Display font settings |
@ColorWords | Color coding rules |
@Window | Window position/size |
These headers are present in the .cha file but invisible in CLAN’s editor.
CLAN’s line numbers skip them entirely. A file that starts with @UTF8 on
line 1 will show @Begin as “line 1” in CLAN’s display, even though it’s
actually line 2 in the file.
What Chatter Does
Chatter automatically adjusts line numbers before sending to CLAN:
- Compute the error’s line number in the source file
- Count how many hidden headers appear before that line
- Subtract the hidden count to get CLAN’s line number
- Send the adjusted line number to CLAN
This happens transparently, you don’t need to do anything.
Edge Case: Errors on Hidden Lines
If an error is on a hidden header itself (e.g., a malformed @UTF8 line),
CLAN cannot navigate to it because CLAN doesn’t display that line. In this
case, “Open in CLAN” will show an error message explaining why.
For Developers
The shared resolution logic lives in talkbank_model::resolve_clan_location().
Both the TUI and the desktop app call this function, it resolves line/column
from byte offsets when needed and adjusts for hidden headers.
See clan_location.rs
for the implementation and tests.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Batch Workflows
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
The chatter CLI is designed for processing large CHAT corpora efficiently. This page covers common batch workflows.
Validating a Corpus
Validate all .cha files in a directory tree:
chatter validate /path/to/corpus/
The validator recursively discovers .cha files and processes them in parallel. Results are cached, subsequent runs skip unchanged files.
Forcing Revalidation
To bypass the cache and revalidate everything:
chatter validate /path/to/corpus/ --force
Filtering Output
Show only errors (hide warnings):
chatter validate /path/to/corpus/ --quiet
Stop after the first reported error:
chatter validate /path/to/corpus/ --max-errors 1
The limit is at least 1 (--max-errors 0 is a usage error), and only errors
count: a warnings-only corpus is never stopped. When the count reaches the
limit, no new file starts, and if files were left the run says
Stopped after reaching --max-errors N; M file(s) were not validated. (on
stderr; a stop record in JSON mode) and exits 1. A run whose limit was
reached by its last file stopped nothing and never says it. The TUI honours
the limit too.
Write a JSONL audit file while validating:
chatter validate /path/to/corpus/ --audit validation.jsonl
CHAT-JSON Roundtrip
Convert an entire corpus to JSON and back:
# CHAT → JSON
for f in corpus/**/*.cha; do
chatter to-json "$f" > "${f%.cha}.json"
done
# JSON → CHAT
for f in corpus/**/*.json; do
chatter from-json "$f" > "${f%.json}.roundtrip.cha"
done
The roundtrip is designed to preserve the ChatFile model. In regression
tests, compare normalized output rather than assuming byte-for-byte identity
after parser or serializer changes.
Cache Management
The validation cache stores results for previously validated files
(keyed by content hash). The cache database file is named
talkbank-cache.db and lives in the OS cache directory:
- macOS:
~/Library/Caches/talkbank-chat/talkbank-cache.db - Linux:
~/.cache/talkbank-chat/talkbank-cache.db - Windows:
%LocalAppData%\talkbank-chat\talkbank-cache.db
It can hold results for large file collections.
To relocate the cache (a different disk, a per-project cache, or an
isolated cache for scripted runs), set the TALKBANK_CHAT_CACHE_DIR
environment variable to a directory; the database is created directly
inside it. This is the supported override on every platform, and the
only effective one on Windows, where the default location comes from
the system Known Folder API rather than environment variables.
chatter cache stats # Show the cache's location, size and entry count
chatter cache clear --all
Do not delete the cache file manually while chatter is running.
Reference Corpus Validation
This repository includes a reference corpus at
corpus/reference/ (currently ~100 .cha files; verify by
find corpus/reference -name '*.cha' | wc -l). The parser must
handle every file in this corpus at 100%:
cargo test -p talkbank-parser-tests reference_corpus_parses
This runs the parser equivalence test; each .cha file is its own test, so reports individual failures.
Integration with batchalign
The Batchalign pipeline uses the same Rust core (via PyO3) for CHAT
parsing and serialization. Batchalign itself is the upstream batchalign3
project, outside this repository. Files processed
by Batchalign produce valid CHAT that passes chatter validate.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CI Integration
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
How to use chatter in continuous integration pipelines.
Exit Codes
| Code | Meaning |
|---|---|
0 | All files valid / command succeeded |
1 | Validation errors found or command failed |
2 | Invalid arguments or missing required options |
All examples below rely on exit code 1 to signal validation failure.
Basic Usage
chatter validate corpus/ --quiet --tui-mode disable
--quietsuppresses per-file success output--tui-mode disablekeeps the interactive TUI away even on a terminal. A run whose stdout is not a terminal, or that asks for--quiet,--format jsonor--audit, never opens the TUI anyway.- Exit code 0 means all files valid; 1 means errors found
--format json, --quiet and --audit each name one output: --audit
with --format or --quiet, --format json with --quiet, and
--tui-mode force with any of them are usage errors (exit 2).
GitHub Actions Example
- name: Validate CHAT corpus
run: |
chatter validate corpus/ --audit results.jsonl
- name: Upload validation report
if: failure()
uses: actions/upload-artifact@v4
with:
name: validation-report
path: results.jsonl
The --audit results.jsonl flag streams one JSON line per diagnostic to a
file, each carrying its severity ("Error" or "Warning"), which is
useful for archiving or downstream analysis even when the step fails. The
record shape is in the diagnostic contract.
JSON Output for Automation
chatter validate corpus/ --format json --tui-mode disable 2>/dev/null
Each file produces a JSON object on stdout with status, error_count,
and errors array. The exit code still reflects overall pass/fail.
Pre-commit Hook
#!/bin/sh
# .git/hooks/pre-commit
chatter validate . --quiet --tui-mode disable
This blocks commits that introduce invalid CHAT files. The hook runs quickly on cached files; only modified files are re-validated.
Suppressing Specific Errors
Some corpora have known issues that should not block CI. Use --suppress
to ignore specific error codes or named groups:
chatter validate corpus/ --suppress E726,E727,E728 --tui-mode disable
Or use the named group shorthand:
chatter validate corpus/ --suppress xphon --tui-mode disable
Suppressed errors do not appear in output and do not affect the exit code.
Audit Mode for Large Corpora
For bulk corpus validation where you want a full error database without caching overhead:
chatter validate corpus/ --audit errors.jsonl --tui-mode disable
The --audit flag streams one JSON object per error to the specified file.
A summary is printed to stderr at the end.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CHAT Processing Playbook for Editors and Analysts
Status: Current Last updated: 2026-03-24 00:01 EDT
Objective
Provide practical guidance for non-compiler users who create, edit, and validate CHAT files, with emphasis on error interpretation and correction workflow.
Who This Is For
- Transcript editors,
- corpus curators,
- QA reviewers,
- linguists using tooling outputs but not parser internals.
Core Editing Workflow
- Open file in editor with CHAT diagnostics enabled.
- Run validation (single file first, then batch).
- Fix highest-severity structural issues first (headers, tier markers, unmatched delimiters).
- Re-run validation and inspect warnings.
- Only then address style and normalization suggestions.
Error Triage Heuristic
- Errors at file start: likely header formatting or encoding issues.
- Errors at tier prefix: likely malformed
*/%tier syntax. - Errors inside words: likely symbol, marker, or annotation boundary issues.
- Repeated same error class: likely one systemic rule violation pattern.
Fast Interpretation Guide
Error: parser/validator could not accept structure; must fix.Warning: valid but suspicious or non-canonical; review strongly recommended.Info: advisory normalization or convention hints.
Common Fix Recipes
- Header spacing problems:
- Ensure expected separators and avoid accidental tabs/spaces drift.
- Unclear language/form markers:
- Confirm
@susage and suffix ordering with house style guide.
- Confirm
- Duration/annotation confusion:
- Verify bracketed annotation form and avoid malformed punctuation.
- Dependent tier attachment issues:
- Ensure
%tiers follow intended main tier and keep indentation consistent.
- Ensure
Batch Validation Workflow
- Validate a small sample first.
- Group failures by error code.
- Fix by pattern, not file-by-file random order.
- Re-run and confirm error count decreases monotonically.
- Save run report for audit trail.
Collaboration Workflow with Developers
When reporting parsing issues, include:
- exact file path,
- minimal excerpt around failing span,
- observed diagnostic code/message,
- expected behavior (if known).
This reduces back-and-forth and speeds defect triage.
Quality Checklist Before Publishing Corpus Updates
- No unresolved error-level diagnostics.
- Warning classes reviewed and accepted or fixed.
- Participant headers and IDs internally consistent.
- Roundtrip serialization check passes for representative samples.
- Changelog note recorded for major normalization edits.
Training Recommendations
- Maintain short examples for each common error class.
- Provide editor cheat sheet for tier prefixes and marker syntax.
- Run periodic QA calibration sessions across editors.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Sanitize (chatter debug sanitize)
Status: Current Last updated: 2026-09-24 00:21 EDT
chatter debug sanitize strips contributor lexical content from a CHAT
file while preserving structure (timing bullets, %wor per-word bullets,
speaker codes, dependent-tier scaffolding, structural counts, POS tags,
language markers). It replaces supported lexical fields, but is not a
complete de-identification guarantee. Preserved identifiers, dates, gem labels,
display metadata and unknown recovery headers can retain identifying text. Review output under the
applicable data-sharing rules before disclosing it.
The command exists so engineering tooling, including LLM-assisted
debugging, can operate on protected-corpus files (aphasia/,
dementia/, rhd/, fluency/Password/, clinical-children corpora,
etc.) without exposing contributor speech to commercial LLM services.
When to use it
Run chatter debug sanitize as one preparation step before manual privacy
review. Do not send its output to another tool or person merely because the
command succeeded; unsupported fields and intentionally preserved metadata can
still identify participants.
When you need to ask a contributor for help debugging a specific file, frame the request as “run the sanitizer locally and send me the output” rather than asking for the raw file.
Usage
# Write sanitized output to stdout
chatter debug sanitize input.cha
# Write sanitized output to a file
chatter debug sanitize input.cha --output sanitized.cha
Working location for sanitized files: prefer a stable, non-/tmp
scratch directory (e.g. set TB_SCRATCH_DIR to a per-project dir
under your workstation’s persistent storage) for any state that
should outlive a single command. macOS clears /tmp on reboot.
What is preserved (byte-exact)
- Timing bullets
•start_end•on the main tier. %worper-word bullets (•start_end•after each word); the words beside them become the main tier’s placeholders, below.- Speaker codes (
*PAR,*INV,*CHI, …). - Utterance count and main-tier word structure; phonological tiers listed below are deliberately dropped, so dependent-tier count is not preserved.
- Structural markers: compound
+, clitic~, CA elements, overlap points, lengthening, stress markers, syllable pause, underline begin/end, proper-noun@nmarkers. - Language markers (
@s:LANG), form types (@a,@b), POS tags ($adj,$n). - Headers:
@Languages,@Birth,@Date,@Media,@PID,@L1Of,@Begin/@End/@UTF8. %morPOS categories and morphological features (e.g.,n|,-Past).%gra(numeric grammatical relations) and%tim(timing).- Untranscribed tokens
xxx/yyy/www, preserving them changes semantic meaning, so they pass through unchanged.
What is replaced or redacted
| Source | Replacement |
|---|---|
WordContent::Text | wN placeholder, indexed by document position |
WordContent::Phonetic (@u) | wN placeholder; phonetic speech can contain names |
Shortening text | (x) |
%mor lemmas (MorWord.lemma) | lemmaN; POS + features preserved |
%wor words | judged as timing recovery judges the tier, before the main tier is rewritten (count match, then word-by-word corroboration): a corroborating tier has each word become its paired main-tier word’s display text, now that word’s placeholder (wN, the same N; w1w1 for a compound), so it still corroborates the main tier; a drifted or uncorroborated tier takes fresh placeholders rather than a manufactured agreement. Bullets preserved |
%pho / %mod / %modsyl / %phosyl / %phoaln / %sin | tier dropped |
Free-text dependent tiers (%com %add %exp %sit %spa %int %gpx %act %cod %eng %gls %ort %flo %def %coh %fac %par %alt %err) | [redacted] |
@Comment, @Transcriber, @Birthplace, @Activities, @Situation, @RoomLayout, @Location, @TapeLocation, @Warning, @Bck | [redacted]; header identity and any speaker reference are retained |
@Participants participant-name field | dropped (Participant_<SPEAKER_CODE> is implied by speaker code + role) |
@ID custom_field and education | cleared |
Event event_type (&=imitates:Mary → &=redacted) | redacted |
Freecode text ([^ aside] → [^ redacted]) | redacted |
OtherSpokenEvent text | redacted |
Inline replacements use a delimiter-free token. The brackets in the free-text
marker [redacted] would introduce invalid nested CHAT syntax in these fields.
Header handling matches the complete typed header enum without a preservation
catch-all, so a new variant requires an explicit policy decision. The canonical
reference corpus witnesses all nine free-text payload types alongside comments,
and checks that header kind, speaker references and preserved metadata survive.
Determinism + Idempotence
Placeholder generation uses one monotonic counter in deterministic document traversal order. Two consequences:
- Deterministic: sanitizing the same input twice produces byte-identical output.
- Idempotent: sanitizing a sanitized file produces the same file again, no double-replacement, no shifting placeholder numbers.
Pipeline
flowchart LR
Input["Source .cha\n(protected corpus)"] --> Parser["TreeSitterParser\n(talkbank-parser)"]
Parser --> Model["ChatFile model\n(talkbank-model)"]
Model --> Sanitize["sanitize()\n(talkbank-transform::redact)"]
Sanitize --> Walker["walk_words_mut\n+ header walker\n+ dep-tier walker\n+ scoped-annot walker"]
Walker --> WordMutation["Typed Word replacement\n+ derived-text refresh"]
WordMutation --> Mutated["Mutated ChatFile\n(placeholders + redactions)"]
Mutated --> Writer["WriteChat\n(byte-exact bullets)"]
Writer --> Output["Sanitized .cha\n(scratch path)"]
The walker step replaces lexical segments in the typed word-content sequence,
mutates MorWord.lemma fields, redacts free-text
header / dep-tier / scoped-annotation strings, and drops phonological
tiers. WriteChat then re-serializes, and because it serializes from
typed content (not from Word.raw_text), every CA element, compound
marker, clitic boundary, and timing bullet round-trips byte-exact.
Word content is externally read-only. Replacement must use Word’s mutation
methods, which invalidate the derived cleaned_text cache. The sanitizer also
rebuilds raw_text from the complete sanitized typed word on every path. That
includes parser-recovery words containing only structural elements and the
xxx/yyy/www pass-through path. This prevents either JSON-facing string
field from retaining source lexical material or losing nonlexical word markers.
Out of v1 scope
Documented for transparency; v2 work:
- Speaker-code anonymization (graph rewrite across
@Participants,@ID,*SPK:,@Birth,@L1Of). @Birth/@Datefuzzing (exact birth dates can be identifying).@Mediafilename redaction.- Audio-side sanitization. (Audio bytes are never touched by the sanitizer; the audio stays at its original path.)
- “Unsanitize” or round-trip mapping. Explicitly not built, the sanitizer is one-way, the mapping table that would reverse it is the exact artifact we don’t want to exist.
Implementation
Library module: talkbank_transform::redact. CLI surface: chatter debug sanitize.
The strict policy is the only public preset in v1; future variants can
grow on SanitizationPolicy.
This page last changed: 2026-09-24 (commit 0089c3eb). The whole book last changed: 2026-10-07 (commit 5e895791).
Speaker-ID (chatter speaker-id)
Status: Draft Last modified: 2026-10-07 (commit 5e895791)
chatter speaker-id assigns CHAT-conformant speaker codes and role
tags to a CHAT file whose speakers carry anonymous or placeholder
labels (typically the output of an ASR system that labels speakers
as PAR0, PAR1, …). It is the bridge between an ASR pipeline that
does not understand speaker roles and a CHAT pipeline that does.
The command is structural: it does not modify utterance content,
does not run audio analysis, does not infer speaker identity from
voice features. Its inputs are the CHAT file to relabel plus an
identification signal (reference transcript, explicit mapping, or
saved override record); its output is the same CHAT file with
speaker codes rewritten and @Participants / @ID headers
reconciled.
When to use it
Whenever you have a CHAT file with placeholder speaker codes that need to become CHAT-conformant codes before downstream tooling can process the file meaningfully. The canonical case is an ASR system that emits CHAT but does not know which speaker is the child, parent, clinician, etc.
A complete pipeline that consumes ASR output and produces a publishable CHAT file goes:
flowchart LR
Media --> Transcribe
Transcribe["batchalign3 transcribe<br/>ASR"] --> AsrAnon["asr.cha<br/>PAR0, PAR1, ..."]
Ref["reference.cha<br/>target speakers only"] -.->|reference signal| SpkId
AsrAnon --> SpkId
SpkId["chatter speaker-id<br/>(this page)"] --> AsrLabeled["asr-labeled.cha<br/>CHI, INV, MOT, ..."]
AsrLabeled --> Merge["Structural assembly (library)"]
Ref --> Merge
Merge --> Aligned["batchalign3 align"]
The speaker-id stage is the single point in the pipeline where
“which anonymous speaker corresponds to which CHAT role” is
decided. Downstream stages (structural assembly, batchalign3 align,
batchalign3 morphotag) all trust that the labels they receive
are correct.
Identification modes
Three mutually-exclusive modes, exactly one of which must be selected:
1. Reference mode
The most common case: a separate CHAT file already exists that
covers the same media and contains an authoritative speaker
(typically the hand-transcribed target speaker). The reference
file’s anchor speaker tells us what that speaker’s content looks
like; speaker-id finds the matching speaker in the input by
text similarity.
The matching algorithm is multiset Jaccard over bags of content
tokens, see “Algorithm” below for the full specification. The
ASR speaker whose bag-of-words best matches the reference anchor’s
bag-of-words is taken as the same speaker, and is marked for
drop in the output (because the reference file authoritatively
covers them, a downstream structural assembly stage will pull
their utterances from the reference, not from this file). The
remaining speakers are renamed to the role specified by
--inserted-role.
If the Jaccard margin between the winning speaker and the
runner-up is below --confidence-threshold, the command refuses
to auto-decide. The operator must either lower the threshold
(not recommended without spot-checking), supply an explicit
mapping (--mapping), or load a previously-adjudicated override
(--override-file).
2. Explicit-mapping mode
The operator already knows the mapping (typically because they listened to the audio, or because the contributor’s data sheet documents it). They supply it directly.
chatter speaker-id input.cha \
--mapping "PAR0=INV:Investigator,PAR1=drop" \
-o relabeled.cha
The grammar for --mapping:
- One or more comma-separated assignments.
OLD=CODE:ROLErenames OLD to CODE with role tag ROLE.OLD=dropremoves OLD’s utterances entirely.- Every speaker present in the input must be named in the mapping (no defaulting). This is intentional, we want operator decisions to be explicit.
3. Override-file mode
The operator has previously adjudicated this session (perhaps
through an interactive review tool) and saved the decision to a
shared override file. speaker-id reads the file, finds the entry
for this session, and applies it. See “Override file format”
below.
chatter speaker-id input.cha \
--override-file batch-2026-05-27.overrides.toml \
--session-id S01-1 \
-o relabeled.cha
This mode is the production substrate for batch workflows: the
orchestrator first runs chatter speaker-id in reference mode for
every session; for any session that exits with low-confidence, the
operator works through an adjudication tool that writes to the
override file; the orchestrator then re-runs chatter speaker-id
in override-file mode for those sessions.
CLI contract
chatter speaker-id <INPUT> [OPTIONS]
ARGUMENTS:
<INPUT> Path to the CHAT file to relabel.
OPERATION MODES (exactly one required):
REFERENCE MODE:
--reference <FILE>
--anchor <SPEAKER>
--inserted-role <CODE>:<TAG>[,<CODE>:<TAG>...]
EXPLICIT-MAPPING MODE:
--mapping <SPEC>
OVERRIDE-FILE MODE:
--override-file <FILE>
--session-id <ID>
REFERENCE-MODE OPTIONS:
--confidence-threshold <FLOAT>
Minimum Jaccard margin (winner_score / loser_score) for the
command to auto-decide. Below threshold: exit code 4. The
command prints per-speaker scores to stderr so the operator
can inspect. Default: 2.0.
--write-match-report <NEW-FILE.json>
Write a typed report for the complete reference-mode attempt.
Matched outcomes record reference, donor, shared, and union token
counts plus the derived score and margin. Structural and input
refusals record their own outcome-specific evidence. The report
never overwrites an existing file.
--write-override <FILE>
When auto-decide succeeds, append the decision to FILE in
override-file format (creates if missing). Captures the
audit trail of a batch run.
COMMON OPTIONS:
-o, --output <PATH>
Write relabeled CHAT to PATH. Default: stdout.
The operator identity and any free-text note for a session are set
when an operator confirms it through `chatter adjudicate` (see the
merge workflow), not on this command.
Exit codes:
| Code | Meaning |
|---|---|
| 0 | Success, relabeled file written |
| 1 | Invalid input (parse error, missing file, unreadable) |
| 2 | Semantic precondition violated (reference has no utterances for anchor; mapping covers a speaker not in input; etc.) |
| 3 | Internal error |
| 4 | Reference mode: confidence threshold not met. Per-speaker scores printed to stderr; no output written |
What the output guarantees
These are testable invariants. Every release verifies them against the reference corpus.
Speaker codes match the supplied mapping
For every speaker in the input file:
- If the mapping marks the speaker for drop, none of their
utterances appear in the output, AND their
@IDrow (if any) is removed from the headers, AND their entry is removed from the@Participantsheader. - If the mapping marks the speaker for rename, every main-tier
line
*OLD:\t...becomes*NEW:\t...byte-stable except for the speaker code prefix. The@IDrow’s third pipe-separated field (speaker code) and eighth field (role tag) are rewritten; other@IDfields are preserved. The@Participantsentry’s code and role-tag tokens are rewritten; any intervening tokens (corpus ID, participant name) are preserved. - Speakers not in the mapping are passed through unchanged. (In modes 1 and 3, all speakers are assigned automatically; in mode 2, “all speakers must be in the mapping” is a precondition.)
Utterance content is byte-stable except for the speaker prefix
For every retained utterance, every byte EXCEPT the leading
*CODE:\t prefix is preserved verbatim. Dependent tiers attached
to the utterance are preserved exactly. NAK-delimited time
bullets, CHAT markup, special-form annotations, paralinguistic
codes, retracing scopes, all untouched.
Headers reconcile per a fixed table
| Header | Behavior |
|---|---|
@UTF8, @Begin, @End, @Window, @Languages, @Media | Pass-through unchanged |
@Participants | Drop entries for dropped speakers; rewrite code + role-tag for renamed speakers; entries for unaffected speakers preserved |
@ID | Drop rows for dropped speakers; rewrite field 3 (code) and field 8 (role) for renamed speakers; other fields preserved |
@Comment | Pass-through unchanged (provenance-carrying comments survive) |
Provenance is captured if --write-override is set
When --write-override <FILE> is supplied AND the command succeeds
in reference mode, an entry is appended to FILE recording the
session ID (derived from the input filename stem unless overridden),
the per-speaker Jaccard scores, the chosen mapping, the operator, and
an ISO 8601 timestamp. The format is specified in “Override file
format” below. The operator identity and any free-text note are set
later, when a session is confirmed via chatter adjudicate.
This is the audit-trail mechanism: a year from now, a researcher who asks “why was PAR0 labeled INV in this session?” can read the override entry and see the scores, the operator, and any notes the operator added.
Algorithm (reference mode)
Token cleaning
Both the reference anchor’s bag of words and each input speaker’s
bag of words are built by walking the typed CHAT AST. Word leaves supply
their model-owned cleaned_text; the matcher does not independently strip
CHAT markup or run a parallel text parser. It trims each value, applies ASCII
lowercasing, and retains only entirely ASCII-alphabetic tokens of at least
two characters. Non-word items do not enter the bag.
Comma, tag-question and vocative separators contribute no tokens. The reference separator control checks exact support counts and verifies that a stricter confidence threshold refuses the same evidence without changing it.
The current matcher also skips an entire ReplacedWord (word [: replacement])
node, including its original word. Unlike punctuation, that node does contain
lexical content. This is an unresolved matching-policy limitation, not a claim
that replaced speech contains no words; inspect the support counts when using
replacement-annotated transcripts.
The ASCII restriction is a limitation of the matcher, not a CHAT validity rule:
accented and non-Latin words are excluded. For example, the Mandarin reference
transcript yields no lexical information even when compared with itself, and
the operation refuses with low_confidence / no_information; it must not
promote the first speaker in document order into an accepted identity match.
Mixed-language speech is scored only on the eligible tokens, so inspect the
reported token counts rather than treating a high score as multilingual evidence.
Both sides use the same policy, and identification does not rewrite the input.
Multiset Jaccard
For two bags-of-words A and B (counted multisets):
J(A, B) = sum_w min(A[w], B[w]) / sum_w max(A[w], B[w])
Range [0, 1]. The multiset (rather than set) form rewards
speakers who say similar things to the anchor in similar volume,
not just speakers whose vocabulary happens to intersect.
Decision
scores = { speaker: J(anchor_bag, speaker_bag) for speaker in input }
winner = argmax(scores)
loser = argmax(scores - {winner})
margin = scores[winner] / scores[loser] # ∞ when loser score = 0
winneris the input speaker whose content matches the reference anchor’s content best → marked for drop (the reference authoritatively covers them).loser(and any other lower-scoring speakers, in the multi-speaker case) → renamed to the role given by--inserted-role.
Match-evidence report
--write-match-report preserves the observations behind the scalar score and
the typed reason when matching could not begin.
This matters because a ratio alone cannot distinguish a winner supported by
one shared token from one supported by hundreds. The JSON schema records, for
each donor speaker, reference_tokens, donor_tokens, shared_tokens,
union_tokens, and the derived score. Its margin is one of
no_information, finite, or unbounded; no-information and a zero-scoring
runner-up are not represented by floating-point sentinels.
The outer outcome is accepted, low_confidence,
reference_missing_anchor, donor_too_few_speakers, or input_rejected.
Matched outcomes contain the lexical report; other outcomes cannot pretend to
contain one. Input rejection records donor versus reference, the typed
pipeline-failure category, and TalkBank diagnostic codes when available.
Within input_rejected, failure_kind: "incomplete_validation" means missing
or recovered parser provenance prevented complete validation. This differs from
"validation" (the model failed validation) and "internal_failure" (the tool
failed). Imported JSON may have unchanged model content but lack parser
provenance, so an incomplete result can have an empty diagnostic_codes array.
It still cannot supply a match_report. Consumers must handle this distinct
category rather than treating every input rejection as invalid CHAT. The outer
outcome and schema version remain unchanged.
The option is intended for audit and calibration tooling. It does not change which speaker wins or the confidence threshold. Chatter stages the complete JSON beside its destination and persists it without clobbering, so a later run cannot silently replace the evidence used for an earlier decision and a failed write cannot leave a partial final report. The no-clobber persist is not guaranteed to be atomic on every platform; an interrupted persist can leave the staging link behind, but it never replaces an existing report.
If margin < --confidence-threshold (default 2.0), the command
exits with code 4 and prints per-speaker scores to stderr. The
operator must inspect, adjudicate, and re-run with
--mapping or --override-file.
Why this algorithm
The choice was empirical, not theoretical, and was made against a calibration set of CHAT files paired with their corresponding ASR output. Two earlier candidates were tested first and rejected:
- Raw temporal-overlap (sum of ms of an input speaker’s activity inside the anchor’s bullet windows): too weak on real data. Hand transcripts often place per-utterance time bullets as end-to-end segmentation boundaries covering 95-99% of the session timeline, rather than as tight “speaker active here” windows. Both input speakers fall almost entirely “inside” the anchor’s bullet windows and the signal disappears.
- Speaker purity (fraction of each input speaker’s activity falling inside anchor windows): same root cause, same failure.
Multiset Jaccard over content tokens succeeded on every
session of the calibration set. The borderline cases (margin
below 2.0x) clustered around tasks where the non-anchor speaker
shares vocabulary with the anchor by the structure of the task,
e.g. a clinician describing the same scene the child is also
describing in a picture-narrative task. These borderline cases
are the reason for the conservative threshold and the
--mapping/--override-file escape hatches; the algorithm
correctly refuses to auto-decide them rather than silently
picking wrong.
Override file format
The override file is a UTF-8 TOML document with one
[<session_id>] table per decision. A minimal entry:
schema_version = 2
[session-101-t1]
mode = "auto"
adult_roles = { PAR0 = { code = "INV", tag = "Investigator" } }
mapping = { PAR0 = "rename", PAR1 = "drop" }
scores = { PAR0 = 0.1931, PAR1 = 0.7347 }
margin = 3.81
operator = "alice"
decided_at = "2026-05-27T08:41:00-04:00"
The complete schema specification, every field, every type, every mode-semantics rule, the strict refuse-with-clear-error versioning policy, and worked examples for auto/explicit/replay/diarization-mixed cases, is on the dedicated reference page: Merge Override File Format.
Highlights from the reference:
mode = "auto" | "explicit" | "override"records how the decision was made (informational for audit trail; behavior at apply time is the same).adult_rolesmaps each renamed speaker’s donor code to its own role assignment:adult_roles[<donor_code>].codeis the CHAT speaker code (INV,MOT,FAT,PAR, …);.tagis the CHAT role-tag (Investigator,Mother, …). Renamed speakers in one entry may share a role or each carry a distinct one.mappingmust cover every speaker in the input, no defaulting.scoresandmarginare optional but the writer always records them when an auto attempt produced them (even when the final decision was operator-supplied).flagscarries operator-supplied markers like"diarization-mixed"for unusual cases. Unknown strings are preserved verbatim.
Preconditions
chatter speaker-id refuses (exit code 2) if any hold:
Reference mode
- The reference file has no utterances for
--anchor - The reference file fails to parse
- The input file has fewer than 2 distinct speakers (no discrimination problem)
Explicit-mapping mode
- A speaker in the mapping is not present in the input
- A speaker in the input is not covered by the mapping (no defaulting)
Override-file mode
- The override file does not contain a
<session-id>entry - The entry’s
mappingreferences a speaker not in the input - The entry’s mapping does not cover every speaker in the input
What chatter speaker-id is NOT
- Not voice diarization. Use Batchalign’s ASR pipeline upstream; the labels this command consumes are the labels Batchalign emits.
- Not content correction. If the speaker the command identifies has been mis-transcribed by ASR, this command does not fix that , re-run ASR with a better engine.
- Not a merge. This command operates on a single CHAT file. To
combine it with a reference, applications must resolve event correspondence
before using structural assembly. The
mergeCLI was removed. - Not interactive.
chatter speaker-idis batch-only: it succeeds, refuses, or fails. The interactive review that resolves a low-confidence refusal into an override-file entry is a separate command,chatter adjudicate, run as part of the merge workflow.
Worked example
A typical fully-automated reference-mode call from an orchestrator script:
chatter speaker-id asr-anonymous.cha \
--reference hand-transcript.cha \
--anchor CHI \
--inserted-role INV:Investigator \
--confidence-threshold 2.0 \
--write-override batch.overrides.toml \
-o asr-labeled.cha
For a session this refused (e.g., shared-vocabulary narrative task with margin 1.82x), the orchestrator captures the failure and the operator later resolves it:
# Inspect the scores the command emitted to stderr:
# PAR0=0.6286 PAR1=0.3457 margin=1.82x threshold=2.0
# Operator listens to a few seconds of audio and confirms PAR0 is
# the child:
chatter speaker-id asr-anonymous.cha \
--mapping "PAR0=drop,PAR1=INV:Investigator" \
--write-override batch.overrides.toml \
-o asr-labeled.cha
Later, if anyone re-runs the batch, they use override-file mode:
chatter speaker-id asr-anonymous.cha \
--override-file batch.overrides.toml \
--session-id S02-1 \
-o asr-labeled.cha
The same asr-labeled.cha content is produced; the audit trail
remains intact.
Implementation notes (for contributors)
- Source:
crates/talkbank-transform/src/speaker_id/. - CLI surface:
crates/chatter/src/commands/speaker_id/. - CHAT-domain types such as
SpeakerCodeandParticipantRolelive intalkbank-model. Speaker-identification types such asMappingSpec,MergeOverride,LexicalMatchEvidence,JaccardScore,ConfidenceThreshold, andConfidenceMarginlive beside the algorithms intalkbank-transform::speaker_id. - The Jaccard cleaner walks
talkbank-model::ChatFiledirectly via the existing content walker (talkbank-model::walk_words); it does NOT re-implement CHAT parsing or use regex on raw bytes for tokenization. - Spec entries for the cleaner and the algorithm live in
spec/constructs/speaker-id/. Every invariant on this page has a spec; regenerate them with the currentspec/toolscommands from Spec Workflow. - The override-file reader/writer is a typed
serderound-trip on a TOML representation owned bytalkbank-transform::speaker_id, so consumers use one shared parser rather than duplicating the format.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Rediarize (chatter rediarize)
Status: Draft Last updated: 2026-09-24 00:21 EDT
chatter rediarize re-attributes utterance speakers in a CHAT file
from an external diarization. Given a transcript whose utterances
carry media time bullets and a JSON file of timestamped speaker turns
produced by a dedicated diarizer (for example pyannote), it reassigns
each utterance’s main-tier speaker to the diarization track that
covers the utterance’s time span the most, keeping the utterance
content (the words) byte-stable.
The command exists for a specific, common failure shape: ASR systems
with bundled diarization (Rev.AI and others) auto-detect the speaker
count and can under-count on hard material such as child-adult
overlap, collapsing three or four real voices into two tracks. The
ASR words are usually fine; the attribution is what is wrong. A
dedicated diarizer recounts the voices correctly, and rediarize
reconciles its turns with the existing transcript so you keep the
good words and replace only the bad attribution.
The command is structural and audio-free: it never touches the recording. The diarizer runs elsewhere (any tool, any model) and hands its result across a documented JSON boundary.
Pipeline position
flowchart LR
Media["recording\n(audio)"] --> Diarizer["external diarizer\n(e.g. pyannote)"]
Diarizer --> Turns["turns.json\n(documented format below)"]
Media --> Asr["ASR with bundled\ndiarization"]
Asr --> AsrCha["asr.cha\ngood words,\nsuspect speaker tracks"]
AsrCha --> Rediarize["chatter rediarize\n(this page)"]
Turns --> Rediarize
Rediarize --> Fixed["rediarized.cha\nPAR0..PARn correctly\nseparated tracks"]
Fixed --> SpkId["chatter speaker-id\n(assign real roles)"]
rediarize fixes WHICH anonymous track owns each utterance; it does
not decide who each track is. Role assignment (child, mother,
investigator, …) is chatter speaker-id’s job,
downstream.
Usage
chatter rediarize INPUT.cha --turns TURNS.json -o OUTPUT.cha
Omitting -o prints the rewritten CHAT to stdout.
A summary is reported on stderr after the rewrite (stderr so that a
stdout CHAT stream stays clean when -o is omitted):
rediarize: 214 reassigned, 671 unchanged, 7 flagged
Flagged utterances (see below) are listed individually with their utterance index, kept speaker, and reason.
--contested-at SHARE additionally reports utterances whose time is
split between tracks; see Contested utterances.
Machine-readable summary (--summary-json)
Batch drivers looping rediarize over a corpus should not scrape the
stderr text. --summary-json PATH additionally writes the outcome as
JSON:
chatter rediarize INPUT.cha --turns TURNS.json \
-o OUTPUT.cha --summary-json SUMMARY.json
{
"source": "pyannote/speaker-diarization-community-1",
"reassigned": 747,
"unchanged": 145,
"flagged": [
{"utterance_index": 12, "kept_speaker": "PAR1",
"reason": "no_overlapping_turn"}
],
"contested": [
{"utterance_index": 41, "assigned": "PAR2",
"ownership": {"shares": [["PAR2", 600], ["PAR1", 400]],
"total_ms": 1000}}
]
}
-
source: the turns file’s provenance, passed through (nullif the turns file carried none). -
reassigned/unchanged: utterance counts.unchangedincludes flagged utterances (they kept their speaker), so the file’s total bulleted-tier utterance count isreassigned + unchanged. -
flagged: every declined reattribution, never truncated (the stderr listing caps at 20 detail lines; this list is complete).utterance_indexis the 0-based position among main-tier lines;reasonis"no_bullet"or"no_overlapping_turn". -
contested: utterances whose time was meaningfully split between tracks, empty unless--contested-atwas given. These were still reattributed, toassigned, so they are NOT inflagged, which means “declined”.ownership.sharesis every overlapping track with its union-held milliseconds, descending; overlapping turns for the SAME track count their shared interval once.ownership.total_msis the sum of those per-track values and is the denominator. Simultaneous DIFFERENT tracks each retain the shared interval, sototal_mscan exceed the utterance bullet’s duration. The whole distribution is emitted rather than a winner and a runner-up, because that narrower shape cannot tell a 55/45 split from 55/23/22 and the difference is the point.
Field names and the reason strings are a stable output contract.
The summary is written only on exit 0, after the CHAT output.
Contested utterances (--contested-at)
An utterance’s bullet can overlap turns from more than one track. The tool assigns it to the track holding the most of it, which is the best available answer, but “most” can mean 95% or 34%, and those are different situations that the output otherwise reports identically.
chatter rediarize INPUT.cha --turns TURNS.json --contested-at 0.25
Reports an utterance as contested when the RUNNER-UP track holds at least that share of the total track-held time. Two rivals at 20% each is a different situation from one at 40%, and this is the latter question. Same-track duplicate coverage never inflates the denominator; cross-track overlap remains evidence for both simultaneous speakers.
There is no default, deliberately. Omit the flag and nothing is reported as contested. What share makes an utterance genuinely mixed has not been measured against human listening, so shipping a number here would hand every user a constant wearing this tool’s authority. Supply one you can defend, or none.
The flag changes reporting only: placement is byte-identical with
and without it. A value outside 0.0 to 1.0, or NaN, fails the
command before any file is read, rather than silently meaning “nothing
is ever contested”.
Known limitation: per-track union-held totals cannot distinguish a speaker change INSIDE an utterance (one track holds the first half, the other the second) from crosstalk (both across the whole), and those want opposite remedies. Contested says ownership is divided, not how it is arranged on the timeline.
The turns JSON format
The --turns file is the corpus-agnostic seam between the diarizer
and chatter. Producing it from any given diarizer’s native output is
the caller’s concern; the format is:
{
"source": "pyannote/speaker-diarization-community-1",
"turns": [
{"track": "PAR0", "start_ms": 12063, "end_ms": 17024},
{"track": "PAR1", "start_ms": 13379, "end_ms": 14375}
]
}
source(optional): free-form provenance, typically the diarizer model name. Not interpreted, but useful in audit trails.turns(required): the timestamped segments. Each has:track: the anonymous CHAT speaker code this segment belongs to (PAR0,PAR1, …). The producer chooses the codes; a deterministic mapping from diarizer-native labels (for example pyannote’sSPEAKER_00) is recommended.start_ms/end_ms: the segment’s media time span in integer milliseconds, half-open[start_ms, end_ms), withend_ms >= start_ms.
Turns MAY overlap each other (diarizers that permit overlapping speech produce such turns). Overlap between turns for the same track is unioned; overlap between different tracks is retained for each track as crosstalk evidence. Input order is immaterial: chatter admits the turns to a typed, start-ordered timeline before attribution. Unknown fields anywhere in the file are rejected, so a misspelled field fails loudly instead of being silently ignored.
Behavior contract
The Rust rediarize API requires an error sink. After reconciling headers it
uses the canonical participant join, reporting inconsistencies before exposing
the resulting participant map. The returned mutable model is not a full
validation certificate. The content-level wrapper refuses to serialize when
that join reports errors; it does not silently discard them.
- Every utterance with a time bullet is assigned to the track with the
greatest union of millisecond coverage against the bullet’s span.
An utterance already on its max-overlap track counts as
unchanged. - An utterance with no bullet, or whose bullet overlaps no turn at all, keeps its existing speaker and is flagged in the summary. Ambiguity is surfaced, never silently guessed.
@Participantsand@IDheaders are reconciled to declare exactly the set of tracks the output actually uses: new tracks get entries cloned from an existing participant (same role), declarations for tracks no longer used are dropped.- Header-only transcripts preserve their declarations: no utterances means there is no attribution evidence for pruning participants.
- Utterance content, dependent tiers, and all other headers are preserved as-is.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Rewrite completed and output written. Flagged utterances do not fail the command; check the summary. |
| 1 | Invalid input: unreadable file, CHAT parse failure, malformed turns JSON. |
| 2 | Precondition violation: the turns JSON parsed but is semantically defective (for example a turn with end_ms < start_ms). |
On any non-zero exit, no output file is written.
Worked example
A recording of one child and two parents, transcribed by an ASR
whose bundled diarization auto-detected two speakers (the two adults
were merged into one track). A dedicated diarizer found three voices
and produced turns.json with PAR0/PAR1/PAR2. Then:
chatter rediarize session.cha --turns turns.json -o session-3spk.cha
chatter validate session-3spk.cha
splits the merged adult track by time, declares PAR2 in the
headers, and leaves every word as the ASR wrote it. The output then
flows into chatter speaker-id (or the merge workflow) to name the
three tracks.
This page last changed: 2026-09-24 (commit 0089c3eb). The whole book last changed: 2026-10-07 (commit 5e895791).
Removed commands: merge, pipeline, and batch
Last modified: 2026-09-28 20:59 EDT
The experimental chatter merge, chatter pipeline, and chatter batch
commands have been removed from the CLI.
Earlier releases exposed structural interleaving of two CHAT transcripts under
this name. That operation assumed the caller had already resolved which speech
events and speakers to retain. It did not match uncertain transcriptions to
recorded speech, reconcile different segmentation, or correct speaker attribution.
The command name encouraged a broader interpretation than the operation supported.
There is no drop-in CLI replacement for fuzzy transcript-to-recording matching.
The former pipeline and batch commands composed speaker mapping and
structural assembly; they did not supply that missing event-matching step.
For applications that have already resolved correspondence and source selection,
the typed structural APIs remain in talkbank_transform::transcript_merge.
They preserve their validated input, draft, and reported-output transitions.
Removing this CLI does not remove or change those library contracts.
Explicit review drafts in the library
Structural assembly normally refuses unresolved cross-source ordering. A caller
can explicitly opt in through
SourceBoundDonorSelection::with_flagged_draft_order, bound to the exact
reference document. This uses reference-first serialization at unresolved
frontiers and preserves both source sequences. Uncertainty is returned only as
structured DraftOrderReview records; no generated review @Comment lines are
inserted. Contributor comments remain unchanged. Callers present review information
outside the transcript using MergeDraft::draft_order_reviews() before validation
or Merged::draft_order_reviews() afterward.
Each record carries an OutputUtteranceBoundary, not a CHAT line index. Its
utterances_before() count excludes all headers and comments; zero means before
the first utterance. The reason distinguishes competing utterances, a section
against an utterance, and competing sections. Optional section navigation bounds
are milliseconds from neighboring recorded speech, not inferred section times.
That convention does not establish chronology or task membership. It does not authorize speech deletion, invented timestamps, or overriding contradictory timing evidence. Model validation remains a separate required transition; validation success does not adjudicate the recorded ordering uncertainties. The canonical untimed conversation and disjoint-timing reference tests exercise these distinctions without changing their source transcripts.
This page last changed: 2026-09-28 (commit 2cb42a45). The whole book last changed: 2026-10-07 (commit 5e895791).
Review Tools (adjudicate, sanity-scan)
Last modified: 2026-10-02 (commit 2d7e886b)
These experimental tools review speaker decisions and inspect existing output.
merge, pipeline, and batch are not commands (see the
removal notice). These review tools do not replace event correspondence.
Use speaker-id --write-pending to prepare unresolved speaker decisions.
chatter adjudicate (the operator step)
Reads the pending file a pass produced, walks the operator through the unresolved sessions, and appends the resolved decisions to the override file. On success the pending file is rewritten to drop the entries that were resolved, so re-running adjudicate only ever shows what is left.
chatter adjudicate <PENDING> --override-file <FILE> [--interactive | --scripted <TOML>]
ARGUMENTS:
<PENDING> The pending-adjudications TOML a pass wrote.
REQUIRED:
--override-file <FILE> Override file to append resolved decisions to
(created if absent). This is the same file pass 2
reads back.
DECISION SOURCE (one of):
--interactive Prompt per pending entry on stdin. See "The
interactive decision language" below for the
three decision verbs and their syntax.
--scripted <TOML> Pre-canned operator decisions, for replayable /
tested runs. Mutually exclusive with --interactive.
--operator <NAME> Recorded in each override entry (defaults to $USER).
The interactive decision language
Each pending entry is printed with its full context (the sessions, the suggested mapping, the engine’s confidence scores and reasoning), then one line is read from stdin. Three decision verbs are accepted:
| Verb | Form | Meaning |
|---|---|---|
accept (or a) | accept [note...] | Take the suggested mapping exactly as proposed |
choose | choose SPK:CODE:TAG [SPK:CODE:TAG ...] [note...] | Supply the speaker mapping yourself: each group maps a donor speaker to a CHAT code and role tag |
override | override SPK:CODE:TAG [SPK:CODE:TAG ...] SPK=action [SPK=action ...] [note...] | Supply the mapping AND per-speaker actions (for example SPK=drop to exclude a donor speaker entirely) |
SPK:CODE:TAG groups are repeatable, so multi-adult sessions are
expressed naturally, one group per speaker:
choose A:CHI:Target_Child B:INV:Investigator C:MOT:Mother reviewed against the recording
Anything after the structured arguments is recorded verbatim as the operator’s note. Every decision (verb, mapping, note, operator, and the engine’s original scores) is appended to the override file, so the audit trail survives the session.
This is the interactive review tool the speaker-id page
refer to: the audit trail (who decided, the scores, any note) lands in
the override file so a later reader can see why a session was labeled
the way it was. The decision schema is the same override-file format
used everywhere in the workflow; see
Merge Override File Format, and the
Adjudication Workflow
architecture page for the design.
chatter sanity-scan (post-merge QA)
A confident auto-decision can still be wrong, the runner-up was simply
even further off. sanity-scan re-reads the merged output and the
pass-1 audit file and flags sessions that pass an out-of-band check: the
mean utterance word count of the anchor speaker versus the inserted
speaker. In a typical child-language recording the adult out-talks the
child, so an anchor (child) mean that is much higher than the inserted
(adult) mean is suspicious, possibly the two were swapped.
chatter sanity-scan <MERGED_DIR> \
--override-file <FILE> --anchor <SPEAKER> --write-pending <FILE> [OPTIONS]
REQUIRED:
--override-file <FILE> The pass-1 audit file. Only auto-decided sessions
are scanned; explicit-mode entries are skipped (the
operator already signed off).
--anchor <SPEAKER> Anchor code in the merged files (typically CHI).
--write-pending <FILE> Flagged sessions are appended here as
sanity-scan-misclassification pending entries for
`chatter adjudicate`. Required.
--threshold <F> Flag when anchor_mean >= inserted_mean * threshold
(default 1.5).
A flag is a question, not a verdict: the session goes back into the adjudication queue for an operator to confirm or correct. Whether to run the scan at all is a judgment about the corpus. It assumes the typical “adult out-talks child” shape, and is unreliable where that inverts (e.g. a clinical-interview corpus where children out-narrate the adult); there, review speaker identity using appropriate evidence instead.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CHAT Format Overview
Status: Reference Last updated: 2026-05-11 21:51 EDT
CHAT (Codes for the Human Analysis of Transcripts) is a standardized transcription format for spoken language data, developed by MacWhinney as part of the CHILDES and TalkBank projects. It is the most widely used format in child language research and conversational analysis.
File Anatomy
Every CHAT file follows this structure:
@UTF8
@Begin
@Languages: eng
@Participants: CHI Target_Child, MOT Mother
@ID: eng|corpus|CHI|2;6.||||Target_Child|||
@ID: eng|corpus|MOT|||||Mother|||
*MOT: what do you want ?
%mor: ADV|what AUX|do PRON|you VERB|want ?
%gra: 1|4|OBJ 2|4|AUX 3|4|NSUBJ 4|0|ROOT 5|4|PUNCT
*CHI: I want cookie .
%mor: PRON|I VERB|want NOUN|cookie .
%gra: 1|2|NSUBJ 2|0|ROOT 3|2|OBJ 4|2|PUNCT
@End
A CHAT file consists of:
@UTF8: required first line, declares UTF-8 encoding@Begin: marks the start of the transcript- Headers: lines starting with
@that provide metadata (participants, languages, IDs, etc.) - Utterances: blocks consisting of:
- A main tier (line starting with
*SPEAKER:) containing the transcribed speech - Zero or more dependent tiers (lines starting with
%tier:) containing annotations
- A main tier (line starting with
@End: marks the end of the transcript
Key Conventions
- Tab separation: a tab character separates the tier prefix from its content (e.g.,
*CHI:⟶content) - Terminators: every utterance ends with a terminator (
.,?,!, or special forms like+...) - Line continuation: long lines wrap with a tab at the start of continuation lines
- Speaker codes: short identifiers; the validator accepts up to seven characters from
A-Z,0-9,_,-,'; three uppercase letters is the convention (e.g.,CHI,MOT,FAT,INV) - Media linking: timestamps link transcripts to audio/video via bullet markers
CHAT vs Other Formats
| Feature | CHAT | Praat TextGrid | ELAN EAF |
|---|---|---|---|
| Morphological tiers | Built-in (%mor, %gra) | No | No |
| Dependency syntax | Built-in (%gra) | No | No |
| Standardized POS | UD-style via %mor | No | No |
| Word-level alignment | %wor tier | Interval-based | Interval-based |
| Error recovery | Tree-sitter GLR | N/A | N/A |
References
- CHAT Manual: the canonical reference
- TalkBank: the data repository
This page last changed: 2026-07-27 (commit 905227c9). The whole book last changed: 2026-10-07 (commit 5e895791).
Headers
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
Headers are lines beginning with @ that provide metadata about the transcript. They appear between @Begin and the first utterance (though some headers like @Comment can appear anywhere).
Required Headers
@UTF8
Must be the very first line of every CHAT file. Declares UTF-8 encoding.
@UTF8
@Begin / @End
Mark the start and end of the transcript body. Every CHAT file must have exactly one @Begin and one @End.
@Participants
Declares all speakers in the transcript. Format:
CODE [Name] Role, comma-separated. The role is required; the name
is optional, so each entry is either CODE Role or CODE Name Role.
@Participants: CHI Target_Child, MOT Mother, FAT Father
@Participants: CHI Alex Target_Child, MOT Mary Mother
In the first line, Target_Child, Mother, and Father are roles,
not names. In the second line, Alex and Mary are optional names
sitting between the speaker code and the role.
Role labels use their canonical, case-sensitive spelling in both
@Participants and the role field of @ID. For example, Mother is a role;
Mom is not an alias, and Target_child is not Target_Child. E532 can offer
a likely correction, but validation and serialization preserve the original
spelling rather than silently rewriting it. Suggestions are heuristic advice,
not additional entries in the accepted vocabulary.
Speaker codes are short identifiers; the validator accepts up to
seven characters from A-Z, 0-9, _, -, and '. The convention
is three uppercase letters; the most common codes are:
CHI: target childMOT: motherFAT: fatherINV: investigatorOBS: observer
@ID
Provides detailed metadata for each participant. One @ID line per participant.
@ID: eng|corpus|CHI|2;6.||||Target_Child|||
Fields (pipe-separated): language, corpus, speaker code, age, sex, group, SES, participant role, education, custom field.
Age format: years;months.days (e.g., 2;6. = 2 years, 6 months).
SES field: ethnicity (White, Black, Asian, Latino, Pacific, Native, Multiple, Unknown), socioeconomic code (UC, MC, WC, LI), or combined with comma separator (e.g., White,MC).
Optional Headers
@Languages
Declares the language(s) used in the transcript.
@Languages: eng, fra
@Date
Recording date in DD-MON-YYYY format.
The day and year require exactly two and four ASCII digits respectively;
numeric signs are not allowed. @Birth of CODE uses the same format checks.
E518/E545 report malformed components without rewriting the source value.
@Date: 15-JAN-2024
@Location
Where the recording took place.
@Location: Boston, MA, USA
@Situation
Description of the recording context.
@Situation: free play with toys in lab
@Activities
Activities during the recording.
@Activities: toyplay, reading
@Comment
Free-form comments. Can appear anywhere in the file (before, between, or after utterances).
@Comment: child was tired during this session
@Media
Links the transcript to an audio or video file.
@Media: session01, audio
When the transcript name is known, Chatter checks it against the media name.
A name that differs in more than ASCII letter case is E531. A name that
differs only in letter case (Session.cha declaring @Media: session) is
W110: CLAN’s CHECK accepts it, but a case-sensitive filesystem will not find
the recording, so the names must be made identical. Different Unicode
spellings of the same name, such as
composed and decomposed accents, produce W109 normalization advice rather than
E531 filename mismatch. The warning identifies whether the media name, the
transcript name, or both need normalization; validation does not rename files
or rewrite the header. Identical decomposed spellings warn about both sides.
Disk-based commands use the name stored in the directory, not the spelling
typed at the command line. An unreadable or unusable stored name is reported,
not silently treated as an anonymous transcript.
To normalize only the @Media name, run
chatter fix --code W109 --apply <file>. This changes only the filename token
and never renames the transcript. A remaining file-name warning requires a
separate rename using a tool that preserves NFC (not Finder). Remote media URLs
are opaque and exempt from both the comparison and this repair.
File Names and @Media gathers these rules in one place: what must match, why a Mac hides the difference, and how to fix each diagnostic.
@Transcriber / @Coder
Identifies who created or coded the transcript.
@Transcriber: JDS
@Coder: ABC
Header Ordering
Headers should follow this conventional order:
@UTF8(required, first line)@Begin(required)@Languages@Participants(required)@IDlines (one per participant)- Other metadata headers (
@Date,@Location, etc.) @Commentlines (can also appear later)
Validation
The parser validates header structure including:
@UTF8must be the first non-empty line@Beginand@Endare required and must appear exactly once@Participantsis required and must declare all speakers used in utterances@IDparticipant codes must match@Participantsdeclarations- Age format validation in
@IDlines
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
File Names and @Media
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
A CHAT transcript names its recording twice: once in its own file name, and
once in the @Media header. A tool that needs the recording then looks for a
file with that name. All three spellings must be the same, character for
character, or the transcript finds its recording on some computers and not on
others.
The rule
For a transcript session01.cha:
@Media: session01, audio
- The
@Medianame equals the transcript’s file name without.cha. - The recording is that name plus its format’s extension, in lower case:
session01.mp3,session01.mp4orsession01.wav. - Names are spelled in Unicode NFC, the standard composed form, where an accented letter is one character rather than a letter plus a separate accent mark.
A remote URL in @Media is exempt: it points at media elsewhere, so there is
no local file name to compare.
Why “the same” means exactly the same
Filesystems disagree about when two names are the same file:
| Difference | macOS (default) | Linux |
|---|---|---|
Letter case: Session01 vs session01 | same file | different files |
Unicode form: composed vs decomposed ü | same file | different files |
A transcript whose names differ only in one of these ways works on a Mac and fails on Linux, including a Linux server that publishes recordings. Both kinds of difference are easy to introduce without noticing: letter case by hand, and decomposed accents by editors, operating systems and transfer tools that default to that form. The two spellings look identical on screen.
What chatter reports
When chatter knows the transcript’s file name (chatter validate on a file or
directory, to-json, the language server), it compares that name with the
@Media name:
| Code | Severity | When | Example |
|---|---|---|---|
| E531 | error | The names differ in more than ASCII letter case and Unicode form | session01.cha with @Media: session02 |
| W109 | warning | The same name, but one or both spellings are not NFC | Schlüssel.cha with a decomposed ü in @Media |
| W110 | warning | The same name except for ASCII letter case | Session01.cha with @Media: session01 |
E531 compares without regard to ASCII letter case, as CLAN’s CHECK does
(its error 157), so a case-only difference is not an E531 error. W110 reports
it instead, because the difference still breaks the lookup on a case-sensitive
filesystem. W109 and W110 are independent: one name can draw both. A
difference in the case of a non-ASCII letter (É vs é) is E531.
Chatter compares the transcript name as it is stored in the directory, not as it was typed on the command line.
Fixing each one
- E531: decide which name is right. Usually
@Mediais updated to the file name, which the suggestion in the diagnostic spells out. - W109:
chatter fix --code W109 --apply <file>rewrites the@Medianame in NFC and changes nothing else. A file name that is not NFC must be renamed with a tool that keeps NFC; Finder does not. - W110: make the spellings identical: update
@Media, or rename the transcript. There is no automatic fix, because chatter cannot see how the recording itself is spelled.
In every case, rename the recording to match as well. Chatter validates the
transcript, not the recording, so a recording spelled differently from its
@Media name is invisible to it. Tools that look the recording up report
that; batchalign3 matches recording names exactly on every host.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Utterances
Status: Reference Last updated: 2026-05-11 23:22 EDT
An utterance is the fundamental unit of a CHAT transcript. It consists of a main tier (the transcribed speech) followed by zero or more dependent tiers (annotations).
Main Tier
The main tier begins with *SPEAKER: followed by a tab and the utterance content, ending with a terminator.
*CHI: I want a cookie .
Speaker Codes
Speaker codes are short identifiers (up to seven characters from A-Z, 0-9, _, -, '; three uppercase letters is the convention) matching a code declared in @Participants:
@Participants: CHI Target_Child, MOT Mother
*MOT: what do you want ?
*CHI: cookie .
Terminators
Every utterance must end with a terminator:
| Terminator | Meaning |
|---|---|
. | Declarative (period) |
? | Question |
! | Exclamation |
+... | Trailing off |
+..? | Trailing-off question |
+/. | Interruption |
+//. | Self-interruption |
+/? | Interrupted question |
+!? | Broken question |
+"/. | Quotation follows on next line |
Line Continuation
Long utterances wrap to the next line with a leading tab:
*MOT: well I think that we should probably go to
the store and get some more cookies .
Content Items
The content between *SPEAKER: and the terminator consists of content items separated by whitespace:
- Words: regular words, potentially with annotations
- Groups: bracketed content like
<word word>for overlap, retrace, etc. - Special forms: pauses
(.), events&=laughs, fillers&-uh - Separators: commas
,and other punctuation
Words
Words are the primary content unit. See Word Syntax for full details.
Groups
Angle brackets < > group words for annotations:
*CHI: <I want> [/] I want cookie .
Common group annotations:
[/]: partial retrace (speaker repeats the same words)[//]: full retrace (speaker restarts with different words)[///]: multiple retracing (multiple false starts)[/-]: reformulation (speaker rephrases with different structure)[?]: uncertain transcription
Special Forms
*CHI: um (.) I want &-uh cookie .
(.): short pause(..): medium pause(...): long pause(1.5): timed pause in seconds&=laughs: paralinguistic event&-uh: filler
Media Linking
Utterances can include media timestamps (bullets) that link to audio/video:
*CHI: I want cookies . •1234_5678•
The numbers represent start and end times in milliseconds. The bullets
delimiting the pair render as • in most editors; on disk they are
the NAK control character (U+0015). See grammar/grammar.js rule
bullet.
Dependent Tiers
See Dependent Tiers for documentation on %mor, %gra, %pho, %wor, and other annotation tiers that follow the main tier.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Retraces and Repetitions
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
Retraces mark content that the speaker said but then corrected, repeated, or abandoned. They are one of the most consequential constructs in CHAT because they affect how every dependent tier aligns to the main tier.
CHAT Syntax
A retrace has two parts: the retraced content (what the speaker said first) and the correction (what follows). The retraced content is marked with a trailing bracket code:
| Marker | Name | Meaning |
|---|---|---|
[/] | Partial repetition | Speaker repeats the same words |
[//] | Full correction | Speaker restarts with different words |
[///] | Multiple correction | Multiple false starts |
[/-] | Reformulation | Speaker rephrases with different structure |
Single-Word Retraces
When only one word is retraced, no angle brackets are needed:
*CHI: I [/] I want that .
*CHI: ana [//] an .
*MOT: the book [/-] the magazine is here .
Group Retraces
When multiple words are retraced, angle brackets delimit the scope:
*MOT: <the dog> [//] the cat ran .
*CHI: <I want> [/] I need cookie .
*CHI: <I want the> [///] give me that .
Retraces with Replacements
A retraced word often has a replacement [: target] and/or error code
[* code]. This is common in aphasia and child language corpora where
the speaker produces an incorrect form:
*PAR: tika@u [: kitty] [* p:n] [//] kitty is nice .
%mor: noun|kitty aux|be-Fin-Ind-Pres-S3 adj|nice-S1 .
*PAR: lɛɾɪ@u [: later] [* p:n] [//] later in the day .
%mor: adv|late adp|in det|the-Def-Art noun|day .
*CHI: male [: female] [* s:r] [/] male [: female] [* s:r] .
%mor: adj|female-S1 .
In each case, the retraced word (before the [//] or [/]) is excluded
from %mor alignment. Only the correction (after the marker) is counted.
Data Model
Retraces are a first-class variant of UtteranceContent:
flowchart TD
UC["UtteranceContent"]
UC --> Word
UC --> RW["ReplacedWord"]
UC --> Retrace
UC --> AG["AnnotatedGroup"]
UC --> Other["...20 other variants"]
Retrace --> BC["BracketedContent"]
Retrace --> RK["RetraceKind"]
BC --> BIW["BracketedItem::Word"]
BC --> BIRW["BracketedItem::ReplacedWord"]
style Retrace fill:#f96,stroke:#333
The Retrace struct wraps the retraced content in a BracketedContent
container, which can hold any combination of words, replaced words, and
other content items:
// crates/talkbank-model/src/model/content/retrace.rs
pub struct Retrace {
pub content: BracketedContent, // the retraced words
pub kind: RetraceKind, // Partial, Full, Multiple, Reformulation
pub is_group: bool, // <word> [/] vs word [/]
pub span: Span,
}
Where the annotations live
There is deliberately no annotations field. A content item followed by a run
of scoped markers is a LEFT-ASSOCIATIVE CHAIN: each marker scopes over
everything to its left, so these two lines are different claims about the same
two words.
dog [* p:w] [/] dog the error is on the abandoned attempt
dog [/] [* p:w] dog the error is on the retrace
Annotations written BEFORE the marker therefore annotate the retraced material
and live inside content; annotations written AFTER it annotate the retrace and
live on an AnnotatedRetrace(Box<Annotated<Retrace>>) wrapper, exactly parallel
to Group / AnnotatedGroup.
A single flat annotations field holding both would make the two
indistinguishable: dog [* p:w] [/] would be silently written back as
dog [/] [* p:w], and a second adjacent marker would overwrite the first, so
на [//] [/] на would become на [/] на. Neither would be visible to
validate, to --roundtrip (which tests idempotence of serialize(parse(x)),
not fidelity to the input) or to SemanticEq (the two orderings would be the
same model).
Why First-Class?
Representing retraces as annotations on words or groups would mean every
match on content had to inspect annotation lists to determine whether a word
was retraced, a class of bugs where retraced content is accidentally included
in alignment counting, word extraction, or retokenization.
Making Retrace a top-level UtteranceContent variant means:
- The compiler enforces handling, WHERE the match is exhaustive. Every
matchonUtteranceContentmust have aRetracearm. The caveat is real: a site matching a retrace behind a_ =>arm compiles unchanged when a variant such asAnnotatedRetraceis added and silently answers wrong (at five sites it did, one of them the gate in front of all retrace validation). A first-class variant guarantees exhaustiveness only where the matches are already exhaustive, which is why this codebase bans catch-alls over content enums. - Domain-aware gating is centralized. The content walker checks the
Retracevariant once, not at every annotation-inspection site. - Alignment counting is simple. The count function returns
0forRetracein Mor domain, no annotation inspection needed.
Parser Conversion
The tree-sitter grammar parses retrace markers ([/], [//], etc.) as
annotations on word_with_optional_annotations. The Rust parser converts
them to structural Retrace nodes in parse_word_content():
flowchart LR
subgraph "Tree-sitter CST"
WOA["word_with_optional_annotations"]
SW["standalone_word"]
BA["base_annotations"]
RP["retrace_partial / retrace_complete / ..."]
WOA --> SW
WOA --> BA
BA --> RP
end
subgraph "Rust Model"
RET["UtteranceContent::Retrace"]
BC2["BracketedContent"]
W2["Word or ReplacedWord"]
RET --> BC2
BC2 --> W2
end
WOA -->|"parse_word_content()\n(word.rs)"| RET
Three cases in parse_word_content():
- Word + retrace (
I [/]), wrapWordinBracketedItem::WordinsideRetrace - Word + replacement + retrace (
tika@u [: kitty] [* p:n] [//]), buildReplacedWord, then wrap inBracketedItem::ReplacedWordinsideRetrace - Word + replacement, no retrace (
tika@u [: kitty]), emit bareReplacedWord
Group retraces (<content> [/]) are handled in group/parser.rs via
the same structural wrapping.
Alignment Behavior
Retraces interact differently with each dependent tier domain:
flowchart TD
RT["Retrace node\n(e.g. 'tika@u [: kitty] [* p:n] [//]')"]
RT -->|"Mor domain"| SKIP["SKIP\n(return 0)\nNot morphologically analyzed"]
RT -->|"Pho domain"| COUNT["COUNT\nPhonologically produced"]
RT -->|"Sin domain"| COUNT2["COUNT\nGesturally produced"]
RT -->|"Wor domain"| COUNT3["RECURSE\napply retrace-aware %wor leaf rule"]
style SKIP fill:#faa,stroke:#333
style COUNT fill:#afa,stroke:#333
style COUNT2 fill:#afa,stroke:#333
style COUNT3 fill:#afa,stroke:#333
Why %mor skips retraces: The %mor tier represents the morphological analysis of what the speaker meant to say. Retraced content is a false start or error; it was produced phonologically but is not part of the intended linguistic structure. The correction after the retrace marker carries the morphological analysis.
Why %pho/%sin/%wor include retraces: These tiers document what was actually produced, the sounds, gestures, and timing of the speech as it happened, including false starts. The retrace was physically spoken, so it appears in these tiers.
For %wor, retrace ancestry does not change leaf-level membership:
- spoken word tokens count both inside and outside retrace
- that includes fillers, fragments, nonwords, and untranscribed placeholders
- overlap annotations do not affect
%wormembership
Exact corpus-shaped contrast:
*CHI: <one &+ss> [/] one play ground .
%wor: one •321008_321148• ss •321148_321368• one •321809_321969• play •322049_322310• ground •322390_322890• .
*CHI: &+ih <the what> [/] what's letter &+th is this ?
%wor: ih •49063_49103• the •49103_49163• what •49183_50205• what's •50205_50405• letter •50405_50685• th •50886_50946• is •50946_51046• this •51086_51586• ?
Implementation
Both the counter and the walker ask one owner, alignment/helpers/descent.rs,
what a domain does with a retrace: %mor excludes it, every other domain
enters it. The counter takes a PositionalDomain and converts it for the
descent rule; the walker takes the TierDomain directly.
Counting and extraction: walk_alignable_item() in
alignment/helpers/count.rs, one walk whose sink either counts or collects:
UtteranceContent::Retrace(_) | UtteranceContent::AnnotatedRetrace(_) | /* other containers */ => {
match descend(item.structure(), Some(domain.into())) {
Descent::Into(entered) => walk_alignable_bracketed(entered.content(), domain, sink),
Descent::Atomic(unit) => sink(AlignablePosition::Atomic(unit)),
Descent::Excluded => {} // %mor: a retrace is not aligned
}
}
Walking: walk_words() in alignment/helpers/walk/mod.rs:
UtteranceContent::Retrace(_) | UtteranceContent::AnnotatedRetrace(_) | /* other containers */ => {
if let Some(into) = descend(item.structure(), domain).entered() {
walk_bracketed_content(&into.content().content, domain, f);
}
}
%wor generation and overlap counting still use dedicated recursive helpers,
but only for %wor-specific sequencing details like replacement handling, not
for retrace-sensitive membership.
Validation
The three retrace rules
validation/retrace/ runs three checks over ONE traversal of the tier
(visit.rs), which reaches every retrace including those nested inside another
retrace’s content.
| Code | Rule | Example rejected |
|---|---|---|
| E370 | A marker must be FOLLOWED by the repeated or corrected material. | <the> [/] . |
| E377 | A marker’s content may not be nothing but another marker. | на [//] [/] на |
| E378 | A marker’s content must contain a word, at any depth. | &=laughs [//] water |
E377 and E378 are disjoint despite the neighbouring names. In a [//] [/] a the
inner retrace still holds a word, so E378 stays silent; in <&=sigh> [/] &=sigh
there is no second marker, so E377 stays silent. The repairs differ too: drop a
marker for E377, retrace the words rather than the vocalization for E378.
Both were adjudicated with the CHAT maintainer on 2026-08-07. On adjacent markers: “clearly a mistake … It’s an error.” On a marker over an event: “No, not legal. You can’t retrace a laugh.” He gave the legal alternative half an hour earlier in the same thread, and it is why E378 tests for absent WORDS rather than for a present event:
*PAR: <the floor on the &=laughs water> [//] the floor on the xxx .
E378 recurses because 205 corpus retraces hold their words one level down, in an
annotated group or a quotation (<<the dog> [?]> [/] the dog), and a rule
testing the immediate children would reject every one of them. Untranscribed
material counts as words on purpose: xxx, yyy and www lower as words, so
retracing speech nobody could make out stays valid.
Which variants are containers is owned in one place for these rules,
model::content::structure::ContentStructure. It has one owner because
two hand-written copies of that knowledge can disagree, as about PhoGroup and
SinGroup, which silently stops E377 firing inside ‹...› with no test able
to see it.
Be precise about the scope: the retrace validators classify through it, and the alignment walkers do not.
They need to know WHICH container they are in, so a tier domain can skip a
phonological group but not a quotation, and Container deliberately does not
carry that. Settling those payloads is the prerequisite for migrating them.
Alignment Validation (E705)
E705 fires when the main tier has more alignable items than %mor. If
retraces are correctly parsed as Retrace nodes; they are excluded
from the count and E705 does not fire. If a retrace is accidentally
parsed as a bare ReplacedWord (the bug fixed in c90b9bf), it is
counted and triggers a false E705.
Regression Tests
tests/retrace_replaced_word_regression.rs contains 6 targeted tests:
| Test | Pattern | Verifies |
|---|---|---|
single_word_retrace_with_replacement_full | word [: repl] [* err] [//] | Retrace wraps ReplacedWord |
single_word_retrace_with_replacement_partial | word [: repl] [* err] [/] | Partial retrace with replacement |
single_word_retrace_with_replacement_multiple | word [: repl] [* err] [///] | Multiple retrace with replacement |
single_word_retrace_with_replacement_no_error_marker | word [: repl] [///] | No [*] still produces Retrace |
single_word_retrace_without_replacement | word [//] | Baseline (no replacement) |
retrace_with_replacement_does_not_cause_e705 | Full pipeline with %mor | No false E705 |
Reference corpus entries: corpus/reference/annotation/retrace.cha
See Also
- Alignment Architecture: full alignment system docs
- The %mor Tier: morphological tier format and alignment rules
- CHAT Manual: Retracing
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Replacements
Status: Current Last modified: 2026-05-29 17:47 EDT
A replacement is a CHAT annotation [: ...] that pairs a single
spoken word on the main tier with one or more “intended” words. It
records both what the speaker actually said and what the analysis should
treat the utterance as containing.
*CHI: wanna [: want to] go .
*CHI: dis [: this] is fun .
*CHI: rocking+house [: rocking+horse] [*] ?
This page is the canonical reference for what replacements mean in TalkBank, both as a CHAT-manual construct and as a typed AST in this repo. The most important load-bearing fact, which the rest of the page expands on:
Replacements are word-level, not group-level. Each tier domain chooses one side of the pair:
%moranalyzes the replacement (right side);%wor,%pho,%sinalign to the original (left side).%grafollows%mor.
CHAT Syntax
Word-Level Scope
A replacement attaches to a single standalone_word on the main tier
and contains one or more replacement words inside the brackets:
*CHI: gonna [: going to] eat lunch .
*CHI: dis [: this] toy .
*CHI: rocking+house [: rocking+horse] [*] ?
The grammar rules are
word_with_optional_annotations and replacement in
grammar/grammar.js
grep for the rule names rather than line numbers so this stays
accurate as the grammar evolves. Replacement words can be separated
by whitespace, so [: going to] is a single replacement of gonna
with two words.
There Is No Group-Level Replacement
<dat is> [: that is] is not valid CHAT. A replacement does not
attach to a group; it attaches to a single word. The grammar enforces
this by typing: ReplacedWord.word: Word, never Group. To replace
words inside a group, attach the replacement to the inner word:
*CHI: <dat [: that] is> [/] is broken .
This shape, replacement inside a group inside a retrace, is legal because each annotation operates at its own scope.
There Is No [::] Form
Some literature on CHILDES tooling references a [::] annotation; it
does not exist in this repo’s grammar, parser, or model, and is not
defined by the current CHAT manual. Only [:] exists. If you encounter
[::] in legacy data, treat it as a parse error to investigate, not a
construct to support.
The Per-Domain Alignment Rule
This is the rule contributors most often get wrong. Different tier domains align to different sides of a replacement pair:
| Tier | Side aligned to | Rationale |
|---|---|---|
%mor | replacement (right) | Morphosyntactic analysis annotates the target form, not the error |
%gra | replacement (right) | Grammatical relations align to %mor’s structure |
%wor | original (left) | Word-level timing is for what was actually spoken |
%pho | original (left) | Phonological transcription describes what was actually spoken |
%sin | original (left) | Spelling-in-actual describes the original surface form |
The mnemonic: the replacement encodes the intended form (what the
speaker meant or what a corrected transcript would read). Tiers
analyzing intent (%mor/%gra) use the replacement; tiers
documenting realization (%wor/%pho/%sin) use the original.
flowchart LR
spoken["Original word\n(left of [:)\n'dis'"]
target["Replacement words\n(inside [: ])\n'this'"]
spoken -->|"%wor (timing)"| wor["%wor: dis"]
spoken -->|"%pho (phonology)"| pho["%pho: dɪs"]
spoken -->|"%sin (spelling)"| sin["%sin: dis"]
target -->|"%mor (UD parse)"| mor["%mor: pron|this"]
target -->|"%gra (paired with %mor)"| gra["%gra: 1|0|ROOT"]
For multi-word replacements like gonna [: going to], the rule
generalizes consistently:
%wor/%pho/%sinproduce one entry, forgonna.%morproduces two entries, forgoingandto.%graproduces two entries, paired to the two%moritems.
The alignment-counting code that enforces this is in
alignment/units.rs
look for the UtteranceContent::ReplacedWord arm. The full table
of per-domain rules is in
spec/docs/ALIGNMENT_RULES.md.
Rust AST
A replacement is modeled as a first-class UtteranceContent variant,
not as a flag on Word:
// crates/talkbank-model/src/model/annotation/replacement.rs
pub struct ReplacedWord {
pub word: Word, // left side: original spoken word
pub replacement: Replacement, // right side: 1+ intended words
pub scoped_annotations: ReplacedWordAnnotations,
}
Two consequences of this shape:
- A replacement is a wrapper around a
Word, not a kind ofWord.ReplacedWordlives as its own variant ofUtteranceContent(andBracketedItem), holding an innerword: Wordplus the replacement payload. Contrast with retraces:Retraceis also a variant ofUtteranceContent/BracketedItem, but it wraps a group of content (a single word or a<...>group), not a singleWord. Different mechanism, different scope, same top-level slot in the AST. - The
walk_words()content walker yieldsWordItem::ReplacedWordas a distinct leaf (defined incrates/talkbank-model/src/alignment/helpers/walk/mod.rs). Domain-aware extraction code branches on this leaf type and chooses original or replacement per the table above.
Validation
Each Replacement Word Is Validated Like a Main-Tier Word
The replacement is a Vec<Word>. Each Word inside it goes through
the same validator that runs on main-tier words:
*CHI: dog [: C-3PO] .
This produces [E220] "C-3PO" is not a legal word in language(s) "eng": numeric digits not allowed, exactly as if C-3PO had appeared on the
main tier directly. The replacement does not provide an escape from
word-level validation. The implementation is in
replacement.rs.
This is critical for any code generating replacements programmatically:
do not assume [: ...] lets you smuggle arbitrary text past the word
validator. If your producer emits a replacement, both sides must be
CHAT-legal under the utterance’s declared language.
Replacement-Specific Error Codes
Three error codes are specific to replacements and do not apply to main-tier words:
| Code | Meaning |
|---|---|
E208 | Empty replacement [:] (no words provided between : and ]) |
E390 | Replacement contains an omission (0prefix form), disallowed inside replacements |
E391 | Replacement contains untranscribed material (xxx, yyy, www), disallowed inside replacements |
The principle: a replacement must be a concrete intended form. Empty, omitted, or unintelligible content defeats that purpose.
Interactions with Other Annotations
Replacements and Retraces Are Orthogonal
A retrace ([/], [//], [///], [/-]) and a replacement
([:]) are distinct annotations operating at different structural
levels:
- Retraces wrap content (a single word or a group). They are first-
class
UtteranceContentvariants and represent post-hoc speaker correction. - Replacements attach inside a
Wordslot viaReplacedWord. They are editorial metadata about an individual spoken word.
Both can coexist:
*CHI: <dat [: that] is> [/] is broken . (replacement inside retrace)
A retrace cannot live inside a replacement (the grammar wraps
replacements around standalone_word, not arbitrary content).
Replacements and Error Coding
Error codes follow the replacement and operate on the replaced word as a unit:
*CHI: rocking+house [: rocking+horse] [*] ?
Here [*] marks rocking+house as containing a phonological/lexical
error; the [: rocking+horse] records the intended form. The two
annotations cooperate: the replacement encodes what was meant, the
error code classifies how it deviates. Implementation:
scoped_annotations field on ReplacedWord.
Common Misconceptions
These are bugs we have repeatedly written down then forgotten, recording them here so future contributors don’t reinvent them.
- “
[: ...]lets me put any text I want.” No. Each replacement word is validated.[: C-3PO]fails E220 in English just asC-3POwould. - “
[:]is the right mechanism for ASR sanitization.” Usually no. ASR-introduced normalization typically wants[% ...](free- form comment) or[= ...](free-form explanation), neither of which validates word grammar. Use[:]only when you have a concrete CHAT-legal intended form. - “
%moranalyzes the original.” No.%moranalyzes the replacement. This is the correction’s morphology, not the error’s. - “
%worcount must equal%morcount.” No. Forgonna [: going to],%worhas 1 entry and%morhas 2. They align to different sides. The validator’s per-domain rule respects this. - “
<a b> [: c d]is a group-level replacement.” No. Group-level replacements don’t exist. Either replace inside (<a [: c] b [: d]>) or rephrase the transcription.
Source Citations
| Concern | File:line |
|---|---|
Grammar rule (replacement) | grammar/grammar.js:1341-1352 |
| Word-with-replacement rule | grammar/grammar.js:1063-1071 |
ReplacedWord struct | crates/talkbank-model/src/model/annotation/replacement.rs (search pub struct ReplacedWord) |
| Per-domain alignment | crates/talkbank-model/src/model/file/utterance/metadata/alignment/units.rs (search UtteranceContent::ReplacedWord) |
| Replacement validation | crates/talkbank-model/src/model/annotation/replacement.rs (search impl ... Validate for ReplacementWords) |
| Reference corpus example | corpus/reference/annotation/errors-and-replacements.cha |
| CHAT manual | https://talkbank.org/0info/manuals/CHAT.html#Replacement_Scope |
See Also
- Retraces and Repetitions: the orthogonal post-hoc correction mechanism.
- The %mor Tier: UD-syntax morphosyntactic analysis that aligns to the replacement form.
- Word Syntax: the word grammar that replacement words must satisfy.
- Dependent Tiers: overview of
%mor,%wor, etc., with their alignment relationships.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Untranscribed Markers: xxx, yyy, www
Status: Reference Last updated: 2026-06-14 19:57 EDT
CHAT reserves three short word-level markers for material the human
transcriber cannot or chose not to render as words on the main tier.
Each one has a specific meaning. Tools that emit CHAT, including ASR
pipelines, format converters, and editor heuristics, must respect
those meanings, because every downstream consumer (researchers,
validators, and aggregate-statistics tools like CLAN’s freq,
kideval, mlu) reads them at face value.
| Marker | Meaning | Emitter |
|---|---|---|
xxx | Transcriber listened to the audio and could not make out what was said. The speech is unintelligible to the human ear at this point. | Human transcriber only. |
yyy | Transcriber heard a discrete utterance but could not write it as ordinary CHAT words. Used when the surface form resists orthography (mumbled, slurred, foreign with no equivalent). The phonetic content typically appears on the %pho tier. | Human transcriber only. |
www | Transcriber chose not to transcribe this stretch, usually for privacy, off-topic content, or because the segment is irrelevant to the corpus’s purpose. | Human transcriber only. |
The shared property: each marker is the human transcriber telling later readers something specific about their experience listening to the audio. None of them mean “tooling could not process this token”.
Why this matters
When a researcher loads a CHAT corpus and counts xxx occurrences, the
result is a measure of human listening difficulty: it tells them how
much of the audio resisted human transcription. That number feeds into
methodology decisions (“can we get reliable MLU from this corpus?”,
“what’s the noise floor on this child’s speech?”, “should we re-record
in a quieter environment next time?”). It is a load-bearing signal in
language-development research.
If an ASR pipeline emits xxx whenever it can’t sanitize a token,
for example, substituting xxx for any word that fails CHAT
validation under a strict language profile, every xxx count in the
corpus becomes a meaningless mixture of “human couldn’t tell” and
“pipeline gave up”. Researchers then reading those counts are silently
misled. The signal is destroyed for the entire history of that
corpus, because the corruption is indistinguishable from real
unintelligibility once committed.
The same reasoning applies to yyy and www. A converter or
post-processor that emits any of these three markers because the
tooling couldn’t handle a token is committing semantic vandalism
against the whole field.
Rules for tooling
- Never emit
xxx,yyy, orwwwfrom a tool to mean “could not process”. These markers are reserved for human transcriber judgment. - When a token cannot be validated as legal CHAT under the
declared language, prefer one of:
- Pass the token through verbatim and let the CHAT validator
(or CLAN’s
check) flag it for human review. The transcriber listens, decides, and corrects. - Fail loud, abort the file rather than emit corrupted output.
- Apply only purely orthographic, semantically null repairs
(e.g., stripping a stray boundary quote mark from
"My). These are safe because no information is lost.
- Pass the token through verbatim and let the CHAT validator
(or CLAN’s
- Never sanitize a token by replacing it with one of the three markers. That is exactly the corrupting behavior this document prohibits.
- Never delete a token to “fix” a validation failure. Deletion loses data without any flag.
What tools synthesizing CHAT should do instead
Any tool that builds CHAT from an external source (ASR output, an importer, a format converter) should follow the same division of labor:
- Silently fix only orthographically inarguable problems (for
example, stripping a stray boundary quote mark from
"My). - For tokens that fail language-level validation but are
structurally legal CHAT (e.g.,
C-3POunder English: tree-sitter accepts the digit-hyphen compound butWord::validatefires E220 “numeric digits not allowed”), ship the token verbatim. The full-file validator andcheckfire E220 on the same word, the file ends up in the human review queue, and the transcriber listens to the audio and decides what was actually said. - For tokens that fail structural parsing (tree-sitter rejects), fail loud: emitting malformed CHAT would corrupt the file beyond the validator’s ability to flag it.
The division of labor is: the tool fixes only what is mechanically unambiguous; CHECK and the human transcriber handle everything that requires judgment about what the speaker said.
Related rules
xxx/yyy/wwwsurvive the transcript through all NLP passes (morphotag, utseg, translate, coref) without re-interpretation. Tools that walk the AST treat them as opaque tokens; they have no POS tag, no lemma, no dependency parent, no translation.%worexcludes all three (no phoneme sequence to align).%phomay referenceyyydirectly because the phonetic content is the whole point of the marker.- See
word-syntax.mdfor grammar; this document is the policy reference for who is allowed to emit them and why.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Postcodes ([+ ...])
Status: Reference Last updated: 2026-06-25 07:30 EDT
A postcode is a tagged annotation token that attaches to an
utterance as a whole and appears after the terminator. The
canonical CHAT syntax is [+ <text>]. Postcodes carry researcher /
analysis tags about the utterance, whether it should be excluded
from analysis, how it should be coded, what kind of speech act it
represents, without modifying the utterance’s word content.
Syntax and Scope
*CHI: I want cookie . [+ exc]
*MOT: what did you say ? [+ imp]
*CHI: no I don't want it ! [+ neg] [+ trn]
Three structural facts to internalize:
- Postcodes attach to the utterance, not to a word. They sit
after the terminator, on the main tier, alongside (but distinct
from) any utterance-level bullet. Unlike word-scoped annotations
(
[: ...]replacement,[% ...]comment,[= ...]explanation,[* ...]error code), a postcode does not modify the interpretation of any single word, it tags the whole utterance. - Multiple postcodes may follow a single terminator. They are ordered, but the order is not semantically privileged.
- The body is free-form text. The CHAT word grammar is not applied to postcode contents. Researchers can write arbitrary tags, codes, descriptions, comments, or analytic notes. The model stores the raw text and leaves interpretation to downstream tooling and conventions.
Common Postcodes, Empirical Survey
The postcode vocabulary is open-ended: the CHAT format imposes no
closed set, and an audit of every [+ ...] token across a
JSON-mirrored snapshot of the TalkBank corpora (~99k files, 23+
data-repo families) found 488 distinct values in active use.
The findings split into three tiers ranked by repo spread (in how many distinct corpus families the code appears), the more useful ranking than raw count, because high-count codes can be concentrated in a single corpus.
Tier 1, Cross-corpus codes (in 7+ repos)
These are the conventions every CHAT consumer should expect to encounter across collections:
| Postcode | Repo spread | Total occurrences | Meaning |
|---|---|---|---|
[+ gram] | 13 | ~3,100 | Grammatical, utterance is grammatically well-formed for purposes of the analysis. |
[+ exc] | 9 | ~26,900 | Exclude utterance from analysis. The utterance is preserved in the transcript but tagged so analytic tools (CLAN’s freq, mlu, etc.) skip it. |
[+ bch] | 9 | ~10,000 | Backchannel, listener-side acknowledgement (mhm, yeah) that should not be counted as a substantive turn. |
[+ trn] | 7 | ~3,800 | Translation utterance. |
Tier 2, Multi-corpus protocol codes (in 4-6 repos)
Codes deployed across several CHILDES sub-collections, typically encoding picture-narration / story-reading / imitation experimental conditions. Substantial raw counts (often tens of thousands), but their meaning is set by the originating protocol, consult per-corpus documentation rather than assuming a global definition:
| Postcode | Repo spread | Total occurrences |
|---|---|---|
[+ SR] | 5 | ~31,000 |
[+ IN] | 5 | ~24,500 |
[+ PI] | 5 | ~22,700 |
[+ R] | 4 | ~16,200 |
[+ I] | 4 | ~10,500 |
[+ nv] | 4 | ~3,300 |
[+ imit] | 4 | ~3,200 |
Tier 3, Single-corpus and long-tail codes
About 80% of the 488 distinct values appear in one repo only. The
single-corpus codes include high-volume protocol vocabularies (e.g.
[+ uncued] ~19,500 in one repo, [+ NAC] ~3,500 in one repo,
[+ diary] ~2,800 in a Romance/Germanic diary-study collection,
[+ noatt] ~2,300 in one repo, [+ inter-utter-switch] ~720
flagging code-switching turns).
The long tail also includes researcher-private notes, typos that
survived check, and per-study coding schemes. Tooling MUST treat
any unknown postcode value as opaque text, the corpus author may
know what it means, the format does not.
Caveats
- Numbers are from a snapshot audit and will drift as corpora are added or revised. Treat the broad shape (open vocabulary, ~4 truly cross-corpus codes, ~10 multi-corpus protocol codes, ~hundreds of single-corpus or long-tail codes) as the load-bearing finding, not the exact counts.
- “Repo spread” counts data-repo families, not individual files. Two corpora curated by the same group inside one data-repo count as one for spread; researchers using the same code in two different family-of-corpora packages count as two.
- The CHAT manual remains the source of truth for standard conventions. The empirical survey above shows what is actually deployed; when ingesting a new corpus, consult its own documentation for the postcodes in use.
What Postcodes Are NOT
Postcodes are easy to confuse with several other CHAT annotation forms because they all use square brackets. The differences are substantive and load-bearing.
| Form | Scope | Body validation | Purpose |
|---|---|---|---|
[+ ...] | Utterance-level (this doc) | None, free text | Researcher / analysis tag attached to the whole utterance |
[: ...] | Word-level | Replacement words ARE validated as CHAT words | Sanctioned-form correction of the preceding word (see replacements.md) |
[% ...] | Word-level | None, free text | Free-form comment about the preceding word or local span |
[= ...] | Word-level | None, free text | Explanation of unclear / non-standard speech (often paired with xxx / yyy placeholders) |
[* ...] | Word-level | None, error code text | Error coding for the preceding word, optionally with a structured code |
Two consequences worth pinning down explicitly:
- A postcode cannot carry per-word semantics. If you want to attach a comment, replacement, or error code to a single word, use the appropriate word-scoped form. Stretching a postcode to mean “this word is X” loses the per-word position downstream tools depend on.
- A word-scoped annotation cannot tag an utterance. If you want
to mark an entire utterance for exclusion or translation, use a
postcode. A
[% exclude this]after a word does not mean “exclude the utterance” to any consumer.
Not Postcodes: Quotation Markers
Quotation marking in CHAT is not a postcode form. The constructs
+"/. (quotation end), +"/, and +" (quotation linkers /
continuations) are tier-level terminators and linkers, not
[+ ...] postcodes, the grammar rule postcode in
grammar/grammar.js is strictly [+ <text>], and the quotation
forms live under separate grammar rules (quoted_new_line,
linker_quotation_follows).
See Utterances → Terminators for the
syntactic forms, and the
talkbank-model::validation::cross_utterance validator family
(gated by ValidationContext::enable_quotation_validation) for the
cross-utterance balance checks.
A walker in talkbank-model::validation::utterance::quotation
(check_quotation_balance) does scan the postcode list for text
"/ and "/., but a sweep over the data-json corpus mirror
(101,414 files, 2026-05-11) returned zero such postcodes, that
code path is effectively dead, retained presumably as defence
against hand-edited oddities. The real quotation-balance work
happens in the cross-utterance family above.
Position in the AST
An utterance’s main tier is MainTier, whose content: TierContent
field carries the actual tier payload, including postcodes, as a
typed list:
pub struct MainTier {
pub speaker: SpeakerCode,
pub content: TierContent,
// spans omitted for brevity
}
pub struct TierContent {
pub linkers: TierLinkers, // utterance-leading +<, ++, etc.
pub language_code: Option<LanguageCode>, // [- code]
pub content: TierContentItems, // word-level items (newtype over Vec<UtteranceContent>)
pub terminator: Option<Terminator>, // ., ?, !, +..., etc.
pub postcodes: TierPostcodes, // [+ ...] tokens after the terminator
pub bullet: Option<Bullet>, // optional terminal media bullet
// content_span omitted for brevity
}
(See talkbank-model/src/model/content/main_tier.rs and
tier_content.rs for the exact shape. The postcodes slot lives on
TierContent, which is the main-tier payload. Dependent tiers do
not use TierContent: each has its own type (for example %com is a
text tier, %wor is a list of timed items), and none carries a
postcode slot. So a [+ ...]-shaped token on a dependent tier is
parsed as ordinary tier content, never a Postcode. This is why
chatter does not, and structurally cannot without banned raw-text
scanning, reproduce CLAN CHECK 109 (“postcodes are not allowed on
dependent tiers”); that deliberate divergence is recorded in the
CHECK Parity Audit.)
Because postcodes live at the utterance level, the per-word
traversal helpers (walk_words, walk_words_mut) do not visit
them. Code that needs to read or rewrite postcodes accesses the
list directly.
The model stores postcode text as SmolStr and preserves it
verbatim through CHAT roundtrips. Downstream tooling, including
CLAN command implementations such as freq, mlu, kideval, is
responsible for interpreting individual postcode values per its own
conventions.
Tooling Rules
Tools that emit or consume CHAT must respect the scope distinction.
- Emitters: when adding a researcher tag to an utterance, attach
a
Postcodeto the utterance’sMainTierContent, not aContentAnnotationto a word. Both serialize, but only the former reaches downstream consumers as utterance-level metadata. - Consumers: when reading utterance-level tags (e.g.,
implementing an “exclude” filter), iterate
main.content.postcodeson each utterance, not the word-level annotations inUtteranceContent. The two lists are populated by different parser branches and have different semantics. - Round-trip preservers (extract→modify→inject pipelines such as
the NLP injection passes in
crates/batchalign-*): preserve the postcode list unchanged. None of the standard NLP passes have a reason to add, remove, or reorder postcodes.
References
- CHAT manual: Postcodes
- CHAT manual: Excluded Utterance Postcode
- CHAT manual: Included Utterance Postcode
- Model:
talkbank-model/src/model/content/postcode.rs - Quotation validator:
talkbank-model/src/validation/utterance/quotation.rs
This page last changed: 2026-07-07 (commit ca8c5b8e). The whole book last changed: 2026-10-07 (commit 5e895791).
Dependent Tiers
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
Dependent tiers appear on lines beginning with % immediately after an utterance. They provide annotations linked to the main tier content.
CHAT defines four structural categories of dependent tiers:
- Structured linguistic tiers: parsed into typed AST nodes with word-level alignment
- Phon phonological tiers: syllabification and segmental alignment from the Phon project
- Bullet-content tiers: free-form text with optional inline timing markers
- Text tiers: plain text with no structural alignment
Structured Linguistic Tiers
These tiers have rich, parsed representations in the data model. Each token aligns 1-to-1 with an alignable word on the main tier (excluding retraces, pauses, and events). Terminators (., ?, !) must match the main tier terminator.
%mor, Morphological Analysis
The %mor tier carries part-of-speech tags, lemmas, and morphological features for each word on the main tier. See The %mor Tier for full documentation covering the UD-style format, data model, divergences from Universal Dependencies, and migration from traditional CHAT MOR.
Format: POS|lemma[-Feature]*, with ~ separating post-clitics.
*CHI: she's eating cookies .
%mor: PRON|she~AUX|be-Pres-S3 VERB|eat-Prog NOUN|cookie-Plur .
%gra, Grammatical Relations
The %gra tier encodes dependency syntax using Universal Dependencies relation labels. Each entry has the format index|head|relation, where indices are 1-based and head 0 indicates ROOT.
*CHI: I want cookies .
%mor: PRON|I VERB|want NOUN|cookie-Plur .
%gra: 1|2|NSUBJ 2|0|ROOT 3|2|OBJ 4|2|PUNCT
The %gra tier aligns with %mor chunks (clitics expand into multiple chunks). Validation checks sequential indices (E721), ROOT structure (E722 missing root, E723 multiple roots), and circular dependencies (E724). Two of those describe the tier AS A WHOLE and are withheld when the tier in hand is not the one you wrote: E721 and E722, when the parser had to reject a relation it could not represent, or when %mor-to-%gra alignment has already reported a count or index fault. E723 and E724 are always reported, because dropping a relation cannot create a second root or close a cycle, so a violation among the relations that survive is one the transcript contains.
%pho / %mod, Phonological Transcription
The %pho tier records actual pronunciation; %mod records target/model pronunciation. Both use the same format: space-separated phonetic tokens aligned 1-to-1 with main tier words.
*CHI: I want three cookies .
%pho: aɪ wɑnt fwi kʊkiz .
%mod: aɪ wɑnt θri kʊkiz .
Phonological tiers support IPA, UNIBET, X-SAMPA, or custom notation systems. They are used for child language, speech disorders, L2 learning, and dialectal variation studies.
Parsing strategy: We deliberately parse only the minimal word/group-level structure in
%phoand%modneeded for coarse alignment with the main tier. The full IPA phoneme content is stored as opaque strings, deep phonological analysis is handled by Phon, and we avoid duplicating that work. The Phon extension tiers (%modsyl,%phosyl,%phoaln) follow the same strategy.
%sin, Gesture and Sign Annotation
The %sin tier codes gestures and signs aligned with speech. Each token is either 0 (no gesture) or g:referent:type (e.g., g:ball:dpoint for a deictic point at a ball).
*CHI: that ball .
%sin: g:ball:dpoint 0 .
Multiple simultaneous gestures use bracket grouping: 〔g:toy:hold g:toy:shake〕.
%wor, Word Timing
The %wor tier carries word-level timing annotations for media synchronization.
Words may include inline bullets with millisecond timestamps. Timing data comes
from the bullet fields. Word text is not lexical authority, but consumers may
compare it with chatter’s canonical generated display sequence to refuse stale
same-count timing reuse.
⚠ IMPORTANT:
%worword text is the cleaned form, by design. When chatter serializes a%worword it writes the word’s cleaned text, the spoken form with surface markers removed, NOT the raw main-tier surface form. This is a deliberate convention (seeWorTier::write_chatincrates/talkbank-model/src/model/dependent_tier/wor.rs), chosen for human readability and because%worexists to anchor timing, not to re-state the main tier’s orthography. The generated%wortext and the TextGrid export both use this cleaned form.Consequence you must know: surface markers carried on a word, prosodic lengthening (
wabe:), and similar in-word notation, are not preserved in%woroutput. A main-tier wordwabe:becomeswabeon%wor. This means a%worline containing such words does not byte-roundtrip (parse, serialize, reparse changes the surface text), and that is expected, not a bug.%woris a cleaned, timing-only view; the main tier remains the faithful record of surface forms. Do not “fix” the%worserializer to emit raw text without an explicit decision to change this convention.
%wor is not a flat “all tokens except punctuation” tier. It follows a
word-level alignment rule:
- Regular words count.
- Fillers (
&-um,&-uh,&-you_know) count; they are real spoken words with known phoneme sequences. - Fragments (
&+...) do NOT count: incomplete phoneme sequences; the FA engine cannot reliably anchor partial phonological material. - Nonwords (
&~...) do NOT count: interactional/gestural sounds without stable lexical phoneme content for alignment. - Untranscribed placeholders (
xxx,yyy,www) do NOT count: they have no known phoneme sequence; CTC forced alignment cannot produce timings for unknown material. - Replacements keep the original spoken word slot for
%wor; the replacement text matters for%mor, not%wor. If the original slot is untranscribed or a fragment/nonword, it is still excluded. - Retrace scope does not change
%wormembership. - Overlap markers do not change
%wormembership.
%wor is a timing-annotation tier. Its word count equals the number of Wor-domain
words and may differ from a naive main-tier word count. CHAT validation does not
require the %wor count to match the current main tier. Timing consumers first
require equal counts and then require every %wor display token to match the
canonical token derived from its main-tier position. A mismatch refuses timing
reuse without making the legacy CHAT file invalid.
*CHI: I want cookies .
%wor: I want cookies .
Exact corpus-shaped contrast:
*CHI: <one &+ss> [/] one play ground .
%wor: one •321809_321969• play •322049_322310• ground •322390_322890• .
# &+ss is a fragment, excluded from %wor regardless of retrace context.
*EXP: &+ih <the what> [/] what's letter &+th is this ?
%wor: the •49103_49163• what •49183_50205• what's •50205_50405• letter •50405_50685• is •50946_51046• this •51086_51586• ?
# Fragments &+ih and &+th excluded; regular words remain.
*EXP: what's is dis [: this] ?
%wor: what's •37050_37471• is •37491_37631• dis •37631_38131• ?
*CHI: xxx snack .
%wor: snack •884668_885168• .
# xxx has no phoneme sequence, excluded from %wor; only snack appears.
*CHI: &~um a boat .
%wor: a •1073779_1073799• boat •1076861_1077361• .
# &~um is a nonword, excluded from %wor.
*CHI: &-mm [<] bananas are good .
%wor: mm •1949506_1949566• bananas •1949566_1949766• are •1949846_1949987• good •1950067_1950567• .
# &-mm is a filler, included in %wor (real spoken word with alignable phoneme sequence).
flowchart TD
A["Main-tier word candidate"] --> B{"Timestamp token /\nomission / empty?"}
B -->|Yes| OUT["Excluded from %wor"]
B -->|No| C{"Untranscribed?\n(xxx/yyy/www)"}
C -->|Yes| OUT
C -->|No| D{"Fragment or nonword?\n(&+ or &~)"}
D -->|Yes| OUT
D -->|No| IN["Counts for %wor\n(word or filler &-)"]
style IN fill:#afa,stroke:#333
style OUT fill:#faa,stroke:#333
Phon Phonological Tiers
These tiers originate from the Phon
project and provide syllable-annotated phonological transcription and segmental
alignment. Older exports serialize them as %x-prefixed user-defined tiers
(%xmodsyl, %xphosyl, %xphoaln); current Phon exports use the official
unprefixed tier names (see Phon Tiers). Phon stores phonological data in its own XML format. As of Phon
4.0.0-beta.9 (2026-06-25), Phon reads and writes CHAT natively.
%modsyl / %phosyl, Syllabified Phonology
%modsyl is a syllabified version of %mod (target pronunciation); %phosyl
is a syllabified version of %pho (actual pronunciation). Each phoneme is
annotated with a syllable position code (N=nucleus, O=onset, C=coda,
etc.). Words are space-separated and align 1-to-1 with the corresponding
%mod or %pho tier.
*CHI: the best .
%mod: ðə bɛst .
%modsyl: ð:Oə:N b:Oɛ:Ns:Ct:C .
%pho: ðə bɛs .
%phosyl: ð:Oə:N b:Oɛ:Ns:C .
Alignment: Content-based, stripping position codes (:N, :O, :C, etc.)
and stress markers (ˈ, ˌ) from %modsyl should yield the same phonemes
as %mod. Same for %phosyl → %pho.
%phoaln, Phone Alignment
%phoaln provides segmental alignment between target and actual IPA,
showing phoneme-by-phoneme correspondence. Each pair uses source↔target
notation; ∅ marks insertions or deletions.
*CHI: the best .
%phoaln: ð↔ð,ə↔ə b↔b,ɛ↔ɛ,s↔s,t↔∅
Alignment: Positional, word-by-word, word N in %phoaln aligns with
word N in both %mod and %pho.
Parsing strategy: Same as %pho/%mod, we parse just enough structure
for alignment (word boundaries for %modsyl/%phosyl, alignment pairs for
%phoaln). IPA phoneme content is treated as opaque strings.
Validation (E725-E728)
Because these are derived views, word counts must match between each syllabification tier and its parent IPA tier:
| Check | Error code |
|---|---|
%modsyl word count ≠ %mod word count | E725 |
%phosyl word count ≠ %pho word count | E726 |
%phoaln word count ≠ %mod word count | E727 |
%phoaln word count ≠ %pho word count | E728 |
These checks are gated on ParseHealth, if either tier in a pair has parse
errors, the alignment check is suppressed to avoid false positives.
Known Export Word-Count Mismatch (Existing Corpus Files)
A subset of existing corpus CHAT files map %mod/%pho to orthography words
one to one, silently dropping extras, while their syllabification tiers
(%modsyl, %phosyl, %phoaln) carry the full IPA word set, undropped. In
child phonology data where children produce more IPA words than orthographic
targets (~4% of Phon corpus files), this creates tier-to-tier word count
mismatches. The mismatches originate in the Phon XML source data
(orthography vs. IPA word count discrepancies) and are handled inconsistently
across tiers in the resulting CHAT file.
As of Phon 4.0.0-beta.9 (2026-06-25), Phon reads and writes CHAT natively; we have not yet seen output from that native export and do not know whether it reproduces this inconsistency. Files already affected remain in the corpora and must continue to validate against the behavior described above.
Bullet-Content Tiers
These tiers contain free-form text with optional embedded timing markers (•START_END•) and picture references (•%pic:"file.jpg"•). They do not align word-by-word with the main tier.
| Tier | Purpose |
|---|---|
%act | Physical actions, gestures, non-verbal behaviors |
%cod | Research-specific coding (semantic roles, thematic coding, error classification) |
%com | Comments, annotations, and contextual notes |
%exp | Explanations or expansions of ambiguous/incomplete speech |
%add | Addressee identification in multi-party conversations |
%spa | Speech act coding (request, assertion, question, directive) |
%sit | Situational context or setting description |
%gpx | Extended gesture position coding |
%int | Intonational contours and prosodic patterns |
%cod is bullet-content in the shared TalkBank AST. In the %cod coding
convention, a word selector such as <w4> scopes the code that follows it
(it names which main-tier word the code applies to) rather than being a code
in its own right.
Example with timing:
*CHI: gimme that .
%act: reaches toward shelf
%com: child is pointing to picture
Text Tiers
These tiers contain plain text with no bullets, timing, or structural alignment:
| Tier | Purpose |
|---|---|
%alt | Alternative transcriptions |
%coh | Cohesion annotation |
%def | Definitions |
%eng | English translations (for non-English transcripts) |
%err | Error annotations |
%fac | Facial expressions |
%flo | Flow annotation |
%gls | Glosses |
%ort | Orthographic representations |
%par | Paralinguistic information |
%tim | Timing information |
User-Defined Tiers
Tiers prefixed with %x (e.g., %xcod, %xact) are user-defined dependent tiers. They are preserved during parsing and roundtrip but receive no structural validation beyond basic format checks. Any %x-prefixed tier is always accepted, this is the open extension point for project-specific annotation.
The Supported Set Is Closed
A dependent tier is valid in chatter only if it is one of the standard tiers documented above (the structured, Phon, bullet-content, and text tiers) or a %x-prefixed user-defined tier. Any other %-tier is invalid CHAT, and chatter rejects the file with error E605 (UnsupportedDependentTier). This is a closed set by design: chatter validate is the binding judgment on CHAT validity, so an unrecognized dependent tier is an error, not a warning.
Deliberate Divergence from CLAN: Retired Legacy Tiers
When TalkBank standardized morphology on a single Universal Dependencies %mor tier (plus %gra for relations), several legacy dependent tiers were retired. CLAN’s check still accepts three of them, so on these chatter is intentionally stricter, a deliberate, documented divergence:
| Retired tier | CLAN check | chatter |
|---|---|---|
%trn | accepts | rejects (E605) |
%tra | accepts | rejects (E605) |
%grt | accepts | rejects (E605) |
%umor | rejects | rejects (E605) |
The modern UD-%mor workflow has one morphology tier (%mor) plus %gra; the older training/translation/variant tiers are no longer part of the format chatter validates. %umor is rejected by both validators and is listed only for completeness. Note that %xtra (with the %x prefix) is a perfectly valid user-defined tier; only the bare %tra is retired.
This is one instance of a general principle: where chatter intentionally departs from CLAN/CHECK behavior, the divergence is documented rather than left implicit. See CHECK Parity Audit.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
The %mor Tier: Morphological Analysis
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
The %mor (morphological) dependent tier provides word-by-word morphosyntactic annotation aligned with the main tier. Each main-tier word receives a morphological code specifying part of speech, lemma, and grammatical features.
Format Overview
*CHI: I want cookies .
%mor: pron|I-Prs-Nom-S1 verb|want-Fin-Ind-Pres-S1 noun|cookie-Plur .
Each %mor item has the structure POS|lemma[-Feature]*, where:
- POS: part-of-speech category (
noun,verb,pron,det,aux, etc.) |: pipe separator (always present)- Lemma: base form of the word (
cookie,be,I). May contain language-specific compound or derivational boundary markers (see Compound Lemma Boundaries below) - Features: zero or more morphological features, each preceded by
-(-Plur,-Fin-Ind-Pres-S3)
Items are space-separated and terminate with a punctuation marker (., ?, !, etc.).
The UD MOR Format
TalkBank’s %mor tier uses a format inspired by Universal Dependencies (UD) but adapted to CHAT conventions. We call this the UD MOR format to distinguish it from the older CLAN-era MOR format.
The UD MOR format comes from batchalign’s Stanza-based morphosyntax pipeline. Stanza produces standard UD analysis (UPOS, lemma, morphological features, dependency relations), and the Rust mapping layer converts this to CHAT %mor and %gra tiers. It is the format used for all new corpus annotation.
Structure: Flat POS|lemma[-Feature]*
Every morphological word is flat, a single POS tag, a single lemma, and a linear chain of features:
POS|lemma[-Feature1][-Feature2][-Feature3]...
There are no compounds, prefixes, subcategories, or nested structures in the UD MOR format. The entire morphological analysis of a word is captured by the POS+lemma+features triple.
Examples:
| Word | %mor code | POS | Lemma | Features |
|---|---|---|---|---|
| dog | noun|dog | noun | dog | (none) |
| dogs | noun|dog-Plur | noun | dog | Plur |
| running | verb|run-Part-Pres-S | verb | run | Part, Pres, S |
| is | aux|be-Fin-Ind-Pres-S3 | aux | be | Fin, Ind, Pres, S3 |
| I | pron|I-Prs-Nom-S1 | pron | I | Prs, Nom, S1 |
| the | det|the-Def-Art | det | the | Def, Art |
Multi-Word Tokens (Clitics)
English contractions and similar multi-word tokens (MWTs) are represented using the tilde (~) separator for post-clitics:
*CHI: it's red .
%mor: pron|it~aux|be-Fin-Ind-Pres-S3 adj|red .
Here it's is a single main-tier word that expands to two morphological words: pron|it (main) and aux|be-Fin-Ind-Pres-S3 (post-clitic). The ~ indicates the two MOR words are fused into one orthographic token.
Each clitic counts as its own chunk for %gra alignment; pron|it~aux|be-Fin-Ind-Pres-S3 produces 2 chunks, each needing its own grammatical relation.
Terminator
The %mor tier ends with a terminator that matches the main tier’s utterance terminator:
*CHI: what is that ?
%mor: pron|what aux|be-Fin-Ind-Pres-S3 det|that ?
The terminator (., ?, !, +..., etc.) counts as one chunk for %gra alignment.
How It Diverges from UD
The UD MOR format is UD-inspired but not UD-compliant. Several deliberate adaptations make it fit CHAT conventions while preserving most UD information. This section catalogs every divergence.
1. POS Tags Are Lowercased UPOS
UD uses uppercase UPOS tags (NOUN, VERB, PRON). CHAT uses lowercase (noun, verb, pron). This is a lossless, trivially reversible surface change.
| UD UPOS | CHAT POS |
|---|---|
| NOUN | noun |
| VERB | verb |
| AUX | aux |
| PRON | pron |
| DET | det |
| ADJ | adj |
| ADV | adv |
| ADP | adp |
| PROPN | propn |
| INTJ | intj |
| CCONJ | cconj |
| SCONJ | sconj |
| NUM | num |
| PART | part |
| X | x |
2. Feature Values Are Flat, Not Key=Value (Currently)
UD represents morphological features as key=value pairs: Number=Plur, Tense=Past, Person=3. The current CHAT convention drops the keys and uses only the values: -Plur, -Past, -S3.
This is the most significant divergence from UD, because:
- Information loss:
Plurcould in principle beNumber=PlurorDegree=Plur(though in practice the UD feature value set has no real ambiguities). - Collapsed person/number: UD
Person=3|Number=Singbecomes-S3, a combined code that cannot be mechanically decomposed back to its UD components. - Feature ordering: Features appear in a conventional order determined by the generation pipeline, not in UD’s alphabetical order.
The data model supports key=value features. The MorFeature type has an optional key field, when present, the feature serializes as Key=Value (e.g., -Number=Plur); when absent, it serializes as just the value (e.g., -Plur). This is forward-compatible: existing flat features parse and serialize identically, and if batchalign’s mapper begins emitting Key=Value features, they flow through the parser and model without any format changes.
3. Multi-Value Features: Commas Preserved
UD encodes multi-value features with commas: PronType=Int,Rel (the word is both interrogative and relative). In CHAT %mor, the comma is preserved within the feature value:
-Int,Rel
This is treated as a single feature value "Int,Rel". The grammar accepts commas within feature values, and the model stores them as-is. No decomposition occurs; the model faithfully records the string that appears in the %mor tier.
Existing corpus data using a concatenated form (PronType=Int,Rel written as -IntRel) also parses correctly; it is simply treated as the flat value "IntRel".
4. Dependency Relations Are Uppercase with Dash Subtypes
The %gra tier (not %mor, but closely related) uses uppercase relation names with dashes for subtypes, where UD uses lowercase with colons:
| UD | CHAT %gra |
|---|---|
nsubj | NSUBJ |
acl:relcl | ACL-RELCL |
obl:tmod | OBL-TMOD |
This is lossless; case and separator are trivially reversible.
5. ROOT Head Convention
In UD, the root word has head=0. In %gra, two conventions coexist:
- UD convention:
head=0(e.g.,3|0|ROOT), the standard emitted - Legacy TalkBank convention:
head=self(e.g.,3|3|ROOT), found in older corpus data
The parser and validator accept both forms. New output uses head=0.
6. No XPOS, No DEPREL Subtypes in %mor
UD provides both UPOS (universal POS) and XPOS (language-specific POS). CHAT %mor uses only UPOS-equivalent tags; there is no XPOS field. Language-specific POS distinctions are not represented.
Similarly, UD’s fine-grained dependency relation subtypes (e.g., nsubj:pass) appear in %gra as NSUBJ-PASS, but the %mor tier itself contains no dependency information.
7. No Morpheme Segmentation
Traditional CHAT MOR formats (CLAN-era) supported morpheme-level segmentation with compound markers (+), prefix markers (#), and suffix chains (-SUFFIX&type). The UD MOR format does not use any of these, each word is analyzed as a flat POS+lemma+features triple.
The grammar still accepts some of these legacy markers for backward compatibility with older corpus data, but the canonical UD MOR format does not produce them.
Compound Lemma Boundaries
Several UD treebanks use special characters inside lemmas to mark morphological boundaries. These are meaningful linguistic annotations preserved in the CHAT %mor lemma field when possible.
Known Markers Across Languages
| Language | Marker | Meaning | Example Lemma | In %mor |
|---|---|---|---|---|
| Estonian | = | Compound boundary | maja=uks (house-door) | noun|maja=uks, preserved |
| Basque | ! | Derivational boundary | partxi!se (share + derivation) | noun|partxi!se-Ine, preserved |
| Finnish | # | Compound boundary | jää#kaappi (ice-cabinet) | noun|jää_kaappi, mangled (# → _) |
= and ! pass through the cleaning pipeline because they are not reserved CHAT %mor syntax characters. # is reserved in traditional CHAT MOR for prefix markers (e.g., v|#un#do), so the sanitizer replaces it with _.
Gotcha:
=ambiguity with legacy CLAN translation glosses. Legacy CLAN%mortiers use=for translation glosses (e.g.,n|perro=dog), a convention predating UD adoption. The parser treats=identically in both cases; it is preserved as part of the lemma string. This means legacyn|perro=dogparses successfully but the translation semantics are lost: the model storesperro=dogas a single lemma, indistinguishable from an Estonian compound likemaja=uks. Since we cannot reliably disambiguate the two uses without language-specific context, legacy translation glosses are silently absorbed into the lemma. Files with legacy=translationsyntax still parse and round-trip correctly, but the translation information is not semantically accessible. This affects corpora that predate our UD MOR adoption and lack Stanza coverage for their language.
Multi-Word Expression Lemmas (Stanza _ Convention)
Stanza uses underscores in lemmas to represent multi-word expressions across many languages: New_York, parce_que (French), pick_up (English), a_causa_di (Italian). The current cleaning pipeline strips underscores entirely (New_York → NewYork), which is a known data quality issue and should be treated as an open data-quality limitation of the current mapper.
Multi-Value Features (Commas in Feature Values)
UD encodes multi-value features with commas: PronType=Int,Rel means a word is both interrogative and relative. These commas appear in the CHAT %mor feature suffix and are preserved as-is:
pron|wat-Int,Rel
This is sometimes mistaken for a compound lemma marker, but commas in UD always appear in the feature column (CONLLU column 6), never in the lemma column (CONLLU column 3). In CHAT %mor, they appear after the - feature separator, not inside the lemma. The grammar, both parsers, and the data model all accept commas in feature values. See Section 3: Multi-Value Features above.
Future Direction
The current handling of compound lemma boundaries is inconsistent across languages. A possible future improvement is a unified Unicode separator character that would normalize all compound/derivational boundary markers (=, !, #, and potentially _) into a single convention. This is not implemented and requires a design decision on which character to use and whether to preserve the original markers in a structured field.
Data Model
The Rust data model in talkbank-model represents %mor tiers with these types:
MorTier
The top-level tier container:
pub struct MorTier {
pub tier_type: MorTierType, // MorTierType::Mor
pub(crate) items: MorItems, // Vec<Mor> wrapper; accessed via accessor methods
pub terminator: Terminator, // typed terminator (.`, `?`, `!`, `+...`, etc.)
pub span: Span, // source location
}
Mor (Item)
One item aligned with one main-tier word:
pub struct Mor {
pub main: MorWord, // required main word
pub post_clitics: SmallVec<[MorWord; 2]>, // optional ~clitics
}
MorWord
A single morphological word (POS + lemma + features):
pub struct MorWord {
pub pos: PosCategory, // e.g., "noun"
pub lemma: MorStem, // e.g., "dog"
pub features: SmallVec<[MorFeature; 4]>, // e.g., [Plur]
}
MorFeature
A morphological feature with optional key:
pub struct MorFeature {
key: Option<Arc<str>>, // e.g., Some("Number") or None
value: Arc<str>, // e.g., "Plur"
}
Construction examples:
// Flat feature (current convention)
MorFeature::new("Plur") // key=None, value="Plur"
MorFeature::new("S3") // key=None, value="S3"
MorFeature::new("Int,Rel") // key=None, value="Int,Rel"
// Keyed feature (UD-standard, forward-compatible)
MorFeature::new("Number=Plur") // key=Some("Number"), value="Plur"
MorFeature::new("Tense=Past") // key=Some("Tense"), value="Past"
// Explicit constructors
MorFeature::flat("Plur")
MorFeature::with_key_value("Number", "Plur")
Lossless roundtrip guarantee: MorFeature::new auto-detects the = delimiter. Features without = are flat; features with = split into key+value. Serialization reproduces the original format exactly, flat features stay flat, keyed features keep their key.
PosCategory and MorStem
Both are interned Arc<str> newtypes for memory efficiency:
pub struct PosCategory(pub Arc<str>); // interned via pos_interner()
pub struct MorStem(pub Arc<str>); // interned via stem_interner()
Common values (noun, verb, the, a, be, etc.) are pre-populated in the interner. Cloning is O(1), atomic reference count increment.
Memory Layout
The model uses SmallVec for inline storage of common cases:
Mor.post_clitics: SmallVec<[MorWord; 2]>: most words have 0-1 cliticsMorWord.features: SmallVec<[MorFeature; 4]>: most words have 0-4 featuresMorFeaturekey and value areArc<str>, interned for deduplication
For a typical 30-word utterance with %mor, the model allocates approximately 30 Mor items, each with 1 MorWord and 0-4 MorFeature values. The interning system ensures that repeated POS tags, stems, and feature values share a single allocation across the entire file.
Grammar
The tree-sitter grammar for %mor is defined in grammar.js. The relevant rules:
mor_content → mor_word (mor_post_clitic)*
mor_post_clitic → tilde mor_word
mor_word → mor_pos pipe mor_lemma (mor_feature)*
mor_feature → hyphen mor_feature_value
mor_feature_value → /[^\.\?\|\+~\-\s\r\n]+/
Key design decisions:
mor_feature_valueaccepts=and!: The regex[^\.\?\|\+~\-\s\r\n]+matches any characters except the MOR structural delimiters. This meansNumber=Plurparses as a singlemor_feature_valuenode. The split on=happens in the model layer, not the grammar, following the “parse, don’t validate” principle.mor_feature_valueaccepts,: Multi-value features likeInt,Relparse as a single node.- No compound/prefix rules: The grammar has no rules for
+(compounds) or#(prefixes) in the UD MOR format. These are legacy CHAT MOR features not used in UD-style output.
Parser
The tree-sitter parser produces MorTier from CHAT text. It is GLR-based and error-recovering, producing a CST that the Rust talkbank-parser crate walks to construct MorTier. Used by the CLI, LSP, and batchalign. High-frequency values (PosCategory, MorStem) are interned via Arc<str> during construction.
The corpus/reference/ set is the correctness gate for %mor parsing,
every file must parse and round-trip cleanly. The file count grows as
new constructs are added; run find corpus/reference -name '*.cha' | wc -l
to get the live total.
Validation
The %mor tier undergoes several validation checks:
Content Validation (E711)
Every MorWord is checked for:
- Empty POS:
|lemmawith no POS before the pipe - Empty lemma:
pos|with no lemma after the pipe - Empty feature: bare
-separator with no feature text
Main-tier Alignment (E705 / E706)
The %mor tier must align 1-to-1 with the main tier’s alignable
words (excluding pauses, events, and other non-word content). The
number of Mor items must equal the number of alignable main-tier
words. The validator emits E705 MorCountMismatchTooFew when
%mor has fewer items than the main tier and E706
MorCountMismatchTooMany when it has more. Terminator-mismatch
errors are emitted separately as E707 (presence) and E716 (value).
GRA Alignment (E720)
When both %mor and %gra tiers are present, the number of %gra
relations must equal the number of %mor chunks (including
clitics and the terminator). A mismatch emits E720
MorGraCountMismatch. This is computed via MorTier::count_chunks().
(%gra’s own internal validators, E708 malformed relation, E709
invalid index, E712 word-index out of range, E713 head-index out of
range, E721 non-sequential index, E722 no ROOT, E723 multiple
ROOTs, E724 circular dependency, are documented in
Dependent Tiers § %gra.)
JSON Serialization
The MorTier serializes to JSON using serde. MorFeature serializes as a plain string ("Plur" or "Number=Plur"), so the JSON schema is simply "type": "string". Example:
{
"tier_type": "Mor",
"items": [
{
"main": {
"pos": "pron",
"lemma": "I",
"features": ["Prs", "Nom", "S1"]
}
},
{
"main": {
"pos": "verb",
"lemma": "want",
"features": ["Fin", "Ind", "Pres", "S1"]
}
},
{
"main": {
"pos": "noun",
"lemma": "cookie",
"features": ["Plur"]
}
}
],
"terminator": "."
}
When key=value features are present, they serialize with the key included:
"features": ["Number=Plur", "Tense=Past"]
The JSON schema for MorFeature is "type": "string" regardless of whether keys are present.
Migration from Traditional CHAT MOR
What Changed
The traditional CHAT MOR format (CLAN-era) used a complex, hierarchically structured notation:
%mor: pro:sub|I v|want n|cookie-PL .
Key differences from the UD MOR format:
| Aspect | Traditional CHAT MOR | UD MOR |
|---|---|---|
| POS tags | CLAN categories (pro:sub, v, n, adj, adv) | Lowercased UPOS (pron, verb, noun, adj, adv) |
| POS subtypes | Colon-separated (pro:sub, det:art, v:aux) | Flat (subtypes dropped or encoded differently) |
| Features | CLAN suffix system (-PL, -PAST, -3S, -PRES) | UD feature values (-Plur, -Past, -S3, -Pres) |
| Compounds | + separator (`n | +n|black+n|bird`) |
| Prefixes | # separator (`v | #un#do`) |
| Morpheme segmentation | Full segmentation (v|eat&PAST) | Not used (features are abstract, not morphemic) |
| Translations | = separator (n|perro=dog) | Not present in base format (separate mechanism) |
Types the Model Does Not Have
The data model has none of the types that the CLAN-era MOR structure needed:
MorSuffix: suffix with type discriminant (fusional,derivational, etc.)MorCompound: compound word with+separatorMorPrefix: prefix with#separatorMorSubcategory: POS subcategory after colonAnnotatedChunk: chunk with optional translationChunk: enum of word/compound/terminator
The flat MorWord { pos, lemma, features } structure covers them. The model has 4 types: MorTier, Mor, MorWord and MorFeature.
Backward Compatibility
The grammar still accepts many traditional CHAT MOR constructs (colons in POS tags, etc.) because the reference corpus contains files in both formats. The parser produces the same flat MorWord regardless; legacy constructs are mapped to the simplified structure during parsing.
What Stays the Same
Despite the format changes, fundamental CHAT conventions remain:
- Pipe (
|) separates POS from lemma - Hyphen (
-) introduces features - Tilde (
~) marks post-clitics - Space separates items
- Terminator ends the tier
- 1-to-1 alignment with main tier words
Toward Full UD Compatibility
The current format is UD-inspired but not UD-compliant. Here is a roadmap of what would be needed for full lossless UD round-tripping:
Already Supported
- POS tags (UPOS equivalents)
- Lemmas
- Feature values (flat and key=value)
- MWT expansions (clitics)
- Dependency relations (via
%gra)
Gaps Remaining
-
Feature keys: The model supports
Key=Valuefeatures, but batchalign’s mapper currently emits flat values only. When the mapper switches to emittingNumber=Plurinstead of justPlur, the parser, model, and serializer handle it automatically with no code changes. -
Person+Number composites: UD has separate
Person=3andNumber=Singfeatures. CHAT combines them into-S3(3rd person singular). DecomposingS3back toPerson=3|Number=Singwould require a lookup table or a convention change. -
Multi-value feature delimiter: UD uses commas (
PronType=Int,Rel). CHAT preserves these commas in the feature value, but the semantic structure (two separate values) is not explicitly modeled. The model treatsInt,Relas an opaque string. -
XPOS: UD provides language-specific POS tags (XPOS) alongside universal tags (UPOS). CHAT
%morhas no XPOS field. This information is simply not represented. -
Morpheme-level analysis: UD’s
MISCfield can encode morpheme boundaries and glosses. CHAT’s UD MOR format does not attempt morpheme segmentation, features are abstract grammatical categories, not morphemic decompositions.
The Path Forward
The model is designed so that moving toward UD compliance requires no breaking changes:
MorFeaturealready supportsKey=Value, just needs the mapper to emit keysPosCategoryis an opaque string, could hold XPOS in a separate field if needed- JSON schema uses
"type": "string"for features, adding keys doesn’t break consumers - The grammar already accepts
=in feature values, no grammar changes needed
The migration can happen incrementally: the mapper starts emitting key=value features, existing flat data continues to parse identically, and corpus files can be upgraded at their own pace.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Phon Tiers (%xmodsyl, %xphosyl, %xphoaln, %xphoint)
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
The Phon extension tiers provide syllable-level phonological annotation, segmental alignment between target and actual IPA, and per-phone time intervals. They are produced by the Phon application. As of Phon 4.0.0-beta.9 (2026-06-25), Phon reads and writes CHAT natively (see “Data quality notes” below for what this means for existing corpus files).
The authority for these formats is upstream, not this chapter. Phon
generates these tiers, so Phon specifies them; the current specification comes
from Phon’s maintainer. This chapter and the E735-E746 error specs are
chatter’s implementation OF that specification, not a substitute for it. Where
chatter disagrees with it, the disagreement is a finding to take upstream rather
than a rule to settle here.
chatter parses and validates all four tiers as first-class CHAT tiers.
The
xprefix. Phon exports use the unprefixed names. Chatter accepts both thex-prefixed names (%xmodsyl,%xphosyl,%xphoaln,%xphoint) and the unprefixed names (%modsyl,%phosyl,%phoaln,%phoint); the parser and validator key off the tier kind, not the literal prefix. Chatter’s canonical serialized form currently isx-prefixed; input compatibility and output spelling are distinct.
CLAN’s dependent-tier definitions declare %phoint (clan-info commit
f062b58). Acceptance depends on the definitions actually loaded by CHECK; an
older bundled depfile.cut may still reject it. The entry permits arbitrary
tier content and does not specify or validate the internal Phon interval
grammar. It does not affect Chatter’s aliases or canonical output spelling.
The four tiers
| Tier | Source | Carries | Word separator |
|---|---|---|---|
%xmodsyl | %mod | Syllabification of the model/target transcription | space |
%xphosyl | %pho | Syllabification of the actual transcription | space |
%xphoaln | %mod+%pho | Phone-by-phone alignment of model ↔ actual | space |
%xphoint | %pho | Per-phone time intervals (0x15 time bullets) | / |
%xmodsyl, %xphosyl, and %xphoaln are word-aligned to their source tier(s)
with single ASCII spaces. %xphoint uses / (space-slash-space) as its word
separator because single spaces already separate the phone and bullet tokens
inside each word.
Tier formats
%xmodsyl / %xphosyl, syllabification
A word is one or more phone:CODE units concatenated with no internal
whitespace; words are separated by single spaces. The phone is one IPA phone
(IPA length is written with the modifier letter ː, U+02D0, never an ASCII
colon, so the : separator is unambiguous). A leading stress marker (ˈ
primary, ˌ secondary) is part of the phone it precedes.
Pause fillers. Phon keeps every word-aligned phonology tier in index
lockstep with the main tier: when the main tier carries a pause, the pause
token ((.), (..), (...), or numeric (x.x)) is mirrored at the same
word position on %mod, %pho, %xmodsyl, and %xphosyl (and as a
(..)↔(..) pair on %xphoaln). A pause filler is a valid word on the
syllabification tiers; it carries no phone:CODE structure and must mirror
the same pause token as the source-tier word at its position. Numeric pauses
((1.5), minutes:seconds (1:02.5)) are accepted on the same footing as
the three untimed forms, per Greg Hedlund’s spec.
The constituent code is one character. The legal codes are O N C L R E A D U:
| Code | Constituent | Notes |
|---|---|---|
O | Onset | |
N | Nucleus | monophthong nucleus |
C | Coda | |
L | Left appendix | e.g. /s/ in an /s/-stop cluster |
R | Right appendix | e.g. final /z/ in a complex coda |
E | OEHS (onset of empty-headed syllable) | e.g. the stop element of an affricate |
A | Ambisyllabic | |
D | Diphthong | a nucleus member of a diphthong/triphthong; treated as a nucleus |
U | Unknown | Phon could not assign a concrete constituent; common on %xphosyl when the model %xmodsyl is fully syllabified |
The reference corpus exercises all nine codes. The
tiers/phon-constituent-controls.cha fixture supplies explicitly authored E
and A acceptance controls; it is not a production attestation or a claim that
Chatter automatically assigns syllable constituents. The corpus workflow checks
their typed-code conversion, output, parsing, validation and roundtrip.
The remaining Phon SyllableConstituentType mnemonics, B (boundary),
S (stress), W (word boundary), T (tone), are not emitted on these
tiers: boundary, stress, and tone need no per-phone marker.
*CHI: I want three .
%mod: aɪ wɑnt θri
%xmodsyl: a:Dɪ:D w:Oɑ:Nn:Ct:C θ:Oɹ:Oi:N
%pho: aɪ wɑn fwi
%xphosyl: a:Dɪ:D w:Oɑ:Nn:C f:Ow:Oi:N
%xphoaln, phone alignment
A word is one or more comma-separated pairs; a pair is model↔actual (↔ is
U+2194). Either side may be ∅ (U+2205, empty set): ∅ on the left is an
epenthesis (a phone produced but not targeted); ∅ on the right is a deletion.
Both sides are never ∅ at once.
*CHI: the best .
%mod: ðə bɛst
%pho: ðə bɛs
%xphoaln: ð↔ð,ə↔ə b↔b,ɛ↔ɛ,s↔s,t↔∅
The alignment lists segments (phones). Suprasegmental stress (ˈ/ˌ) that
may appear on the %mod/%pho word is therefore not part of the alignment
pairs; the reconstruction checks below compare modulo those stress markers.
%xphoint, per-phone intervals
%xphoint gives the time segmentation of each individual phone on %pho,
effectively phone-level bullets analogous to the word-level timing on %wor.
Groups (one per %pho word) are separated by /. Within a group, each phone
is followed by a CLAN time-alignment bullet: the byte 0x15 (NAK), the interval
start_end, then 0x15.
*CHI: I want . •0_500•
%pho: aɪ wɑnt
%xphoint: aɪ •0_250• / w •250_320• ɑ •320_400• n •400_460• t •460_500•
(Bullets are shown as • above; in the file they are the 0x15 byte.)
Validation
These checks run by default. Pass --suppress xphon to silence the entire
Phon %x validation surface, or suppress an individual code. (The --check-xphon
flag is a deprecated no-op: the checks are always on.)
Word-count cross-checks:
%xmodsyl↔%mod: E725,%xphosyl↔%pho: E726. Always a strict equality: inter-word pauses are mirrored onto both tiers by construction (spec §1 rule 4), so the counts must match exactly.%xphoaln↔%mod: E727, ↔%pho: E728. NOT strict equality: a pause word present on only one of%mod/%phoforms its own%xphoalnalignment word (its other side entirely∅, e.g.(..)↔∅) and consumes a word slot only on the tier bearing the pause (spec §2 rule 5). E727/E728 compare each tier’s word count against%xphoaln’s count after excluding the one-sided pause words that do not consume a slot on that tier, not against%xphoaln’s raw word count.
Reconstruction uses those same independent source positions. A pause on only
one side must not cause subsequent alignment words to be compared with the
wrong %mod or %pho word. Count mismatches remain errors rather than being
hidden by the pause exception.
Content checks:
| Code | Tier | Rule |
|---|---|---|
| E735 | xmodsyl/xphosyl | a non-pause-filler unit is not a well-formed phone:CODE (no :, empty phone, or empty code) |
| E736 | xmodsyl/xphosyl | a constituent code is not one of O N C L R E A D U |
| E737 | xmodsyl | stripping codes and concatenating phones does not reproduce the %mod word (a pause filler must mirror the same pause token) |
| E738 | xphosyl | stripping codes and concatenating phones does not reproduce the %pho word (a pause filler must mirror the same pause token) |
| E739 | xphoaln | a pair is malformed (not exactly one ↔, an empty side, or ∅↔∅) |
| E740 | xphoaln | concatenating the model sides (skipping ∅, modulo stress and ^/. syllable boundaries) does not reproduce the %mod word |
| E741 | xphoaln | concatenating the actual sides (skipping ∅, modulo stress and ^/. syllable boundaries) does not reproduce the %pho word |
| E742 | xphoint | a bullet has start >= end |
| E743 | xphoint | interval start times are not non-decreasing across the tier |
| E744 | xphoint | the first start / last end falls outside the record’s media bullet (1 ms tolerance) |
| E745 | xphoint | a group’s phones do not reproduce the %pho word |
| E746 | xphoint | the number of groups does not equal the %pho word count |
See Alignment Architecture for the word-count implementation.
Parsing strategy
- %xmodsyl / %xphosyl: stored as flat word strings
(
talkbank-model::dependent_tier::phon::SylTier), consistent with how%phoand%modstore flat phone words. The validator tokenizes each word into typedphone:CODEunits (PositionCode) to apply the content rules above; the IPA characters themselves stay verbatim for exact round-trip. - %xphoaln: each word is parsed into a
Vec<AlignmentPair>, whereAlignmentPair { source, target }carries onemodel↔actualmapping (Noneis∅). - %xphoint: parsed into typed groups of
(phone, bullet)pairs (XphointTier/XphointGroup/PhoneInterval), reusing the same0x15bullet machinery as%wor.
Deep phonological analysis is Phon’s domain; chatter parses the structure that validation needs and keeps the IPA content verbatim.
Phon XML source format
In Phon’s native XML format, phonological data is stored as structured elements:
<ipaTarget>
<pho>
<pw>
<ph scType="onset"><base>θ</base></ph>
<ph scType="nucleus"><base>ɹ</base></ph>
<ph scType="nucleus"><base>i</base></ph>
</pw>
</pho>
</ipaTarget>
Each <pw> (phonological word) element contains <ph> elements with syllable
constituent types (scType). The <alignment> element provides phone-level
mappings between target and actual using index-based <pm> (phone map) entries.
Data quality notes
A small percentage of Phon corpus XML records have an orthography↔IPA word-count
mismatch: the number of <pw> elements in <ipaTarget> / <ipaActual> differs
from the number of <w> elements in <orthography>. This is expected in child
phonology data: children may produce extra syllables, partial words, or
over-productions relative to the target.
For current counts on a local CHILDES/TalkBank data tree, run:
python3 scripts/analysis/scan_phon_mismatches.py /path/to/data
A subset of existing corpus CHAT files handle this discrepancy inconsistently between tiers:
%mod/%phomap IPA words to orthography words one to one; extras are silently dropped.%xmodsyl/%xphosyl/%xphoalncarry the full IPA word set, undropped.
This produces CHAT files where %xmodsyl may have more words than %mod,
triggering the E725-E728 word-count errors. As of Phon 4.0.0-beta.9
(2026-06-25), Phon reads and writes CHAT natively; we have not yet seen
output from that native export and do not know whether it reproduces this
inconsistency. Files already affected remain in the corpora and must
continue to validate against the behavior above.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Word Syntax
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
Words are the primary content unit on the main tier. CHAT defines several word types and annotation mechanisms.
Standalone Words
Most words are simple tokens separated by whitespace:
*CHI: I want a cookie .
Words can contain Unicode characters for any language:
*CHI: ich möchte Kekse .
Compounds
Compound words join multiple elements with +:
*CHI: I want ice+cream .
Special Word Forms
Shortened Forms
Parentheses mark omitted portions of a word:
*CHI: (be)cause I want it .
The full form is because; the child produced cause.
Replacements
Square brackets with colon mark what the speaker actually meant:
*CHI: I goed [: went] to the store .
The speaker said “goed” but the intended word was “went”.
Language Markers
The @s: suffix marks a word’s language in multilingual transcripts:
*CHI: I want a Keks@s:deu .
When a whole stretch switches language, annotate the group rather than suffixing every word:
*CHI: ik weet niet <how to do it> [@s] .
*TEA: us samay <kyaa hotaa hai> [@s:hin] .
[@s:code] names the language; bare [@s] resolves the way a bare word@s
does. Every word in the <> scope takes that language, exactly as if each
carried the suffix. As with any scoped annotation, a single item needs no
angle brackets: hallo [@s] is well-formed and means what hallo@s means.
A word inside the span may carry its OWN marker, and the word wins:
*TEA: <rocket@s:eng jaise jaataa hai> [@s:hin] .
That is not redundancy to avoid. It is how a borrowed word is marked inside a switched clause, and it is what transcribers actually write: the span carries the matrix language of the stretch, the suffix carries the donor language of one item. Resolution is innermost-first, and each layer is recorded with its own provenance, so a consumer can tell which mark decided a given word.
A word can also carry one special-form marker naming what kind of form it is
(gumma@c for a child-invented word, b@l for a letter). The complete set,
with meanings and examples, is the table in
Symbols.
A partial copy of a closed set is worth less than a link to the whole one: a
hand-picked subset drifts (glossing @si as “signed word”, which is @sl,
when @si is singing).
Annotations
Words and groups can carry post-positioned annotations in square brackets:
Error Marking
*CHI: he goed [*] to school .
[*] marks an error. More specific error codes can follow: [* m:+ed].
Explanations
*CHI: that one [= the red ball] .
[= text] provides an explanation or gloss.
Replacements
*CHI: I wanna [: want to] go .
[: text] marks the target/intended form.
Best Guess
*CHI: I want the birfer [?] .
[?] marks uncertain transcription.
Events and Actions
Paralinguistic Events
Events marked with &= describe non-speech sounds:
*CHI: &=laughs I want cookie .
*CHI: &=coughs .
Fillers
Fillers are marked with &-:
*CHI: &-um I want &-uh cookie .
Interposed Speech (Other Speaker)
Brief background speech from a different speaker is marked with the
&*SPK:text prefix, it captures the interjection without creating
a full turn line:
*CHI: I want &*MOT:careful a cookie .
This says CHI was speaking and MOT briefly said “careful” mid-turn.
If the intervention is substantial enough to constitute its own turn,
transcribe it as a separate *MOT: utterance instead. Model:
crates/talkbank-model/src/model/content/other_spoken.rs.
(Note: [^ text] is a freecode, a standalone free-form
researcher annotation that sits as its own content item on the main
tier (variant of UtteranceContent::Freecode, sibling of Word and
Group; it is NOT attached to any word). See grammar/grammar.js
rule freecode and
crates/talkbank-model/src/model/content/utterance_content/. Used
for transcriber notes that are independent of any single word; for
notes about a single word use [% text] or [= text] instead.)
Pauses
*CHI: I (.) want (..) a (...) cookie .
*CHI: I (1.5) want a cookie .
(.): short pause(..): medium pause(...): long pause(N.N): timed pause in seconds
Overlap
Overlapping speech between speakers uses angle brackets and overlap markers:
*MOT: do you want <a cookie> [>] ?
*CHI: <cookie> [<] !
[>]: follows the overlap (this speaker started first)[<]: overlaps the previous speaker
Retrace and Repetition
Groups followed by retrace markers indicate speech disfluencies:
*CHI: <I want> [/] I want a cookie .
*CHI: <I want> [//] I need a cookie .
*CHI: <I want a> [///] give me a cookie .
[/]: partial retrace (speaker repeats the same words)[//]: full retrace (speaker restarts with different words)[///]: multiple retracing (multiple false starts)[/-]: reformulation (speaker rephrases with different structure)
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
The CHAT Word
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
“Word” is the most complex and most misunderstood concept in CHAT. This chapter documents what a word actually is, how the grammar parses it, and how the Rust model represents it. If you maintain this codebase, you will encounter word-level bugs. This chapter exists so you can understand them.
The Fundamental Rule
Whitespace delimits words. Contiguous non-whitespace characters form one word token. This applies everywhere on the main tier.
*CHI: hello world .
^^^^^ word: "hello"
^^^^^ word: "world"
The grammar uses extras: $ => [] – no implicit whitespace. Whitespace
nodes (whitespaces, space) are explicit in the CST. Tree-sitter does
not skip whitespace between tokens. This is the foundation of every
tokenization decision in the grammar.
There are no exceptions to this rule. Every ambiguity described in this chapter is resolved by applying this rule consistently.
Word Structure
A word in the grammar is standalone_word – a sequence of an optional
prefix, a required body, optional suffixes, and an optional POS tag.
The following diagram shows the full decomposition. All named nodes are separate CST children.
flowchart TD
sw["standalone_word\n(grammar.js, prec.right 6)"]
zero["zero\n'0' -- omission prefix"]
wp["word_prefix\n'&-' filler | '&~' nonword | '&+' fragment"]
wb["word_body\n(required)"]
fm["form_marker\n@b, @c, @d, @z:label, ..."]
wls["word_lang_suffix\n@s, @s:eng, @s:eng+fra"]
pos["pos_tag\n$n, $v, $adj, ..."]
sw -->|"optional prefix"| zero & wp
sw -->|"required"| wb
sw -->|"optional suffix"| fm & wls
sw -->|"optional"| pos
ws["word_segment\npure spoken text"]
short["shortening\n'(text)' omitted sound"]
sm["stress_marker\nprimary or secondary"]
len["lengthening\n':' one or more colons"]
op["overlap_point\none of four brackets"]
cae["ca_element\nsingle CA marker"]
cad["ca_delimiter\npaired CA marker"]
ub["underline_begin\ncontrol char pair"]
ue["underline_end\ncontrol char pair"]
cm["'+'\ncompound marker"]
wb -->|"children (any order)"| ws & short & sm & len & op & cae & cad & ub & ue & cm
In the grammar (search grammar/grammar.js for the standalone_word
and word_body rules), the structure is:
standalone_word: $ => prec.right(6, seq(
optional(choice($.word_prefix, $.zero)),
$.word_body,
optional($.form_marker),
optional($.word_lang_suffix),
optional($.pos_tag),
)),
word_body: $ => prec.right(choice(
seq(
choice($.word_segment, $.shortening, $.stress_marker),
repeat(choice($.word_segment, $.shortening, $.stress_marker, $._word_marker)),
),
seq(
choice($.overlap_point, $.ca_element, $.ca_delimiter, $.underline_begin),
choice($.word_segment, $.shortening, $.stress_marker),
repeat(choice($.word_segment, $.shortening, $.stress_marker, $._word_marker)),
),
)),
word_body has two branches:
- Standard start: the word begins with
word_segment,shortening, orstress_marker, followed by any number of body children. - Marker-initial: the word begins with a structural marker (overlap,
CA, underline), but that marker must be immediately followed by text
content. This prevents degenerate words like a standalone overlap marker
from forming a valid
standalone_word.
Lengthening and + (compound marker) are excluded from starting a word
body. This is how standalone : falls through to separator(colon) –
see Section 5 (Tokenization Ambiguities) below.
The word_segment Purity Invariant
word_segment contains ONLY pure spoken text. All structural markers
are separate typed children in word_body, never consumed by word_segment.
This is a hard invariant with three consequences:
cleaned_text()never scans for markers. It concatenatesTextandShorteningelements. No stripping needed.- Validation finds ALL markers by type. Overlap markers, CA elements,
and underline pairs are always
WordContentvariants, regardless of position within the word. - Editors get typed CST nodes. Syntax highlighting, bracket matching, and hover info work on individual markers, not opaque substrings.
How it works
word_segment is a DFA token at prec(5) with a regex that excludes
all structural characters. The exclusions are generated from the symbol
registry (grammar/src/generated_symbol_sets.js) – never hand-written.
word_segment: $ => token(prec(5, seq(
WORD_SEGMENT_FIRST_RE, // generated: excludes structural chars + '0' at start
WORD_SEGMENT_REST_RE, // generated: excludes structural chars
))),
Full exclusion table
Every character in this table is excluded from word_segment and
becomes a separate typed node in the CST.
| Category | Characters | CST node type |
|---|---|---|
| Overlap markers | ⌈ ⌉ ⌊ ⌋ | overlap_point |
| CA elements | ↑ ↓ ≠ ∾ ⁑ ⤇ ∙ Ἡ ↻ ⤆ | ca_element |
| CA delimiters | ∆ ∇ ° ▁ ▔ ☺ ♋ ⁇ ∬ Ϋ ∮ ↫ ⁎ ◉ § | ca_delimiter |
| Stress markers | ˈ ˌ | stress_marker |
| Colons | : | lengthening |
| Underline markers | \x02\x01, \x02\x02 | underline_begin / underline_end |
| Brackets | [ ] < > ( ) { } | structural (annotations, groups) |
| Punctuation | . ! ? , ; + | terminators, separators, compound |
| CHAT prefixes | @ $ & * % | headers, events, speakers |
| Intonation contours | ⇗ ↗ → ↘ ⇘ ≈ ≋ ∞ ≡ | content-level markers |
| Group delimiters | ‹ › " " 〔 〕 | pho/sin groups, quotes |
| Control chars | \x01-\x08, \x15 | bullets, underline |
First-character-only exclusion: 0 is excluded from the first
character of word_segment (it is the omission prefix). 0 in
non-initial positions is valid: 200, h0me, abc0 all parse correctly.
The Rust Data Model
Word struct
The Word struct (crates/talkbank-model/src/model/content/word/word_type.rs)
is the canonical typed representation:
pub struct Word {
pub span: Span,
pub word_id: Option<SmolStr>,
content: WordContents,
pub category: Option<WordCategory>,
pub form_type: Option<FormType>,
pub lang: Option<WordLanguageMarker>,
pub part_of_speech: Option<SmolStr>,
pub inline_bullet: Option<Bullet>,
}
Key fields:
raw_text(): a derived spelling, including typed markers, also emitted as the JSONraw_textfield. It is not stored independently and is not an exact source slice. See the computed-field contract.content: aWordContents(SmallVec-backed sequence ofWordContentelements). This is the structured decomposition. Most words have 1-2 elements; SmallVec avoids heap allocation for the common case. Access it withcontent()and typed mutation methods, which invalidate cached cleaned text.category: optional prefix (Omission,CAOmission,Filler,Nonword,PhonologicalFragment).form_type: optional@suffix (@cchild-invented,@ddialect,@z:labeluser-defined, etc.). The user-defined form requires the colon and a label (@z:label); a colon-less marker such as@zzzis not a valid form and is rejected with E203 (matching CLAN CHECK 147).lang: optional@slanguage marker (Shortcut,Explicit,Multiple,Ambiguous).part_of_speech: optional$tag.
WordContent enum
WordContent (crates/talkbank-model/src/model/content/word/content.rs)
is the enum of everything that can appear inside a word body. Each variant
maps directly to a grammar node.
| Grammar node | WordContent variant | Rust type | Example |
|---|---|---|---|
word_segment | Text | WordText(NonEmptyString) | hello, want |
word_segment in a @u word | Phonetic | WordPhonetic(NonEmptyString) | rɛmbə˞ in rɛmbə˞@u |
shortening | Shortening | WordShortening(NonEmptyString) | (be) in (be)cause |
overlap_point | OverlapPoint | OverlapPoint | ⌈, ⌉2 |
ca_element | CAElement | CAElement | ↑, ↓ |
ca_delimiter | CADelimiter | CADelimiter | ∆, ° |
stress_marker | StressMarker | WordStressMarker | ˈ primary, ˌ secondary |
lengthening | Lengthening | WordLengthening { count: u8 } | : = 1, :: = 2, ::: = 3 |
| (caret in word) | SyllablePause | WordSyllablePause | ^ in o^ver |
underline_begin | UnderlineBegin | UnderlineMarker | \x02\x01 |
underline_end | UnderlineEnd | UnderlineMarker | \x02\x02 |
+ (compound) | CompoundMarker | WordCompoundMarker | + in ice+cream |
~ (clitic boundary) | CliticBoundary | WordCliticBoundary | ~ in le~ha |
cleaned_text()
Word::cleaned_text() derives NLP-ready text from content by
concatenating only Text, Phonetic, and Shortening variants:
pub fn compute_cleaned_text(&self) -> String {
let mut result = String::new();
for item in &self.content {
match item {
WordContent::Text(t) => result.push_str(t.as_ref()),
WordContent::Shortening(s) => result.push_str(s.as_ref()),
_ => {}
}
}
result
}
This works because the purity invariant guarantees that Text elements
never contain structural markers. There is nothing to strip.
@u phonetic forms are typed phonetic content
A @u word is a phonetic transcription (UNIBET or, more
usually, IPA) standing in a word slot: used when the orthographic word is
unknown, unintelligible, or a paraphasia, frequently the spoken side of
a [: target] replacement in aphasia data (rɛmbə˞@u [: remember]).
Because its content obeys phonetic conventions rather than orthographic
word conventions, the parsers fold a @u word’s body into a single
WordContent::Phonetic(WordPhonetic) node instead of Text. This makes
“orthographic rules apply to orthographic words only” a property of the
model: word-hygiene rules structurally cannot reach phonetic content.
The phonetic string itself is deliberately lenient (IPA, ASCII UNIBET,
and X-SAMPA all pass; only non-emptiness is enforced), matching the
%pho tier’s PhoWord stance. In to-json output the node appears as
{"type": "phonetic", "content": "..."}, and cleaned_text remains the
phonetic string verbatim. The scope is @u only: sibling special forms
(@b, @o, …) remain orthographic words.
Examples:
| Input | content | cleaned_text() |
|---|---|---|
hello | [Text("hello")] | hello |
(be)cause | [Shortening("be"), Text("cause")] | because |
no:: | [Text("no"), Lengthening(2)] | no |
ice+cream | [Text("ice"), CompoundMarker, Text("cream")] | icecream |
le~ha | [Text("le"), CliticBoundary, Text("ha")] | leha |
ja^ja | [Text("ja"), SyllablePause, Text("ja")] | jaja |
he↑llo | [Text("he"), CAElement(PitchUp), Text("llo")] | hello |
°soft° | [CADelimiter(Softer), Text("soft"), CADelimiter(Softer)] | soft |
ˈhello | [StressMarker(Primary), Text("hello")] | hello |
⌈hello⌉ | [OverlapPoint(TopBegin), Text("hello"), OverlapPoint(TopEnd)] | hello |
The result is cached via OnceLock on first access.
What is included in cleaned_text vs what is stripped
The following table is the complete inventory of how every word-internal
element contributes to (or is excluded from) cleaned_text(). This must
match what NLP pipelines (Stanza, etc.) expect as input.
WordContent variant | Character(s) | In cleaned_text? | Rationale |
|---|---|---|---|
Text | spoken text | YES | The actual word |
Shortening | (be) | YES | Shortened form is still spoken |
CompoundMarker | + | No | Structural boundary, not spoken |
CliticBoundary | ~ | No | Morphological boundary, not spoken |
SyllablePause | ^ | No | Pause between syllables, not spoken |
Lengthening | : :: ::: | No | Prosodic marker, not spoken |
StressMarker | ˈ ˌ | No | Prosodic marker, not spoken |
OverlapPoint | ⌈ ⌉ ⌊ ⌋ | No | Timing marker, not spoken |
CAElement | ↑ ↓ ≠ ∾ ⁑ ⤇ ∙ Ἡ ↻ ⤆ | No | Prosodic annotation |
CADelimiter | ∆ ∇ ° ▁ ▔ ☺ ♋ ⁇ ∬ Ϋ ∮ ↫ ⁎ ◉ § | No | Voice quality annotation |
UnderlineBegin | \x02\x01 | No | Formatting marker |
UnderlineEnd | \x02\x02 | No | Formatting marker |
Characters that stay in word_segment (ARE spoken text):
- Letters (all Unicode)
- Digits (in non-initial position;
0in initial = omission prefix) - Hyphen (
-), part of word text, e.g.,ice-cream,self-conscious - Apostrophe (
'), contractions, e.g.,don't,it's - Hash (
#), appears in some transcription conventions - Underscore (
_), compound boundary in some conventions
Characters NOT in word_segment (excluded by symbol registry): See the full exclusion table in Precedence Decisions in the grammar docs.
Comparison with batchalign2
batchalign2’s annotation_clean() (60 lines of .replace() calls) strips
all the same characters that our grammar excludes from word_segment.
Key differences:
- Parentheses: ba2 COMMENTED OUT the strip. We handle them as
Shortening, the content inside parens IS included incleaned_text. - IPA characters (
ạ ā ʔ ʕ ʰ): ba2 incorrectly strips them. We correctly keep them; they are real phonetic content. - Hyphen (
-): ba2 strips it. We keep it in word_segment because hyphen is a valid word character (contractions, compounds, morphological suffixes in%mortier).
Our design eliminates the need for character-by-character stripping entirely.
cleaned_text() is a simple concatenation of Text + Shortening
elements, with zero scanning.
The Six Tokenization Ambiguities
CHAT was designed for human readability, not machine parsing. Six
characters have context-dependent meanings that the grammar must
disambiguate. Full details with proof grammars are in
grammar/docs/tokenization-rules.md and grammar/docs/precedence-decisions.md.
What follows is a summary for orientation.
1. Overlap markers (⌈⌉⌊⌋)
Adjacent to text = part of the word. Space-separated = standalone
overlap_point. This adjacency rule is a deliberate approximation of
an ideal (edge markers top-level, interior markers in-word) whose full
history, feasibility analysis, and open implementation decision are
documented in Overlap Marker Binding.
Yeah⌋⌈2 hey ONE word: "Yeah⌋⌈2"
Yeah ⌋ ⌈2 hey three tokens: "Yeah", ⌋, ⌈2
Maximal munch at prec(5) makes word_segment consume adjacent overlap
characters. Overlap markers are only recognized as overlap_point when
space-separated on both sides.
2. Zero/omission prefix (0)
Adjacent to word body = omission prefix. Space-separated = action marker.
0die ONE word: standalone_word(zero, word_body("die"))
0 die TWO tokens: nonword(zero), word("die")
standalone_word at prec.right(6) beats nonword at prec(1).
The extras: [] setting prevents whitespace from being skipped between
zero and word_body. The zero token is inlined directly into
standalone_word (not through word_prefix) because tree-sitter’s
precedence does not propagate through intermediate rules. This was proven
empirically with a minimal test grammar – see grammar/docs/precedence-decisions.md.
3. CA parenthetical vs shortening
In CA mode (@Options: CA), a fully parenthesized word (word) is an
uncertain/omitted word (CAOmission), semantically equivalent to 0word.
Partially parenthesized hel(lo) is always a shortening.
@Options: CA
*CHI: (ja) . CAOmission: uncertain "ja"
*CHI: hel(lo) . Shortening: "(lo)" is the shortened part
Distinguishing these requires file-level context (@Options header).
The parser sets WordCategory::CAOmission when the word is fully
parenthesized in CA mode. Isolated parser.parse_word_fragment() calls
cannot determine CA mode – they need a FragmentSemanticContext.
4. Colon – lengthening vs separator
Inside a word (after text): prosodic lengthening. Standalone: separator.
no:: ONE word: Text("no") + Lengthening(2)
hello : world separator(colon)
The DFA always produces lengthening for : (higher precedence). But
word_body rejects lengthening as a first element, so standalone :
cannot form a valid word and falls through to separator(colon). This is
the “constrain the parser, not the DFA” pattern.
5. Plus (+) – compound vs terminator vs linker
Inside a word: compound marker. At line end: terminator prefix. At line start: linker prefix.
ice+cream ONE word with compound marker
and then +... terminator: trailing_off (prec 10 beats prec 5)
+< but I +/. linker: lazy_overlap, terminator: interruption
Terminators and linkers use prec(10), which beats word_segment at
prec(5). No valid CHAT word ends with + – the grammar enforces this
by structure.
6. Bracket annotations vs plain brackets
Bracket annotations ([= text], [=! text], [% text]) use prec(8)
prefix tokens to beat generic bracket handling.
Design: Structured Word Content
standalone_word is a structured grammar, not an opaque token. Every
word-internal marker is a separate child of word_body. A reference copy of this shape is kept in
grammar/docs/pre-coarsening-grammar.js.reference:
word_content: $ => choice(
$.word_segment,
$.shortening,
$.stress,
$.colon,
$.caret,
$.tilde,
$.plus,
$.overlap_point,
$.ca_element,
$.ca_delimiter,
$.underline_begin,
$.underline_end,
),
The design decisions that follow from this:
- All marker characters are excluded from
word_segment, using the symbol registry as the single source of truth for the exclusion sets. - Each marker type is a separate CST child in
word_body, so editors get typed nodes and validation finds structural markers without re-parsing. - The
WordContentenum in the Rust model is aligned 1:1 with the grammar nodes, socleaned_text()reads typed content rather than scanning for and stripping marker characters. - The
word_segmentpurity invariant is a gate: a structural marker is never consumed byword_segment.
The result is one parser, one source of truth for exclusions, and typed markers from grammar through model.
Testing: The word_segment Purity Gate
The purity invariant, each structural marker produces a separate CST
child rather than being consumed by word_segment, is enforced by a
group of tree-sitter corpus tests under
grammar/test/corpus/generated/word/. Each *_in_word_lint.txt file embeds a
structural marker inside a word and asserts the CST splits the word
appropriately:
| Test file | Input | Asserts |
|---|---|---|
overlap_in_word_lint.txt | butt⌈er⌉ | word_segment, overlap_point, word_segment, overlap_point |
ca_element_in_word_lint.txt | CA element inside a word | word_segment, ca_element, word_segment |
ca_delimiter_in_word_lint.txt | CA delimiter pair around a word | ca_delimiter, word_segment, ca_delimiter |
lengthening.txt, lengthening_between_segments.txt | no::, etc. | word_segment, lengthening |
stacked_ca_markers.txt | Multiple adjacent CA markers in one word | Each marker is its own CST child |
Underline and stress invariants are covered by corpus tests elsewhere
in grammar/test/corpus/ and by the parser-equivalence tests in
crates/talkbank-parser-tests/. Each construct has its own test file, as the
spec generators produce from the spec sources.
How to add a new purity-style test
If you add a new structural marker to the grammar:
- Add its characters to the symbol registry
(
spec/symbols/symbol_registry.json). - Run
just symbols-gento regenerate the exclusion sets. - Add a spec in
spec/constructs/that embeds the marker inside a word; regenerate the affected grammar/parser fixtures with the currentspec/toolscommands from Spec Workflow so a per-construct test fixture is created ingrammar/test/corpus/generated/word/. Verify the CST output names each marker as its own child. - Run the full verification sequence:
cd grammar && tree-sitter generate && tree-sitter test cargo test -p talkbank-parser cargo test -p talkbank-parser-re2c --test integration equivalence_reference_corpus cargo test -p talkbank-parser-tests --tests roundtrip_reference_corpus
Key Source Files
| File | What it defines |
|---|---|
grammar/grammar.js | search for standalone_word, word_body, word_segment, _word_marker |
grammar/src/generated_symbol_sets.js | Character exclusion sets (generated, do not edit) |
grammar/test/corpus/generated/word/*_in_word_lint.txt, lengthening*.txt, stacked_ca_markers.txt | Per-construct purity-invariant gate tests |
grammar/docs/tokenization-rules.md | The 6 tokenization ambiguities with full examples |
grammar/docs/precedence-decisions.md | Precedence proofs (zero, colon, purity invariant) |
grammar/docs/pre-coarsening-grammar.js.reference | The structured word grammar shape in word_content form, kept as a reference |
crates/talkbank-model/src/model/content/word/word_type.rs | Word struct |
crates/talkbank-model/src/model/content/word/content.rs | WordContent enum (12 variants) |
crates/talkbank-model/src/model/content/word/word_contents.rs | WordContents (SmallVec-backed sequence) |
crates/talkbank-model/src/model/content/word/category.rs | WordCategory enum (5 variants) |
crates/talkbank-model/src/model/content/word/form.rs | FormType enum (22 variants) |
crates/talkbank-model/src/model/content/word/language.rs | WordLanguageMarker enum (4 variants) |
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Annotations
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
A scoped annotation is a bracketed code written immediately after the thing
it describes: hello [*], <the dog> [//], bobo [= toy], 0 [= ! whining].
It is scoped because it attaches to a specific construct rather than to the
utterance as a whole, which is what separates it from a
postcode (utterance-wide, written before the terminator) and
from a dependent tier (a whole line of its own).
This chapter answers three questions the model makes precise: what can carry annotations, what it means for something to carry none, and why an annotated construct always carries at least one.
What can be annotated
Each of these constructs has exactly two spellings. The list is the count; stating a number beside it is one more thing to keep true, and this line said five above six rows.
| Construct | Bare | Annotated |
|---|---|---|
| Word | Word | AnnotatedWord |
Group <...> | Group | AnnotatedGroup |
| Quotation | Quotation | AnnotatedQuotation |
Event &=laughs | Event | AnnotatedEvent |
Action 0 | Action | AnnotatedAction |
| Retrace | Retrace | AnnotatedRetrace |
Everything else in an utterance is a leaf that takes no scoped annotation: pauses, separators, overlap points, bullets, freecodes, and the long-feature, underline and nonvocal delimiters.
Two constructs are worth calling out because they behave unlike their
neighbours. A replaced word (word [: replacement]) is ReplacedWord
rather than an Annotated<Word>, because the replacement is part of the word’s
identity rather than a comment on it; it carries its own annotations alongside.
And a retrace’s annotations describe the retrace itself, not the material
inside it, which is why a retrace opens no language scope for the words it
contains.
Carrying none is a different variant, not an empty list
The bare and annotated spellings are different variants because they are
different things. hello is a word; hello [*] is a word plus a claim about
it. The model does not represent the first as the second with nothing in it.
This is enforced in the type rather than checked afterwards:
// The only public constructor. `None` when the list is empty.
AnnotatedContentAnnotations::new(annotations) -> Option<AnnotatedContentAnnotations>
So an annotated wrapper cannot be built without an annotation, and that
Option IS the bare-versus-annotated decision. Every place the parser builds
content, it reads:
match AnnotatedContentAnnotations::new(scoped) {
None => UtteranceContent::Event(event),
Some(scoped) => UtteranceContent::AnnotatedEvent(Annotated::new(event, scoped)),
}
TryFrom<Vec<_>> applies the same check, Deserialize rejects an empty list
off the wire rather than accepting one, and there is deliberately no Default.
The type also does not take the crate’s collection-newtype macro, whose take
and retain can empty a collection in place.
Why this is stated so emphatically
An invariant stated only in prose does not hold, so it is enforced by the
types. UtteranceContent has a bare Action beside its bare Event, and
BracketedItem has a bare Group; an action or group with no annotations
therefore has somewhere to go, and nothing is wrapped in an Annotated or
AnnotatedGroup carrying an empty list. The two content enums are symmetric,
and the empty state is unconstructible. The rule is not something a validator
looks for; it is something the compiler refuses.
No error code catches the empty case, because none can: bare [*] is valid
CHAT, and an empty bracket is a parse error. The reasoning is recorded in
Leniency Policy, Decision 1.
What an annotation attaches to when constructs nest
Scoping follows the innermost construct. In <the big dog> [//] [* m] both
annotations attach to the group. In <the [//] dog> the marker attaches to the
word inside it, because that is what precedes it.
One consequence matters for anything reading language: a <...> [@s:spa] group
opens a code-switch scope for the words inside it, and a retrace does not open
one at all. Tools should ask the model for the governing scope rather than
re-deriving it from the annotation list, because the two rules differ and the
difference is invisible if you get it wrong.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Symbols
Status: Reference Last modified: 2026-10-02 (commit 2d7e886b)
CHAT uses a rich set of symbols for transcription conventions. This
page documents the symbol categories and the symbol registry that
drives both the grammar and the Rust crates. The
symbol registry
(spec/symbols/symbol_registry.json) is the source of truth, when
this page and the registry disagree, the registry wins.
Symbol Registry
The authoritative symbol definitions live in spec/symbols/symbol_registry.json. This JSON file is the single source of truth, it generates:
- Character sets for the tree-sitter grammar (
grammar.js) - Rust constants for the model and validation crates
- Validation rules for the spec tool
After any change to the symbol registry, run:
just symbols-gen
Symbol Categories
Terminators
Punctuation that ends an utterance:
| Symbol | Name | Usage |
|---|---|---|
. | Period | Declarative |
? | Question | Interrogative |
! | Exclamation | Exclamatory |
+... | Trailing off | Incomplete utterance |
+..? | Trailing-off question | Question trails off |
+/. | Interruption | Speaker interrupted by another |
+//. | Self-interruption | Speaker interrupts self |
+/? | Interrupted question | Question interrupted |
+!? | Broken question | Exclamation-question |
+"/. | Quoted new line | Quotation continues on next line |
CA and Disfluency Symbols
The tables below are GENERATED from spec/symbols/symbol_registry.json, which
is the single owner of what each symbol means. The Rust types
CAElementType and CADelimiterType, the grammar’s character constants and
these tables all come from the same record, so they cannot disagree.
The category names describe a PARSING ROLE, not a provenance. A
ca_element_symbol attaches to a word token; a ca_delimiter_symbol brackets
a stretch. Ask a symbol’s notation_family() for provenance; never read it off
the name of the array the symbol sits in. That confusion is what once filed the
blocking and segment-repetition disfluency marks as Conversation Analysis
notation.
The Notation column is the symbol’s provenance and is independent of which
category it parses into. A symbol marked disfluency comes from the CHAT
manual’s Disfluency Transcription chapter and is not Conversation Analysis
notation; CLAN classifies those explicitly as NOT CA. They sit in the ca_*
categories purely because of how they parse.
Every example is parsed and validated by a test, so a row here cannot drift from what the grammar accepts.
Word-attached symbols (word_attached_symbols)
These attach to a word, so book↑ is a single token whose content
carries the symbol.
| Symbol | Codepoint | Meaning | Notation | Example |
|---|---|---|---|---|
⁑ | U+2051 | Hardening | CA | ⁑hello there . |
↑ | U+2191 | Shift to high pitch | CA | ↑hello there . |
↓ | U+2193 | Shift to low pitch | CA | ↓hello there . |
↻ | U+21BB | Pitch reset | CA | ↻hello there . |
≠ | U+2260 | Blocking, a word attack | disfluency | ≠hello there . |
∙ | U+2219 | Inhalation | CA | ∙hello there . |
∾ | U+223E | Constriction | CA | ∾hello there . |
⤆ | U+2906 | Sudden stop | CA | ⤆hello there . |
⤇ | U+2907 | Hurried start | CA | ⤇hello there . |
Ἡ | U+1F29 | Laugh inside a word | CA | Ἡhello there . |
Paired delimiter symbols (paired_stretch_symbols)
These are PAIRED: each opens and closes a stretch, and an unmatched one is rejected (E230).
| Symbol | Codepoint | Meaning | Notation | Example |
|---|---|---|---|---|
⁇ | U+2047 | Unsure transcription | CA | he said ⁇hello there⁇ today . |
§ | U+00A7 | Precise articulation | CA | he said §hello there§ today . |
⁎ | U+204E | Creaky voice | CA | he said ⁎hello there⁎ today . |
° | U+00B0 | Softer | CA | he said °hello there° today . |
↫ | U+21AB | Segment repetition, brackets repeated material that is NOT lexical | disfluency | ↫b-b-b↫boy ran away . |
∆ | U+2206 | Faster | CA | he said ∆hello there∆ today . |
∇ | U+2207 | Slower | CA | he said ∇hello there∇ today . |
∬ | U+222C | Whisper | CA | he said ∬hello there∬ today . |
∮ | U+222E | Singing | CA | he said ∮hello there∮ today . |
▁ | U+2581 | Low pitch register | CA | he said ▁hello there▁ today . |
▔ | U+2594 | High pitch register | CA | he said ▔hello there▔ today . |
◉ | U+25C9 | Louder | CA | he said ◉hello there◉ today . |
☺ | U+263A | Smile voice | CA | he said ☺hello there☺ today . |
♋ | U+264B | Breathy voice | CA | he said ♋hello there♋ today . |
Ϋ | U+03AB | Yawn | CA | he said Ϋhello thereΫ today . |
CA arrow separators
These are own-node separators between words rather than word-attachments, and
the parser splits them as their own nodes. They are NOT yet registry-owned, and
this table is still hand-written. They are not untyped: five of them are
Separator variants in talkbank-model, whose glyph table is hand-written
again in WriteChat and in several places across the grammar and the re2c
backend. Bringing them into the registry is the same move the two families
above have already made, and it is outstanding work rather than a decision.
| Symbol | Codepoint | Meaning |
|---|---|---|
→ | U+2192 | Level pitch contour |
↗ | U+2197 | Rising to mid |
↘ | U+2198 | Falling to mid |
⇗ | U+21D7 | Rising to high |
⇘ | U+21D8 | Falling to low |
↖ ↙ ← | U+2196, U+2199, U+2190 | Registered as separators; named in neither the CHAT manual’s symbol table nor CLAN’s symbol enum. |
Word Segment Characters
Characters that are forbidden at the start of words, forbidden in the rest of words, or forbidden throughout. These define the lexical boundaries of what constitutes a “word” in CHAT.
The grammar uses these sets to construct the word-matching regex patterns. Characters like [, ], <, >, (, ) are structural delimiters and cannot appear inside words.
Event Segment Characters
Characters forbidden in event descriptions (&=event content). Events have slightly different lexical rules than words.
Language Codes
CHAT uses ISO 639-3 three-letter language codes in @Languages headers and @s: word markers:
@Languages: eng, fra
*CHI: I want a croissant@s:fra .
Common codes: eng (English), fra (French), deu (German), spa (Spanish), zho (Mandarin), jpn (Japanese).
Special Markers
@ Markers (Word-Level)
The form-marker set has ONE owner:
spec/form_markers/form_marker_registry.json. The FormType enum, both
directions of its marker mapping, the re2c lexer’s code set and the table below
are all generated from it, so a marker cannot exist in one and not another.
| Marker | Meaning | Notes |
|---|---|---|
@b | Babbling | abame@b |
@c | Child-invented form | gumma@c, meaning sticky |
@d | Dialect form | younz@d, meaning you |
@f | Family-specific form | bunko@f, meaning broken |
@fp | Filled pause | um@fp, deprecated, use &-um instead, because filled pauses are excluded from grammatical analysis |
@g | General special form | gongga@g |
@i | Interjection, interaction | uhhuh@i |
@k | Multiple letters | abcd@k, mnemonic is “kana”: a Japanese kana is one symbol for a whole syllable |
@l | Letter | b@l, the letter b |
@ls | Letter plural | p@ls, the plural of a letter |
@n | Neologism | breaked@n, meaning broke |
@o | Onomatopoeia | woofwoof@o, a dog barking |
@p | Phonologically consistent form | aga@p |
@q | Metalinguistic use | if@q, as in no if@q-s or but@q-s, when citing words |
@sas | Sign and speech | apple@sas, signs and says apple |
@si | Singing | lalala@si |
@sl | Signed language | apple@sl, signs apple |
@t | Test word | wug@t |
@u | Unibet transcription | binga@u |
@wp | Word play | goobarumba@wp |
@x | Excluded words | stuff@x |
@z:<label> | User-defined code | word@z:rtfd, any user code |
Where the CHAT manual and chatter disagree, and why chatter is right:
@x: The manual’s Letters column writes@x:*, implying a label, but its own Example column writes barestuff@xand depfile.cut sanctions bare*@xbeside*@s:*and*@z:*.@x:foois rejected (E203); do not “fix” this to match the manual’s table.
Every meaning above is taken from the “Special Form Markers” table in the CHAT
manual, and each links to that marker’s own anchor there. The meanings are the
manual’s, not expansions of the letters: @k is “kana” (multiple letters), @p
a phonologically consistent form, @sl signed language, @sas sign and
speech, @g the general special form, and @ls the letter plural (the letter
sequence is @k). If you find another that disagrees with the manual, the
manual wins.
@a is not a form marker. The corpus authority eliminated it from every file
together with @e and @lp; it has no main-tier occurrences in any corpus,
and appears in neither depfile.cut nor the manual’s table.
The second-language qualifier @s:LANG is a separate construct (see
the L2 morphotag section of the Batchalign book); it is not part of
FormType.
& Markers (Events and Fillers)
| Prefix | Meaning |
|---|---|
&= | Paralinguistic event (e.g., &=laughs) |
&- | Filler (e.g., &-um) |
&+ | Phonological fragment (e.g., &+sh) |
&~ | Nonword (e.g., &~mama) |
&* | Other speaker’s speech event (e.g., &*MOT:word, speech attributed to another speaker) |
Scope Markers
| Marker | Meaning |
|---|---|
[/] | Partial retrace, speaker repeats the same words |
[//] | Full retrace, speaker restarts with different words |
[///] | Multiple retracing, multiple false starts |
[/-] | Reformulation, speaker rephrases with different structure |
[*] | Error |
[?] | Best guess |
[>] | Overlap follows |
[<] | Overlap precedes |
[= text] | Explanation |
[: text] | Replacement |
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Architecture Overview
Status: Current Last modified: 2026-08-21 13:12 EDT
TalkBank/chatter is the standalone home of the TalkBank CHAT specification,
tree-sitter grammar, Rust crates, chatter CLI, LSP server, and desktop app.
It is self-contained: the CHAT-format core builds and runs
without any external TalkBank repository, so downstream consumers can depend on
its crates directly.
Data Flow
Specification is the source of truth. Code is generated downstream from it.
spec/ Source of truth (CHAT specification)
↓
grammar.js Tree-sitter grammar (in grammar/)
↓
parser.c Generated C parser (never hand-edited)
↓
Rust crates Parser → Model → Validation → Transform
↓
Applications chatter CLI, LSP server, desktop app
Two layers
Within this repository, the architecture splits into two layers:
Source-of-truth artifacts. spec/, spec/symbols/, and grammar/ define
the CHAT language and generate downstream parser tests, error docs, and shared
symbol sets.
Consumer crates and applications. The Rust crates under crates/, the
chatter CLI, talkbank-lsp, and the desktop app all consume those
source-of-truth artifacts rather than defining CHAT semantics independently.
Crate Dependency Graph
flowchart TD
derive["talkbank-derive\nProc macros"]
model["talkbank-model\nData model, validation, alignment, errors"]
cache["talkbank-cache\nValidation + roundtrip cache"]
parser["talkbank-parser\nCanonical parser (tree-sitter)"]
re2c["talkbank-parser-re2c\nAlternate parser (equivalence oracle)"]
transform["talkbank-transform\nPipelines, CHAT↔JSON, caching"]
cli["chatter\nCLI: validate, normalize, convert"]
lsp["talkbank-lsp\nLanguage Server Protocol"]
s2c["send2clan\nCLAN app bindings"]
desktop["chatter-desktop\nDesktop validation app (Tauri)"]
tests["talkbank-parser-tests\nEquivalence tests"]
derive --> model
model --> parser & re2c
parser --> transform
re2c --> transform
cache --> transform
transform --> cli & lsp & desktop
s2c --> cli & desktop
parser --> tests
re2c --> tests
Repository Layout
chatter/
├── grammar/ Tree-sitter grammar
├── spec/ CHAT specification (source of truth)
│ ├── constructs/ Valid CHAT examples + expected parse trees
│ ├── errors/ Invalid CHAT examples + claims
│ ├── symbols/ Shared symbol registry (JSON)
│ ├── tools/ Core spec generators
│ └── runtime-tools/ Runtime-aware spec bootstrap/validation tools
├── crates/ Rust crates (model, parser, transform, CLI support, LSP)
├── corpus/ Reference corpus
├── tests/ Integration tests and fixtures
├── schema/ JSON Schema (auto-generated)
├── apps/chatter-desktop/ Desktop validation app (Tauri v2, React)
├── book/ This documentation
└── docs/ Strategy, proposals, investigations
Cargo Workspaces
Two separate Cargo workspaces live here:
- Root workspace (
Cargo.toml), all Rust crates for parsing, model, transform, CLI, LSP, andapps/chatter-desktop/src-tauri. - Spec workspace (
spec/Cargo.toml),spec/toolsfor core generation,spec/runtime-toolsfor runtime-aware spec tooling.
Use the relevant manifest path for the workspace you mean to operate in:
spec/tools/Cargo.tomlfor generatorsspec/runtime-tools/Cargo.tomlfor bootstrap/mining/runtime validation
Where to read next
For per-topic detail (sections being consolidated; see SUMMARY for the authoritative current list):
- Spec System, Grammar, Parser Backends, how CHAT becomes typed AST.
- CHAT model: the AST itself, content traversal, wide-struct rule.
- Alignment: tier alignment, DP, sequence alignment.
- Errors and validation: diagnostics, validation gates, and parser/model invariants.
- Editor/runtime integration:
talkbank-lspand application boundaries layered on top of the CHAT core. - Memory and Ownership, Type-Driven Design (lands during M11 errors-and-validation work).
For per-crate summaries see Crate Reference.
This page last changed: 2026-08-21 (commit b1084eaf). The whole book last changed: 2026-10-07 (commit 5e895791).
Spec System
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
spec/ is the source of truth for what CHAT is and for what chatter rejects.
Tests, fixtures and error documentation are GENERATED from it. You change the
spec; you do not hand-edit what it produces.
This chapter is the reference: what the spec files contain, what each field does, and what checks them. To make a change, follow Spec Workflow.
Start here: ask the system
Before reading further, run:
just spec-status
It reports, derived from the same code the gates use rather than from prose: how many specs exist and what status they declare, how many examples are verified, how many are deferred, how many assert nothing at all, the state of CLAN CHECK parity, and which gate checks which artifact. If this page and that command ever disagree, the command is right.
The two kinds of spec
Construct specs, spec/constructs/
A valid CHAT fragment and the tree it must parse to.
# languages_single
@Languages header with single language code
## Input
```languages_header
@Languages: eng
```
## Expected CST
```cst
(languages_header
(languages_prefix)
...
)
```
## Metadata
- **Level**: header
- **Category**: header
The Input fence label (languages_header, main_tier, utterance,
standalone_word, …) names a template in spec/tools/templates/ that
wraps the fragment in a complete CHAT file, because tree-sitter parses
documents rather than fragments. A label with no matching .tera template is
an error; add the template.
Error specs, spec/errors/
Invalid CHAT, and the codes it must produce. Everything declared lives in
+++ TOML frontmatter; everything published as prose lives in the body.
+++
code = 'E207'
name = 'Unknown scoped annotation marker'
[[example]]
source = 'E2xx_word_errors/E207_multiple_form_types.cha'
level = 'word'
claim = 'violates'
chat = '''
@UTF8
@Begin
@Languages: eng
@Participants: CHI Target_Child
@ID: eng|corpus|CHI|||||Target_Child|||
*CHI: word@zz .
@End
'''
+++
## Description
Unknown scoped annotation marker.
What every field actually does
The fields are not decoration. Each one changes what is checked. The
authoritative list, with types, is talkbank_spec_vocabulary::frontmatter,
which refuses an unrecognised key at load; this table says what each field
DOES, which a type cannot.
| Field | Effect |
|---|---|
code | The code the spec DOCUMENTS; names the generated tests, and is resolved against spec/codes/error-codes.toml at load, so a spec naming an unregistered code does not load. |
name | The human-readable title used in generated error documentation. |
status_note | A human’s adjudication of the code’s current state. Prose, published nowhere, read by people. |
example.chat | The input itself, a whole CHAT file. Required: an example without one is not an example. |
example.source | The fixture the example came from. Its stem NAMES the transcript, see below. |
example.title, example.notes | Prose about this example, read by people. |
example.claim | What the example asserts: violates, legal, or subsumed_by <code(s)>. REQUIRED, and both halves are enforced (absences included), see below. |
example.level | Where THIS example’s fault is (word, tier, utterance, header, file). Required per example: a code like E519 is violated at header level in one example and at utterance level in another, so the fault site is a fact about the example, not the code. The page’s Level line renders the distinct set. |
Two prose sections are published as well as read by humans:
| Section | Effect |
|---|---|
## Description | Published verbatim, markdown and paragraph breaks intact, as the page’s description. Required. |
## CHAT Rule | Published verbatim as the page’s ## CHAT Rule section: what CHAT requires, and therefore what a maintainer must write instead. Optional; a spec without one publishes no such section. |
## Expected Behavior, ## Notes | Prose for whoever opens the spec file. Read by no tool. |
Write the RULE in ## CHAT Rule, not a bare manual link. The pages exist so a
data maintainer can fix a file without reading the validator’s source.
kind and status are facts about a CODE, and live in the registry
Both are properties of the CODE, not of a document about it, so a code with
several spec files has one kind and one status. They live in
spec/codes/error-codes.toml, one entry per code, and a
spec reaches them through the code it names. Because there is one copy, there
is nothing to reconcile: the enum’s status and the specs cannot disagree.
The code registry
spec/codes/error-codes.toml is the source of truth for everything true of a
CODE, as opposed to true of a document about one:
[[code]]
code = 'E202'
variant = 'MissingFormType' # the ErrorCode variant it compiles to
summary = 'Missing form type on special word.' # the variant's rustdoc
kind = 'Invalidity'
status = 'implemented'
[[retired]]
code = 'W601'
reason = 'renumbered to E756 on 2026-07-16; the warning prefix was the bug'
crates/talkbank-model/src/errors/codes/generated_error_code.rs (the
ErrorCode enum) and generated_diagnostic_kind.rs are both GENERATED from
it, and both are under the currency gate, so the enum cannot disagree with the
specs about which checks run.
The schema, with a reason per field, is
talkbank_spec_vocabulary::registry. It refuses, at load: an unrecognised key,
a code registered twice, two codes compiling to one Rust identifier, and a
RETIRED number brought back; the load error names the retirement’s own
recorded reason. Reusing a retired number such as W210, W601 or E754 is
therefore unrepresentable.
What the registry deliberately does NOT own is whether a code is DOCUMENTED.
That is a coverage question, and error_code_specs asks it as one: a variant
with no spec file is a coverage gap, not a vocabulary divergence.
A field has no position
An example is one value that carries its own input and its own declared fields, so there is no fence for a field to be on the wrong side of, and no placement rule to remember. It is the clearest example in this system of a type removing a rule rather than a document restating one.
Every example carries a CLAIM, and absences are assertable
Each example declares one of:
claim = 'violates' # the spec's code MUST appear
claim = 'legal' # the spec's code MUST NOT appear
claim = { subsumed_by = 'E316' } # E316 appears; this code does not
claim = { subsumed_by = ['E246', 'E249'] } # all listed appear; this code does not
The claim is REQUIRED: an example that asserts nothing is unwritable.
Extra emitted codes are fine (one malformed line legitimately raises several
diagnostics); the exact per-stage sets are the observation snapshot’s
business. The NEGATIVE half is assertable: legal and the own-code-absent
part of subsumed_by assert that a code is NOT emitted. A spec whose examples
are all subsumed_by is the parser-specificity worklist, verifiable against
the snapshot; coverage --errors lists it.
Reading a spec: a subsumed_by claim tells you what chatter emits, not
necessarily what the rule is. Read the spec’s Description and title for the
rule, and treat a mismatch between the title code and the example’s codes as
an open question rather than as a specification. subsumed_by E316
(“unparsable content”) means the example does not parse, so the specific rule
is never reached: a parser gap. A specific other code is usually a wrong
fixture. When authoring, an example whose input violates the rule should
produce that rule’s code; if it does not, chatter has a gap, and the gap is
the finding.
There is no layer field, and the runner is total
Which stage catches a rule is not authored; it is an OBSERVATION, recorded per
example in spec/observations/example-diagnostics.json. Every example is a
fixture, and the fixture runner collects BOTH stages’ codes against a real
file, so there is no stage a declared code can hide in (some examples’ codes
are genuinely SPLIT across stages, which no per-stage harness could assert).
Error specs have no string-based per-stage tests; the fixture runner plus the
observation snapshot cover them.
Tree-sitter corpus membership is derived from the snapshot: an example joins iff it produced parse-stage diagnostics, so there is structure to pin.
status controls implementation checking, not every regression contract
| Value | Effect |
|---|---|
implemented | Examples are verified. |
not_implemented | Own-code implementation is DEFERRED and generated tests carry #[ignore]; explicit legal/subsumption claims remain checked separately. |
deprecated, unreachable_from_chat | Implementation is deferred; legal/subsumption regression claims remain checked. |
| absent | REFUSED: spec/codes/error-codes.toml fails to load, naming the entry. |
Declared per CODE, in the registry, with no default: implemented means
someone decided it once for the code, rather than each of its spec files
claiming it separately.
Changing a spec from not_implemented to implemented un-#[ignore]s its
generated tests, and those tests may never have run. Regenerate and run them in
the same change.
Deferred code status is not a count of missing validation rules. Run
cargo run --manifest-path spec/Cargo.toml --bin spec_status -- --deferred
to see each authored claim beside its observed codes. This view distinguishes
verified legal/subsumption claims, contradicted claims and planned violation
claims using the claim’s shared evaluator. The separate deferred-spec regression
test enforces legal/subsumption claims even when their code is not
implemented. A verified alternate diagnostic does not reactivate that code;
an observed emission of the deferred code instead requires status review.
source names the transcript
Some CHAT rules are about the file’s own name: E531 requires the @Media
header’s filename to match the transcript’s stem. The example runner therefore
names each transcript after the stem of its source, and an example with no
source is anonymous, so those rules do not run for it.
The backend-parity harness also preserves this context: its input owns both CHAT text and the declared source path, and performs contextual validation for either parser. A measurement regression checks mismatching, matching, and anonymous source names through both backends.
A legal claim asserts absence of this spec’s own code. It does not assert that
other rules accept the input or that parsing required no recovery. For example,
E758’s malformed-content controls retain their content diagnostics while proving
there is no space directly after the tab. Serialization-equivalence assertions
must distinguish clean parsing from recovery.
The observation snapshot
spec/observations/example-diagnostics.json (generated, gated) records, for
every example of every spec, the exact diagnostic codes the current binary
produces, split by the stage (parse or validation) that emitted them. It
covers every spec regardless of status, because an observation is not an
assertion: for an unimplemented rule the honest record is “nothing fires”.
It is the regression instrument for the spec suite: a diff in this file
is a review event, and every changed entry is adjudicated INTENDED (the
behaviour change was the point; commit the regenerated snapshot in the same
change) or UNINTENDED (a regression; fix the code, never the snapshot). It is
also what makes a subsumed by claim verifiable and what the layer-of-capture
question is answered from, observed rather than authored.
What is generated, and by what
One command regenerates everything committed: just spec-gen. Its registry
(spec/tools/src/artifacts.rs, plus the half in spec/runtime-tools that needs
the live ErrorCode enum) is the only place a destination is written down, and
the same list drives writing, checking and the gate.
| Artifact | Committed at | Directory ownership |
|---|---|---|
| example-diagnostics observation snapshot | spec/observations/ | the whole directory, cleared on every run |
| talkbank-model’s generated sources | crates/talkbank-model/src/errors/ | only the files it produces |
| tree-sitter corpus tests | grammar/test/corpus/generated/ | the whole directory, cleared on every run |
| generated Rust test bodies | crates/talkbank-parser-tests/tests/integration/generated/ | only the files it produces |
| published error documentation | docs/errors/ | the whole directory, cleared on every run |
| validation fixture corpus + manifest | crates/talkbank-parser-tests/tests/error_corpus/validation_errors/ | the whole directory, cleared on every run |
| the book’s artifact table | book/src/architecture/generated/ | only the files it produces |
That table is itself generated from the registry, and the currency gate keeps it true.
One generator sits outside it, deliberately: gen_form_markers has its own
registry and its own drift gate (just form-markers-gen).
docs/errors/*.md is a tracked, committed registry artifact like any other:
just spec-gen writes it and just spec-check compares it.
Two registries under spec/ own closed vocabularies and generate every site
that names them: spec/symbols/symbol_registry.json (just symbols-gen) and
spec/form_markers/form_marker_registry.json (just form-markers-gen). Each
has its own README and its own drift gate.
Shared-directory artifacts (Ownership::NamedFiles) retain byte-identical
outputs and delete only explicitly retired filenames. This preserves generated
Rust inputs across no-op regeneration while leaving other producers’ files
alone. The generator reports the number of files actually written, not the
number it expected to produce. Whole-directory artifacts additionally require the ownership capability
described below before obsolete files may be pruned.
Generated and hand-written tests live in separate trees.
grammar/test/corpus/generated/ retains unchanged files and removes obsolete
ones through GeneratedDir, which requires a .generated-output-dir marker
and refuses human ownership or symlinked entries;
grammar/test/corpus/manual/ is never written by a generator, so a
regeneration can never destroy hand-mined corpus tests.
What checks what
| Gate | Checks | Needs |
|---|---|---|
every_generated_artifact_is_current | every committed generated artifact against what the specs produce now | |
error_spec_codes | every example emits the codes it declares | |
manifest_agrees_with_clan_reference | parity manifest against check.cpp | |
generated_form_marker_sites_are_current | form-marker outputs against the registry | |
generated_symbol_sets_are_current | symbol-set outputs against the registry | node |
clan_check_grounding | fixtures against the REAL CLAN binary | CLAN, CHATTER_CLAN_RUN |
The first four run in CI under
cargo test --manifest-path spec/Cargo.toml --workspace. clan_check_grounding
is #[ignore]d and catches UPSTREAM drift; refresh-unix-clan.sh runs it after
a successful CLAN sync, which is the moment it matters.
CLAN CHECK assessment
See the generated CHECK Assessment
for the current inventory, adjudications, architectural mandate, completion
limits and rules for reopening an obligation. Its authored manifest is the
single authority; this chapter deliberately does not repeat the verdict scheme
or completion claims. just spec-status derives its assessment summary from
the same shared types and manifest. Neither report substitutes for a fresh
runtime observation of a specific CHECK executable.
Related
- Spec Workflow, how to make a change.
- Testing, the wider test strategy.
- Grammar Governance, the grammar side.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Grammar
Status: Current Last updated: 2026-03-24 00:01 EDT
The CHAT grammar is defined in grammar/grammar.js using the tree-sitter parser generator. It produces a GLR parser that handles the full CHAT format with error recovery.
Design Principles
Explicit Whitespace
Unlike most tree-sitter grammars, CHAT does not use extras for whitespace. All whitespace is grammar-visible because CHAT’s structure is whitespace-sensitive:
- Tab separates tier prefix from content
- Newline ends tiers
- Line continuation uses tab-at-start-of-line
- Space separates words and annotations
Two-Level Structure
The grammar has two structural levels:
- Document level: headers, utterances,
@Begin/@End - Tier level: main tier content, dependent tier content (each with distinct rules)
Opaque Lemmas
In the %mor tier rules, lemmas are parsed as opaque Unicode strings. The grammar does not attempt to decompose lemma content, that happens in the model layer. This follows the “parse, don’t validate” principle.
Key Grammar Rules
Document Structure
document → utf8_header, begin_header, lines..., end_header
line → header | utterance
utterance → main_tier, dependent_tiers...
Main Tier
main_tier → star, speaker, colon, tab, tier_body
tier_body → contents, utterance_end
contents → content_item, (whitespace, content_item)...
MOR Tier (UD-style)
mor_contents → mor_content, (whitespace, mor_content)..., terminator
mor_content → mor_word, mor_post_clitic*
mor_word → mor_pos, pipe, mor_lemma, mor_feature*
mor_post_clitic → tilde, mor_word
mor_feature → hyphen, mor_feature_value
POS tags are simple identifiers (no subcategories). Lemmas are opaque strings. Features are hyphen-separated values that may contain = for Key=Value pairs and , for multi-value features.
Grammar Change Workflow
parser.c is generated from grammar.js, never edit it directly.
After any change to grammar.js:
cd grammar && tree-sitter generatetree-sitter test(160 tests)cargo test -p talkbank-parsercargo test -p talkbank-parser-tests(reference corpus equivalence, per-file)- Verify the 78-file reference corpus passes at 100%
Conflict Resolution
The grammar uses tree-sitter’s precedence and conflict mechanisms to handle ambiguities:
- Word tokens use
prec(5)to win over separators - Inline bullets use
prec(10)for their delimiters - CA (conversation analysis) symbols use
prec(3)for colon disambiguation
Generated Artifacts
Running tree-sitter generate produces:
src/parser.c: the C parsersrc/node-types.json: node type metadata
crates/talkbank-parser/src/node_types.rs is GENERATED from it, by
scripts/generate-node-types.js, which also reads scripts/node-type-docs.json
for the per-constant doc comments:
just node-types-check # regenerate and fail if the committed output moved
The doc file is the hand-written half and the only part a human edits; the kind
list comes from the grammar. Both are gated, because both have drifted: the
script was left behind in the repository chatter was extracted from, so for
three months node_types.rs carried a “DO NOT EDIT, auto-generated by
scripts/generate-node-types.js” banner naming a script this repo did not have,
and it drifted to 8 missing kinds and 6 constants for kinds the grammar no
longer had.
This page last changed: 2026-08-27 (commit 8b445304). The whole book last changed: 2026-10-07 (commit 5e895791).
Overlap Marker Binding
Status: Current Last modified: 2026-07-10 12:06 EDT
Overlap markers (⌈ ⌉ ⌊ ⌋) mark simultaneous speech. They are
the hardest tokenization problem in CHAT, because they legitimately
live at two levels: between words (marking a span boundary in the
utterance) and inside words (marking that the overlap boundary
falls mid-word, as in o⌈ne t⌉wo). This page documents the binding
rule the grammar ships, the ideal rule it approximates, why the gap
exists, and the measured options for closing it. It is the permanent
record of a design debate that has run since the project’s earliest
prototypes; read it before touching word_body, contents, or
anything overlap-adjacent.
The shipping rule: adjacency binds into the word
A marker adjacent to text is part of the word; only a
space-separated marker is a standalone overlap_point.
Yeah ⌈2 hey ⌈2 is standalone (spaces both sides)
⌈one two⌉ TWO words: "⌈one" and "two⌉" (markers bound in)
o⌈ne t⌉wo TWO words with interior markers (same rule)
Mechanically: overlap_point and word_segment carry equal token
precedence, and maximal munch plus word_body’s continuation rules
give the word custody of every adjacent marker. See
Word Internals (tokenization
ambiguity #1) and the grammar’s tokenization-rules.md (Exception 1).
The ideal rule this approximates
A marker with spoken text on BOTH sides is word-internal; a marker
at a word’s edge is top-level content. Under the ideal rule,
⌈one two⌉ parses as (⌈) (one) (two) (⌉): the visually obvious
reading: while o⌈ne keeps its interior marker. The shipping rule
diverges exactly at word edges, where it gives the word custody of
markers the ideal calls top-level.
The ideal was the project’s ORIGINAL specification. Early prototype grammars (January 2026) attempted it and produced a substantial decision record; the analysis concluded that the rule “requires bidirectional context that LR parsers cannot naturally handle” (the parser must see both sides of the marker to classify it, and LR(1) has one token of lookahead). The adjacency rule was adopted as the tractable alternative, and every grammar generation since (through the February coarsening campaign and the March re-structuring) has carried it forward.
What 2026-07-10 established
A feasibility experiment revisited the impossibility conclusion with
GLR machinery the January analysis had not combined: an interior-only
word_body (a word may not begin or end with an overlap marker), a
declared conflict, dynamic precedence on interior continuations, and
removal of the static prec.right bias so the conflict genuinely
splits. Results:
- The ideal rule IS expressible: probes and grammar fixtures parse to the ideal shapes with no ERROR nodes, and the grammar’s conflict inventory net-shrinks.
- Corpus reality is the hard part. Conversation-analysis (CA)
transcription layers: overlap points, paired CA delimiters,
underline spans, lengthening, compounds: cross-nest freely at word
edges (
☺you ⌈there⌉☺,∇⌈ho:ney⌉∇,full⌉+grown,⌈drug⌉ [!]). The shipping rule sidesteps every such case by giving the word custody of everything adjacent; the ideal rule must answer a custody question PER MARKER PAIR, each answer costing a grammar rule, a conflict, and an AST-shape decision. Measured against the full kept corpus (763 overlap-bearing files, all of which parse cleanly under the shipping grammar), five iterations of custody rules reduced ideal-rule regressions from 195 files to 105: a converging but long tail.
Two implementation routes therefore exist:
- Grammar route: finish the custody enumeration. Honest estimate: a multi-week grammar project, followed by AST migration across the model, the generated visitor, and the second (oracle) parser.
- Conversion route (recommended by the experiment): keep the shipping grammar, and re-associate edge-bound overlap points to top level during CST-to-model conversion. At that point the CA layers are already resolved into typed word children, so every custody question becomes a deterministic tree transformation rather than a GLR fight. Precedent: CA terminator promotion, which already uses this parse-one-way/normalize-at-conversion pattern. The grammar’s empty-extras design (all whitespace grammar-visible) preserves exactly the facts the transformation needs.
The choice between routes (or deferral) is an open maintainer decision at the time of writing; this page must be updated when it is made.
Why this interacts with whitespace separation
The grammar’s deepest design commitment is extras: []: all
whitespace is grammar-visible, because the worst historical CHAT
parser bug was ACCEPTING glued content items as if properly separated
(the legacy Java implementation tokenized hello(.) correctly as a
word and a pause, and that silent acceptance was precisely the
problem: malformed sources never got cleaned). Overlap markers and a
short list of negotiated exceptions (notably comma-left: one, two
is accepted; one ,two is not) are the only constructs that
legitimately juxtapose with words at all. Whitespace-separation
violations that the grammar tolerates for recovery’s sake are rejected
by validation with precise diagnostics (E749, E750, E751), per the
layer rule: the grammar’s job is SHAPE (parse everything, truest
tree); rejection of recoverable style belongs to validation, where
messages are helpful and recovery graceful.
This page last changed: 2026-07-27 (commit d65b7eee). The whole book last changed: 2026-10-07 (commit 5e895791).
Parsing
Status: Current Last updated: 2026-10-07 (commit 5e895791)
The parsing pipeline converts CHAT text into a typed ChatFile AST.
ParseError::internal constructs E001 producer-failure evidence; it must not
classify a tool fault as invalid CHAT, even after presentation severity changes.
Diagnostic display maps own their text and keep the zero-byte origin implicit.
Only post-origin breakpoints are stored; an empty table is an identity map,
so no caller can omit a required origin sentinel. Display events consume source
bytes monotonically, preserving UTF-8 byte-coordinate mapping.
Display-span normalization bounds incoming source coordinates before offset
arithmetic, so even extreme out-of-source spans clamp safely to display EOF.
Source-span overlap uses half-open byte ranges: an empty insertion span never
overlaps another span, even when its point is inside it. Containment of an
insertion point is a separate operation. Recovery deduplication retains its
explicit same-point/same-code rule rather than treating a point as covered bytes.
Line-end lookup also checks index arithmetic before accessing its line table:
even an extreme out-of-range line request uses the documented source-EOF clamp.
Header lowering rejects an unreadable comment body or participant name/role
word after reporting the internal source-binding failure. It does not invent
an unknown comment or reinterpret a shortened word sequence as a participant.
This is distinct from structural CHAT recovery, whose diagnostics remain intact.
Unsupported-tier lowering likewise reports E001 if a generated anonymous body
span cannot be read from its source. It does not report E330 or substitute empty
content: an unreadable span is a tool failure, whereas authored empty content
remains represented in the model for validation.
The shared transitional raw-node text reader also reports unreadable ranges as
E001 through SourceBindingError::InvalidRange. Freecode, event and
grammatical-relation fields use that one boundary rather than attributing a
failed source read to malformed CHAT. This reader proves readable coordinates,
not tree/source identity; generated source-bound admission remains the stronger
contract. Structural recovery and lexical validation retain their own findings.
Postcode decoding requires SourceBound<PostcodeNode>: final-code groups
project their generated source fields and admit the selected token range before
decoding. The decoder takes no independent source string and retains the token’s
exact span. First and repeated groups share one leaf-slot decoder; displaced
diagnostics retain their existing per-group order. Error/Absent recovery still
contributes no postcode, while source-field failures propagate as CstFailure
rather than silently dropping a code. The ownership test rejects a node from an
independent equal-text parse; shared source admission owns range refusal.
User-tier parse-health classification consumes a source-bound tier and admits
its prefix before comparison; a failed read cannot become successful “no domain”
classification. Childless separator recovery rejects unreadable text through
the source-bound entry boundary, then reads admitted text without a second
decode. Its internal structural extraction still retains all recovery states.
Fragment admission and recovery likewise distinguish
binding/range failures (E001) from missing or malformed CHAT, retaining fragment
coordinates in the reported context.
Main-tier word/replacement wrappers, annotated angle groups, quotations and
phonology/sign groups consume the same compiled canonical-grammar admission.
Their composite word, replacement, annotation, quotation and contents slots
cannot themselves be Missing. The shared contents walker still handles lexical
and mixed-choice recovery; delimiters, absent required content, ERROR nodes,
source failures and empty-group refusal retain their established policies.
Carrier-owned recovery reporting accepts the corresponding admitted group
carriers, preserving body-region diagnostics for displaced children. This is
not a clean-tree proof and does not permit dropping recovery inside a composite.
Standalone-word and word-body extraction also carry canonical grammar admission through prefix, form-marker and segment/overlap-initial piece choices. The required composite word-body slot cannot itself be Missing. Its contents remain a mixed choice, so Missing, Error and Absent handling still applies there and to lexical pieces. The shared piece conversions, suffix policies and source-read failures are unchanged; admission does not certify a recovered word as valid. Replacement extraction uses the same proof for its first and repeated word slots. No word-position counter is needed to describe an impossible composite Missing state. Zero-width words, unreadable source, delimiter recovery, Error/Absent slots and refusal of an empty replacement remain separate checks.
MOR item collection has a private one-way admission state. Items and failure health cannot be supplied independently to finalization. Any item, separator or terminator failure rejects the collection; subsequent traversal still reports diagnostics but cannot revive it. Finalization handles that rejection before checking for an absent terminator, preserving diagnostic precedence. Source failures and lexical/structural recovery remain explicit.
The low-level @ID header entrypoint likewise takes a
SourceBound<IdHeaderNode> rather than an independently supplied node and text.
Header, contents and required/optional field projections retain that source
association. Each present field still admits its readable range; failure flows
as an internal producer error rather than an empty field or invalid CHAT.
Required-field rejection, omitted optional values, displaced-node reporting and
the whole-tree recovery backstop retain their policies. Optional field reading
returns Result<Option<String>, CstFailure>: it cannot separately claim a
CHAT rejection after a successful read. This is not a new grammar, a whole-tree
range prepass, or a change to the public ChatParser trait’s ID-fragment API.
All six dedicated structured-header dispatch paths (Languages, Participants,
ID, Media, Situation, Types) retain source-bound nodes through one
shared diagnostic-forwarding wrapper. No path in this family accepts a separate
source argument. Situation text and the three Types fields are read through
source-bound projections; unreadable ranges are internal failures, not missing
authored text. Missing-field recovery, Types field order and its first-missing-
field short-circuit are unchanged. This does not assert readable child ranges
before their checked reads or certify recovered headers as valid.
The low-level %mor tier entrypoint takes
SourceBound<MorDependentTierNode> from the existing parse owner, with no
independent source argument. Dependent-tier dispatch retains this capability
through tier/contents extraction, choice and repeat projection, and item
admission. A private admission type separates readable items, absence and
reported failures; displaced nodes still use the shared recovery collector.
Word, post-clitic and feature decoding also retain source-bound nodes, and
POS/lemma/feature text comes from admitted fields rather than a second UTF-8
decode. Source-binding/reconstruction faults propagate as CstFailure to the
owning item boundary, without becoming missing POS/lemma diagnostics.
String-based fragment APIs are unchanged. Tier-body, main morphology-word,
post-clitic, contents-item and repeated feature boundaries consume canonical-grammar admission: their
composite-node Missing state is uninhabited in the generated slot type. This
does not prove that required children exist or lexical payloads are nonempty.
Error, absent-child, displaced-node and lexical missing-placeholder recovery
remain. Structural delimiter recovery still uses the shared slot helpers with
the owner’s source. POS, lemma and feature-value tokens still permit Missing;
the enclosing composite feature does not. Source ownership alone would not
justify this narrowing.
The contents alternative remains a mixed choice with its own Missing/Error/Absent
handling. Both admitted terminator choices use the shared exhaustive terminator
conversion; admission does not invent a new terminator or whitespace policy.
The low-level %gra entry likewise requires a source-bound node. Dispatch,
tier/contents extraction, repeated pairs, relations and field text retain the
same source owner. Index, head and label admission retains its existing order
and numeric/empty checks; source faults propagate as CstFailure to dispatch
instead of dropping a relation and misclassifying a producer failure as CHAT
invalidity. Recovery and declared-versus-lowered completeness remain explicit.
String-based grammar-tier and relation fragment APIs are unchanged.
Compiled-grammar admission narrows the tier-body and repeated relation
carriers: composite gra_contents and gra_relation cannot be Missing.
Error/Absent still preserve the existing truncated-body policy; lexical
whitespace, index, head and label recovery remain. Declared-versus-lowered
relation accounting still comes from the same traversal and is not inferred
from this composite-node proof.
The low-level %sin entry carries source ownership through group choices,
repeats, tokens and whole-group fallback. Internal source failures propagate to
dispatch as CstFailure; they cannot become an empty token list or trigger
recovery fallback. Existing empty-token and whole-group recovery policies remain.
This does not alter %phoaln/%phoint or the string fragment APIs.
The shared %pho/%mod decoder also retains source ownership through tier
selection, repeated groups, compound words and whole-group fallback. Its typed
tier variant determines both the extractor and model tag. Low-level entry points
accept SourceBound and return CstFailure; internal faults propagate to dispatch,
never into empty-content or recovery branches. Compound spelling, empty-content
omission, Missing/Error/Absent handling and string fragment APIs are unchanged.
Both families consume compiled canonical-grammar admission at their tier,
group-list, group-choice and grouped-content boundaries. Generated
NonMissingKindSlot carriers rule out Missing for composite bodies, groups,
grouped content, pho_words and sin_word; this is a producer proof, not an
inference from an error-free tier or a lack of failing fixtures. Outer choice
Missing, Error, Absent, displaced nodes and lexical whitespace Missing remain
explicit. Whole-group fallback and empty-token omission are unchanged, as are
the public APIs. Grammar/source admission failures propagate as internal
failures rather than taking any CHAT recovery branch.
%wor marker construction retains its already-admitted source-bound comma,
tag and vocative nodes, so it has no second fallible byte decode that could
silently drop a separator. Its transitional language-token reader and raw
dependent-tier body reader use checked range admission and report internal
failure before lexical or empty-content policy is applied.
The tier, body and word-item extractors also consume compiled canonical-grammar
admission: body, language-code and standalone-word composite slots cannot
themselves be Missing. Mixed word/bullet/marker choices, lexical whitespace and
terminators still retain their recovery states; Error, Absent and displaced
children remain explicit. The admitted terminator choice uses the same exhaustive
mapping as other tiers. No clean-tier assumption substitutes for those types,
and no timing, empty-tier or public API policy changes follow from this proof.
The free-text dependent tiers (the nine bullet-payload tiers, the fifteen
raw-text tiers and %x user-defined tiers) also extract under compiled
canonical-grammar admission. Each body is a named nonterminal, so its slot is a
SelectedNonMissingKindSlot and the body readers have no Missing arm: before
the proof reached them they read a placeholder body’s empty text as content, a
path the grammar can never take. Their tier_sep slot is narrowed too, and the
separator reader accepts it under either kind proof because it reads only a
present separator. Error, an absent optional body and displaced children keep
their diagnostics. Header and field readers whose callers hold both narrowed
and plain kind slots stay generic over the slot’s Missing payload
(AnyKindSlot); a placeholder there only refuses, so one body serves every
reading.
Misplaced-linker diagnostics likewise retain source-bound token text instead of
substituting a grammar name. Recovery collectors and main-tier language precodes
report unreadable coordinates as internal failures; an unreadable word-recovery
range is not an invalid control character and must not advise editing the CHAT.
The same distinction applies to generic, dependent-tier and utterance recovery
admission. An unreadable range produces an internal-failure diagnostic before
the readable-recovery classifier can run; utterance admission still taints every
potentially affected alignment domain. This does not change the diagnostics for
readable malformed CHAT or turn recovery into validity.
Structured-tier recovery keeps its generated SourceBound through
source_slice() and source-associated child traversal. The recursive reporter
accepts no independent source string and passes admitted recovery text directly
to the classifier. It retains ERROR/MISSING traversal boundaries, tier context,
diagnostic order and explicit internal failure on an unreadable child range.
Utterance recovery slots likewise retain their generated SourceField until
read admission. Their reporter accepts no independent node/text pair; equal
source bytes from another parse do not establish ownership. Unidentified
recovery still taints every potentially affected alignment domain, and failed
range admission remains an internal failure rather than CHAT invalidity.
Gem-label and option-name accumulation also rejects a failed source read rather
than returning a shortened label or incomplete flag set. Their outcome types
distinguish rejected admission from a genuinely absent label or empty options.
Participant entry lowering consumes the generated range-admitted carrier:
speaker/name/role slices are checked before decoding and read without another
fallible source operation. Missing/Error/Absent states, name order, role admission
and displaced recovery diagnostics remain. Admission failure reports internal
failure at the entry boundary and rejects the entry; it cannot expose a partial
name whose final word would be mistaken for a role. Only displaced-node evidence
is copied across the consuming transition, then released on success.
This opt-in costs storage: reference code+role and code+name+role entries perform
3/5 checks; repeated reads add none. Inline carrier storage grows 224 to 288
bytes, with 200 to 264 occupied bytes per repeated word group on the measured
64-bit build. The admitted owner occupies 360 bytes. These are transient layout
figures, not peak-memory or throughput claims; malformed/recovery entries may
also need storage for retained displaced evidence. Public APIs and grammar are
unchanged.
Gem headers retain SourceBound through their optional separator/text group
and free-text choices. Label lowering accepts no independent source string;
admitted text pieces need no second byte decode. Bare markers, missing/error
positions, continuation spacing and displaced-child reporting retain their
existing semantics. A node from an independent parse cannot enter this lowering
through another parse’s owner, even when both inputs have identical bytes.
Unknown-header recovery preserves the authored header text, never a grammar-node
name substituted after a failed read. HeaderSite retains its admitted text;
transitional raw-node construction returns SourceBindingError and its callers
propagate internal failure. A rejected header fragment does not construct a
throwaway unknown-header model before returning its diagnostics.
The shared bullet/text-tier reader similarly propagates CstFailure through
its typed adapters to dependent-tier dispatch. A failed source-bound read cannot
become a successful empty payload. Authored empty optional bodies still lower
as empty content for validation, and malformed slots retain recovery diagnostics.
Nested segment reads use the same fallible transition: text, leaf and missing-node
source admission must succeed before accumulation can return BulletContent.
A failed read rejects that content instead of silently omitting a segment.
Inline bullet lowering likewise propagates reconstruction failure separately
from malformed timestamp rejection; ordinary bullet recovery is unchanged.
The %wor timing adapter preserves its existing alignment-owned rejection
policy for malformed timestamps, but explicitly reports reconstruction faults
as internal failures instead of silently discarding them with .ok().
Main-tier body lowering preserves its source-bound ending through extraction.
Failure to read the body content or ending propagates as CstFailure; it cannot
produce a successfully constructed empty body. Structural recovery remains
separate from these producer failures.
All structured timestamp consumers pass SourceBound<BulletNode> to the
shared reader. Start/end fields retain that association through generated field
projection and checked reads. There is no independent source-string argument
to mismatch with a bullet, and unreadable fields report producer failure rather
than masquerading as missing timestamps. This changes neither overflow nor
leading-zero policy.
The default and canonical parser is the tree-sitter parser
(talkbank-parser). A second implementation, talkbank-parser-re2c,
exists alongside it as an experimental, incomplete alternative, not an
authority for CHAT validity. It targets the same ChatFile model and is opt-in via
chatter validate --parser re2c. The LSP and all production paths
default to the tree-sitter parser.
The finite reference workflow checks editor-facing source ownership after lowering: byte offsets at main/dependent-tier starts and interiors select the owning utterance, while headers, EOF and a turn’s exclusive end do not select that finished turn. These are byte coordinates, including multiline tiers, not character indices. This does not prove that source-free reconstructed ASTs have usable spans.
Model access is also distinct from validity. The missing-ID E522 spec still
exposes both declared speakers through declared_speakers(), but the speaker
without an ID has no id_metadata(). The participant-map join contains only
the materialized ID record. Diagnostics remain present; a UI may display the
declaration without fabricating metadata or certifying the file as valid.
Tree-Sitter Parser
Generated recovery capabilities should survive through diagnostic helpers.
For example, the empty-colon reporter requires KindMissing<ColonNode>, not
an arbitrary colon node followed by a zero-width test. The producer already
distinguishes a present token from a missing placeholder. Retaining that state
does not establish that CHAT fixtures reach it: deleting a speaker colon still
produces generic E316 recovery, and E322 remains deferred.
An ERROR child in main-tier contents is recovery evidence, not a word suffix.
The parser reports its source-associated location without appending guessed
text to a preceding word, annotated word or replacement. An @ prefix or a
marker-like character sequence cannot establish lexical ownership. The
whole-tree recovery backstop remains active; removing that text heuristic does
not remove recovery diagnostics or certify a recovered model as valid.
The canonical corpus contract checks the main-tier word traversal against both reference files and error-spec fixtures: each retained word’s raw spelling equals its own UTF-8 source span, including in documents with parse errors. This verifies source ownership, not validity of those recovered documents or exhaustive coverage of all possible recovery trees.
The talkbank-parser crate wraps the tree-sitter C parser and converts its concrete syntax tree (CST) into the ChatFile model.
Full-file parsing is the canonical entry point. TreeSitterParser also
provides fragment methods (parse_word_fragment(), parse_main_tier_fragment(),
parse_chat_file_fragment(), etc.) for parsing isolated CHAT fragments
directly.
CST → AST Pipeline
flowchart LR
chat["CHAT text\n(.cha file)"]
grammar["tree-sitter grammar\n(grammar.js → parser.c)"]
cst["Concrete Syntax Tree\n(all whitespace preserved)"]
walker["TreeSitterParser\n(CST traversal)"]
ast["ChatFile AST\n(semantic model)"]
chat --> grammar --> cst --> walker --> ast
Source text
↓ tree-sitter parse
Concrete Syntax Tree (CST), green tree with all tokens
↓ tree_parsing (Rust)
ChatFile AST, typed model with validation-ready data
The CST preserves every character of the source (whitespace, punctuation, comments). The Rust tree-parsing modules extract semantic information from the CST into the typed model through a generated typed traversal layer, described next.
Source-preserving document and participant lowering
Document classification retains generated source-bound capabilities for both complete and recovered documents. Associated child projections carry the producer’s source through lines and header dispatch without repeated root-membership searches. Header fragments use the same dispatcher with their already admitted wrapped-source slice.
Source association is not a proof that all runtime ranges are readable. The
incremental API still accepts a raw old tree, whose caller must apply every
source edit before reuse. The finite-corpus boundary test deliberately reuses
an unedited reference tree after deleting its input: parsing completes, but
root admission rejects InvalidRange; a cold parse of the same empty input
has a readable root. This is an API-contract violation, not a CHAT syntax rule.
Keep range refusals until the producer models the edit/source relationship;
neither a bound parent nor unreached fixture branches establishes that proof.
The editor uses TreeSitterParser::parse_chat_file_revision and its opaque
ParsedRevision cache. Only the parser constructs that cache from the source
it actually parsed. The next transition accepts new text and a previous
revision, derives a UTF-8-safe edit from the retained baseline, edits a cloned
tree, and parses it. The LSP owns no independent edit calculator and
stores no separately replaceable source/tree pair. Detached trees remain
available for structural queries but cannot be installed back into a revision.
This closes stale pairing in the revision API, not the raw-tree compatibility
APIs described above, and does not certify recovered CHAT as valid.
New editor integrations should use the revision API. The raw CST, strict-model,
and streaming incremental methods remain compatibility entry points whose
callers must edit the old tree correctly; they do not accept a proof of that
transition. Their continued availability prevents treating source-bound child
range failures as impossible across the entire public parser API.
The edit calculator counts shared complete Unicode scalar values, not shared
bytes requiring later UTF-8 repair. Equal characters have equal byte widths;
the resulting prefix is a boundary in both sources. The suffix is counted only
within the remainder, so it cannot overlap the prefix. This removes intermediate
split-code-point states without removing any CST recovery or range admission.
Utterance lowering also takes SourceBound<UtteranceNode> directly from the
document projection. Its entry point accepts no independently chosen
source string. Its generated dependent-tier repeat retains SourceField
association until attachment admits the choice once as SourceBound<Choice>.
Parse-health classification and exhaustive dependent-tier dispatch select from
the generated ChoiceBoundView: each variant retains the same admitted range,
so dispatch does not repeat range checks in every arm. This capability exists
only for single-node choices; composite choices still require individual child
admission. Main/dependent-tier leaf adapters still take
raw typed nodes and the source borrowed from that owner; migrating their own
child projections is separate work. Association alone does not prove their
recovery or decoding branches unreachable. Missing/Error/Absent slot states
remain distinct, with conservative parse-health taint preserved.
User-defined %x* lowering keeps the admitted concrete tier and projects its
optional body through associated slots. Present and typed Missing bodies both
retain their own range admission; text comes from the resulting bound node,
not a second node/source decoding attempt. Empty tiers remain in the model,
and Error/Absent body states retain their existing diagnostics. Prefix handling
and other dependent-tier families remain transitional leaf adapters.
@Participants lowering keeps this association through contents, list groups,
speaker codes, names and roles. Generated SourceSlotView keeps genuine
Missing, Error and Absent outcomes distinct, with impossible payloads remaining
uninhabited. Leaf reads admit their canonical byte ranges at the shared boundary;
there is no independent source argument that can be paired with a participant
node. This removes transitional decoding from that family, not recovery.
@Media keeps the same association through its body, filename, media type
and optional status group. Body extraction consumes the selected carrier through
admit_ranges() before lowering: typed leaves retain checked slices and the
payload reader can refuse only structural recovery, not source admission.
No independent source argument or temporary payload strings are needed.
This opt-in transition checks all selected ranges, even delimiters whose text
lowering does not use. It trades larger transient carriers and eager checks for
a single explicit producer-failure boundary; it is not a throughput claim.
The two reference media headers admit four and seven ranges respectively. On
the measured 64-bit build, their inline body carrier is 888 bytes versus 632
before admission (960 bytes including its owner). These are transient layout
figures, not peak-memory measurements or a reason to migrate every consumer.
Missing/Error/Absent slots, displaced nodes, range refusals and the validated
MediaFilename constructor remain. Header-level recovery still uses the shared
diagnostic site. The thirteen scalar/text headers and the single-value
@Number, @Recording Quality and @Transcription headers use the same
source-associated reader. Their entry points cannot take an independently
selected source, and model constructors receive borrowed checked payloads.
HeaderSite::bound derives diagnostic identity from that same capability.
The two-slot @Birth of, @Birthplace of and @L1 of headers do likewise:
their participant slot requires generated SpeakerNode identity, and the
admitted result distinguishes SpeakerCode from borrowed value text. A failed
speaker read still prevents value decoding. The old owned-string slot reader
has no callers and is removed. Comment headers retain association through body
admission too; the bullet-text adapter receives that same source-bound node,
with no separate source argument. HeaderSite has only the producer-bound constructor,
so diagnostic sites cannot be built from independently paired nodes and input.
Options, bullet-text internals and other structured header internals remain
transitional; this does not prove recovery states unreachable or certify
semantic validity of source-admitted values.
The generated typed traversal (generated_traversal)
The bridge between the tree-sitter CST and the typed model is a single
generated module, crates/talkbank-parser/src/generated_traversal.rs,
produced by the tree-sitter-grammar-utils generator from the grammar’s
own machine-readable description (grammar/src/grammar.json plus
node-types.json). It contains one extract_* function per grammar
rule, each returning a typed view of that rule’s children, so consumer
code dispatches on generated types rather than on node.kind() strings.
Every child position a grammar rule models is exposed as a NodeSlot
with five states, of which each position’s type admits only the ones that
position can produce:
NodeSlot state | Meaning |
|---|---|
Present | The expected node is there; a typed accessor is available |
Missing | Tree-sitter inserted a zero-width MISSING node during recovery |
Error | An ERROR subtree occupies the position |
Unexpected | A node of an unmodeled kind landed here |
Absent | A fixed position has no matching child; optional emptiness is None |
The generator names the position’s kind in the slot’s type. A ChildSlot
(a child taken by kind) is never Unexpected; a SeqSlot (an inline
sequence) is never Missing or Unexpected; a selected ChoiceSlot is never
Unexpected; a ClassifiedSlot (a supertype rule’s own node) is never
Absent. The impossible states have the uninhabited Never as their
payload, so an arm that reads a node out of one does not compile, and a
match by value may omit it. Every generated accessor hands out a
reference, and a match through a reference must still name every variant,
so a consumer matches slot.view(), which copies the recovery states out
by value and borrows only the present payload. The parser’s shared verbs
(expect_present, expect_structure, expect_delimiter, present) are
generic over the slot’s recovery payload types. Repeat elements and present
optional elements use SelectedKindSlot, SelectedChildSlot,
SelectedSeqSlot or SelectedChoiceSlot: their Absent payload is Never.
Nested fixed positions retain their own slot types; no missing/error recovery
is removed by selecting an outer element.
This design makes silent recovery-node loss structurally impossible at
modeled positions: Missing and Error are explicit variants every
call site must handle, not conditions a hand-written walk can forget to
check, and a diagnostic for a state the position cannot reach cannot be
written either. Missing maps to E342 (a MISSING placeholder for a required
element); Error reaches E316, which is the generic “content could not be
parsed” catch-all.
That asymmetry matters when reading a diagnostic. E342 names a specific fact
the parser knows. E316 names the absence of one, so an E316 on input a human
can read is a standing invitation to ask whether the parser, rather than the
file, is at fault. Generic rejection alone is not a specificity defect;
a narrower diagnosis requires structural evidence. Hand-walking the CST
with node.kind() comparisons, and classifying the text of ERROR nodes
to guess what was malformed, are both banned in production parser code
for exactly this reason.
What to write instead, when the question really is “which alternative is
this node?” A grammar rule whose alternatives are each a single named kind
lowers to a <Rule>Choice enum, and the generator emits that enum’s own
classifier, <Rule>Choice::from_node. It returns Option<Self>: None says
the node is not one of the alternatives, which is a fact about the input, not a
default to paper over. Matching the result is exhaustive, so adding an
alternative to the grammar breaks compilation at every site that decides on it,
which a chain of kind() string comparisons never does.
The classifier is emitted exactly when a kind can identify an alternative. A
choice with a sequence, repeat or optional alternative is told apart by
STRUCTURE, so it deliberately gets none: there, the shape is the question, and
a kind-keyed answer would be a guess. When no classifier exists, extract the
rule and match on the carrier rather than reaching for kind().
Bullet-capable free text follows that structural path: BulletTextNode retains
the two grammar carriers, and lowering consumes their generated first/repeated
choices and nested bullet/picture groups. The groups own trailing spaces;
recovery positions and displaced sinks remain explicit. Structured timestamp
parsing requires BulletNode and derives start/end fields from its generated
carrier, rather than accepting arbitrary nodes and looking up string field names.
Checked transitional text reads reject invalid ranges but do not by themselves
prove tree/source identity; the producer-owned source boundary remains distinct.
The internal bullet-text sum type exposes only source-bound carrier conversions,
not an independent raw-node classifier or raw projection. All nine bullet-text
dependent-tier adapters preserve the dispatcher’s binding, and their optional
body slots retain source identity through generated projections. Present and
MISSING bodies still undergo fallible range admission; ERROR and absent bodies
retain their separate diagnostic policies. Generated choice and nested-group
projections carry source identity through the inner segment sink too; it owns
no separately supplied source. Text, inline-picture and inline-bullet leaves
accept admitted source-bound nodes. Picture delimiter/filename admission and
the inline all-zero timestamp policy remain separate from source admission.
The shared timestamp decoder still has a transitional raw interface; this
migration does not prove its recovery guards redundant. Displaced children
still use the shared recovery collector. Reference tests obtain carriers through
the producer-bound descendant API rather than pairing raw nodes with text.
Leading-zero timestamp diagnostics require an admitted LeadingZeroTime
spelling and stream directly to the sink. Both numeric fields must first fit
the timestamp representation; overflow takes precedence over spelling reports.
The two-component case reports start before end, retaining the bullet’s source
span and numeric value. The E748 spec family and its corpus contract check
that multiplicity, ordering and precedence; the observation snapshot alone
records code sets and cannot establish diagnostic counts.
The legacy single-utterance helper probes through an admitted wrapper and the checked parse producer, rather than calling tree-sitter directly and treating failure as fragment shape. Its private classified-input state retains either the complete document envelope or the exact input requiring scaffolding. A valid no-transcription document may contain no utterance; the helper then reports MissingMainTier without reclassifying the document as a fragment.
Inline pictures additionally admit their decoded text into a private nonempty filename value before conversion to an owned string. Its constructor checks both delimiters and the nonempty payload. Tests retain real picture nodes and exercise incompatible source and delimiter boundaries; those are API admission witnesses, not evidence that the grammar produces malformed picture tokens.
Inline bullet times likewise pass a private admission value that excludes the all-zero pair. This is an inline-tier rule, not a time-ordering proof. Parse-backed tests produce separate trees for zero/positive timestamp combinations; source incompatibility is tested separately without treating it as malformed CHAT.
The ban above stood for a year with nothing to point at, which is why the
node_types kind-constant catalogue kept being reached for: a prohibition
loses to whatever the API actually makes easy. This is that missing
affordance, and it lives in the generator so every consuming grammar gets it.
Recovery handling is two-layered by design: the per-position NodeSlot
states cover every position the grammar models, and a whole-tree
recovery backstop (see the recovery discussion below) surfaces recovery
nodes that land where no grammar rule models a slot, such as top-level
junk. The layers are complementary, and both are load-bearing: removing
the backstop demonstrably regresses the CHECK-parity and
recovery-is-not-validity test suites.
Main-tier body sinks are owned by a sealed set of generated carriers. Their
reporter derives the displaced nodes from the carrier and the diagnostic context
from its owning rule’s generated NamedKind; callers cannot supply a different
sink, region or label. ERROR children use the shared body classifier, while other
displaced nodes use the ordinary reporter. This does not prove node/source
identity or remove recovery states. Raw slot-level errors still retain their
explicit body-versus-outside-body policy.
Annotated angle-bracket, phonology and sign groups belong to this sealed body-carrier set. Recovery can place malformed word material beside a group’s contents slot; that material must remain rejected, without requiring a particular diagnostic. The E202 spec pairs a valid suffix with a one-character deletion inside a retraced group after multibyte text, exercising the normal document path. Parallel valid/deleted-suffix pairs exercise phonology and sign groups; their displaced word faults are rejected with generic E316. The parser does not rescan recovery text for a dangling @ or an unclosed replacement opener. E202 applies at the model-validation boundary; E311 has no live producer.
Fragment-corpus tests also extract dependent tiers through model-owned spans and compare public full-line and content-only APIs against file parsing. Timing-tier payload states remain distinct during rebasing: their text and clock values are unchanged, but the source span in every state shifts with the owning fragment. Invalid E603/E756 fixtures cover unsupported and empty timing content without inventing model values.
Linkers likewise use generated first/repeated groups and exhaustive concrete alternatives, not a second kind-name catalogue. A source-ordered accumulator keeps each token’s kind and span together and reinserts displaced recovery in source order. Normal traversal appends directly; the recovery boundary uses the generated supertype classifier to retain any concrete linkers found in a sink.
Main-tier speaker admission retains SpeakerNode until checked text decoding.
Only that admission can construct the private nonempty speaker/span value, which
conversion consumes after reporting the rest of the tier’s diagnostics. Missing
placeholders retain their MissingSpeaker diagnostic; unreadable source ranges
reject with TreeParsingError. This compatibility boundary does not establish
tree/source identity.
Unsupported dependent-tier lowering retains the dispatcher’s
SourceBound<UnsupportedDependentTierNode> instead of separating the node from
its source. Both unsupported and user-defined prefix decoding consume generated
source-associated fields, so callers cannot supply independent text for those
slots. Missing, error, and absent prefix states retain their diagnostics; marker,
nonempty-label, and anonymous body-range checks remain. Parent source admission
does not prove that a generated leaf span fits UTF-8 boundaries.
Prefix decoding retains kind-proven placeholders for star, speaker, colon and
tab through the generated KindSlotValue projection. Speaker and colon
diagnostics receive their respective typed nodes, rather than a raw-node
projection. Present, MISSING, ERROR and absent policies remain distinct; the
colon width check occurs inside the typed-node arm without an intermediate
optional raw node.
Malformed dependent-tier labels use a private nonempty admission value after the leading percent sign. Only colon, tab, space, carriage return and newline terminate a label; unknown labels remain unknown rather than matching a known prefix. This recovery classification does not certify CHAT validity. Both document and utterance recovery share its alignment-domain policy.
The independent re2c text-tier path admits its source-located tokens before building recovered content. All-zero inline media bullets produce E360 at their lexer spans and are omitted, while surrounding tokens remain. Full-file text tiers and fragment entry points share this admission; fragment sinks rebase diagnostic locations through the admitted fragment source. This does not change main-tier bullet timing or establish complete backend parity.
Top-level recovery binds its node to a producer-owned SourceSlice, deriving
text and location from the same input. Binding refusal reports a parsing error.
Admitted recovery reaches structural delimiter evidence or generic rejection;
no headers or tier labels are reconstructed from its text.
Generic file-error analysis consumes that same binding instead of re-reading a
raw node against a separately supplied string. Line-slot ERROR handling also
binds before entering this analyzer; source-binding refusal is a parsing error,
not a diagnostic constructed from a mismatched source.
Both routes share the same binding/refusal function. Its regression test uses
real ERROR nodes from a retained spec fixture and a separately parsed identical
source: the producing owner admits them, while the other owner must refuse.
That is source-identity boundary evidence, not a production recovery-slot witness.
Delimiter-specific diagnoses require grammar evidence. A wrapper proving readable text does not establish an unclosed construct, so no text classifier produces E312 or E313.
No classifier infers an annotation from text prefixes or missing whitespace from the preceding character. Such malformed inputs are rejected by grammar recovery, with generic E316 where no structural evidence supports a narrower fault. Canonical valid controls remain clean; a diagnostic-specific state is not itself proof of parser structure.
Non-ASCII speaker IDs remain invalid. Parsed speaker fields use the model’s E307 assessment; a main-tier ERROR without an admitted speaker field receives E316 rather than reconstructing a speaker ID by splitting text. Supported ASCII forms remain unchanged.
Recovery never reconstructs headers or dependent tiers by scanning an ERROR prefix. Unidentified top-level recovery conservatively taints the preceding utterance’s dependent alignments; unidentified utterance recovery taints the main tier and dependent alignments. A typed dependent-tier choice can still narrow taint to its own domain. Recovery text is retained for diagnostics, not promoted into a fabricated header or reparsed to select an alignment domain.
Postcodes retain opaque labels in Postcode; their text is not reparsed as
quotation syntax during validation. Actual quotation delimiters, typed linkers
and terminators have their own checks. A postcode whose label resembles a
quotation marker cannot create a quotation-balance obligation.
Bracket-to-word spacing validation projects each content item’s following-word and closing-code evidence together with its enclosed sequence. One exhaustive mapping serves both top-level and bracketed content. A private nonzero code-end value excludes the unlocated sentinel before adjacency can be checked. Each recursive sequence starts with no predecessor: flattening an outer wrapper and its children would create false neighbors. E757 controls and single-space deletions cover annotated events, actions, quotations, retraces and phonology/sign groups; the diagnostic contract checks exact byte offsets and serialization back to the paired control. This is source-location evidence, not a replacement for the parser’s tree/source ownership boundary.
Replacement-word producers also retain the whole typed CST wrapper’s span, including trailing scoped annotations. The shared spacing projection admits that wrapper’s code end while retaining the spoken word’s own start. Canonical replacement pairs exercise missing separators at either closing bracket and nested use.
Wrapped-header selection starts at ParsedSource::root() and carries bound
descendants through SourceSlice::children. Admission takes only the owning
parsed wrapper and header ordinal, so it cannot pair a selected node with another
wrapper’s input mapping. Child traversal resets its cursor to the bound parent
and avoids repeated root-wide lineage searches. Header-counting policy and
complete-input checks remain separate from that source-ownership proof.
The lookup result is a privately constructed SelectedHeader: generated
header/pre-begin choices and anchor wrappers admit its kind, with the deliberate
Thumbnail exclusion retained. The grammar’s invisible supertypes do not add
concrete wrapper nodes, so lookup traverses the real line wrapper but never
tries to unwrap an abstract header or pre_begin_header. A selected header
cannot be a generic ERROR node. MISSING placeholders and other lowering
failures remain recovery states, not proof that the fragment is valid.
The old handwritten header-kind predicates and their membership-only test are
removed with their last consumer. Generated choices own subtype membership;
reference/spec fragment tests continue to own behavior and recovery policy.
User-defined dependent-tier taint routing reads the generated Present prefix
slot, not a positional raw child. %xmod identifies the model-alignment domain;
other labels and recovered or unreadable prefixes do not identify one. When
attachment reports errors without a specific domain, all alignment dependents
remain conservatively tainted.
The module is regenerated whenever the grammar changes; the command and its
preconditions are in Grammar Workflow.
It is never edited by hand: generator defects are fixed in
tree-sitter-grammar-utils and regenerated.
The generated cursor owns its remaining-child iterator. Selection retains a
source-bound match plan, and extraction consumes that plan without rematching.
PositionAdmission distinguishes a selected position from an unmatched cursor;
only successful carrier construction commits cursor movement and recovery.
Absolute indices refer to the stable child slice. Exhaustion stays at EOF and
the final sweep visits actual remaining children. A ReconstructionFault
reports a producer invariant failure, not malformed CHAT. A selected choice
therefore has an uninhabited Unexpected payload: it cannot independently
rematch and disagree with selection. Missing, Error, Absent and displaced-node
recovery remain intact, as does Unexpected for supertype classification. This
is a producer-construction invariant, not an inference from corpus coverage.
The fold also retains selected presence for repeat elements and the Some
payload of optional positions. Their extraction consumes a retained element,
so it cannot yield Absent. Empty repeats and optional None remain valid
outcomes. This does not narrow independent fixed positions or prove that a
selected element is free of MISSING/ERROR recovery.
Document, utterance, participant/language header and main-tier/body/ending
reconstruction use compiled-language admission. Complete document and ERROR-root
extraction share this contract; neither route may fabricate a complete document.
The grammar crate compiles TSGU’s generated C metadata bridge against its own
parser header. A cached, fallible CanonicalLanguage capability verifies the
generated narrowing requirements once; extraction checks actual language
identity and retains the proof in the existing selection memo. There is no
second parser, rematching pass or consumer assertion based on a kind name.
Admitted nonterminal slots have an uninhabited Missing payload, including the
participant/language contents, participant entries, tier bodies, body contents,
linkers, utterance endings, document anchors, selected lines/dependent tiers and
final-code/postcode slots. Lexical slots such
as speaker and language codes retain Missing, and all applicable Error, Absent,
displaced recovery and source-read failures remain. Raw extraction stays broad.
A failed metadata admission is an internal tool failure, not invalid CHAT.
Main-tier body location still searches the source-associated recovery sink when
the body is not in its expected slot. A nonmissing proof does not prove that a
required position exists or that displaced children are impossible. Nested
terminator recovery and lexical prefix diagnostics remain unchanged.
Mixed line choices retain their broad Missing state when any alternative lacks
the nonterminal proof. Document recovery ownership failures use E001 and cannot
certify source invalidity. The low-level pre-begin/dependent-tier adapters accept
the generated admitted choice types; their concrete header/tier policies remain
unchanged. The generated broad/raw APIs remain available for other producers.
Wrapper and supertype child positions use KindSlot: extraction matches the
concrete kind before constructing KindMissing<T>. The kind-preserving
known_or_placeholder projection therefore has no unclassified-placeholder
case. Raw partial and composite-choice slots retain their fallible projection
and recovery states. Diagnostic view() still exposes the raw missing node;
generic conformance observes its payload through RecoveryNode. Consumer
foreign-kind fallback branches are removed only for the kind-proven slots.
Generic, word and dependent-tier recovery cross a shared ReadableRecovery admission
boundary before examining text. The admitted value retains its node, source and
checked slice together; the classifier consumes that value instead of accepting
independent text. Its fields remain private and consumers cannot construct it
without checked slicing. This proves range compatibility, not tree/source identity.
Incompatible ranges retain each boundary’s diagnostic family instead of
panicking in tree-sitter’s byte indexing.
Dependent-tier classification derives both coordinates and text from that
carrier; nonempty admitted text therefore needs no second positive-span check.
Delimiter findings remain priority-sensitive: a leading space can bypass the
first-character delimiter finding, so a later trim-based bracket fallback is
not generally redundant.
Empty-POS morphology findings retain both the recognized token and its relative byte range. Both dependent-tier analysis and whole-tree recovery use that same occurrence; searching for the token again could select an identical split tail that the classifier deliberately excluded. UTF-8 and Unicode whitespace retain byte-correct spans, and there is no invented whole-node fallback range.
Language-list admission consumes the generated language-code slot through one policy for first and repeated entries. A position enum preserves their distinct diagnostic wording; only Present enters the fallible model constructor. Missing, ERROR and absence retain their recovery behavior. Sequence-level Missing is eliminated only through its generated uninhabited payload, not corpus absence.
Word-language resolution keeps explicit, mixed and ambiguous candidate sets distinct. Each explicit resolution is consumed into an outcome carrying its registry diagnostics, using one shared transition over every candidate code. An ambiguous candidate is not exempt from ISO validation, and an unresolved shortcut never invents a fallback language.
Participant-list slots use the same admission pattern, with a position type that carries the enclosing header only for the first slot. Absence there still reports an empty header; repeated-slot absence does not. Both positions share entry conversion and Missing reporting while preserving their ERROR wording.
Scoped overlap indices are admitted as 1–9 at construction and JSON decoding;
the schema carries the same bounds. Unindexed markers use None, never a
numeric sentinel. Atomic token decoding distinguishes an absent index from a
malformed one, so malformed content cannot become an accepted unindexed marker.
The experimental backend carries the admitted index with its original lexer
slice into model conversion. CA overlap-point indices retain their separate
recovery representation and 2–9 validation policy.
CA element and delimiter nodes retain their generated kinds into a private decoder trait. Each kind fixes its registry, diagnostic label and model output; callers cannot mix those policies. Source admission uses checked bounds and UTF-8 decoding before character classification. Missing placeholders remain rejected for whole-tree recovery reporting, and unknown symbols still diagnose TreeParsingError. Source compatibility is not proof of original-tree identity.
Utterance construction and main-tier conversion consume the generated
SourceBound<MainTierNode>. This replaces the handwritten ReadableMainTier
range wrapper: the immutable parse owner and its field projections supply
both node identity and checked text. Failed child-field admission reports
an internal producer failure without constructing an utterance. Fragment
conversion retains the same capability; its separate original-input parameter
is diagnostic context, not a substitute source for CST text. Inner word/content
adapters have their own migration boundaries; this parent capability alone does
not prove that every leaf API is source-bound.
Speaker-prefix admission also consumes a SourceBound<SpeakerNode>, projected
from that main tier’s associated children. It cannot accept a separately chosen
source string; even identical text from a different parse owner is not the same
capability. Present text is range-checked at the generated leaf boundary before
nonempty speaker admission. Missing, error and absent slots retain their
distinct recovery handling, and a failed source binding is not CHAT invalidity.
Body selection projects the main tier’s associated body slot and recovery sink.
Its result distinguishes an absent body from a located body whose source range
was refused. The latter reports an internal producer failure and rejects construction after
the remaining main-tier diagnostics are emitted; it is not reported as a missing
terminator. Body decoding consumes a bound TierBodyNode and derives its source
and carrier range from that node. Displaced-body recovery remains available;
its absence from current fixtures is not grounds for deleting it.
Generated raw source fields provide read_typed::<TierBodyNode>() for this sink
selection: None means a different kind, while a matching node retains either
its range refusal or its admitted bound wrapper. This replaces the consumer’s
separate kind check, raw read, and repeated classification. The operation is
owned by TSGU and regenerated into Chatter, not hand-edited in the generated file.
The recursive contents/group cycle retains generated source association: body contents, first/repeated content choices, content-item choices, annotated angle groups, quotations, phonology groups, and sign groups all pass bound nodes or associated children into the same walker. Bound choice views preserve the admitted wrapper without another kind classification. Kind-proven Missing group contents retain their placeholder identity through extraction; other missing content uses source-owned classification and keeps its existing policy. The utterance-body versus inside-brackets distinction still owns E759 behavior, and angle-edge whitespace still owns E750. Range refusals and recovery sinks remain visible.
Standalone-word conversion also requires a producer-bound word. Base-content
choice dispatch retains that association through annotated words and replacement
sequences; fragment lowering never detaches its admitted word. The independent
%wor route carries its bound tier through body, repeated item choices and word
items to the same converter. No leaf reconstructs ownership from a detached node
and arbitrary source. Word-level source text uses the admitted slice; missing and
empty words retain their separate refusals. Document attachment still drops
malformed %wor tiers, while the public bound-node adapter retains recovery
handling because source ownership does not prove syntax validity.
Word bodies retain ownership through both generated sequence shapes and every piece choice. Segments, stress, lengthening and shortening content consume admitted text; standalone and word-internal overlap markers share one source-bound decoder. Missing/error/absent positions and displaced children retain their existing recovery handling. Source admission does not prove nonempty lexical content or valid marker semantics, so those checks remain.
Word suffix and CA adapters, nonword and annotation leaves, and %wor
language/bullet/separator payload decoders still have transitional interfaces.
Their child-range checks are not removed by the outer word’s admission.
Utterance-level recovery consumes ReadableRecovery too, forwarding that
admitted value directly into dependent-tier classification when appropriate.
Unreadable input cannot select a tier: it reports internal failure and taints
Main plus all alignment dependents. Valid-source recovery retains the existing
label-based taint policy; boundary tests do not imply production slot witnesses.
The generated repeat producer retains per-element decisions and boundaries in flat arena links. Consumption checks their continuity rather than rerunning selection or applying a detached count. Producer faults propagate through E001 and block validation/admission; they are never empty successful carriers. No repeat recovery state is narrowed, and element-owned extras are not filtered.
Word and main-tier fragment admission stores SourceBound<GeneratedNode> from
ParsedSource::bind_typed. The binding’s sealed wrapper trait prevents external
adapters from claiming a mutable node identity. Lowering obtains both the node
and source from that one value; there is no separate node field to pair with the
wrong bound slice. Existing completeness, wrapper-offset and recovery policy
checks remain independent and unchanged.
Phonology fallback accepts PhoGroupNode rather than an arbitrary raw node and
uses the shared checked-text admission boundary. Nonempty readable group text
still becomes one fallback word, and empty or unreadable text adds no item.
Unreadable ranges diagnose TreeParsingError. These boundary tests do not prove
that a valid corpus fixture reaches fallback through production extraction;
the recovery branches remain until the producer justifies narrowing them.
The staleness guard proves less than its name suggests.
generated_traversal_is_current recomputes the digests of grammar.json and
node-types.json, so it catches a forgotten regeneration after a grammar
change and nothing else. It cannot see which generator produced the file, so a
module emitted by an older backend passes indefinitely. The generator’s name
and version are stamped in the file’s own header comment; read that when the
question is which backend built it.
The compiled conformance walk checks real CSTs from the reference corpus.
Its slot-state census admits the complete reference and error populations
before reporting; discovery, read or parse failures cannot silently shrink
the measurement. Positions carry the generated carrier type and field name,
so a nested group’s child_0 cannot merge with its outer carrier’s child_0.
A recovery observation is a reachability witness. A state absent from this
finite population is not proof that the producer cannot emit it, and is not
grounds for deleting its handling.
Error Recovery
Tree-sitter’s GLR algorithm provides automatic error recovery. When the parser encounters unexpected input, it:
- Inserts ERROR nodes in the CST
- Continues parsing the rest of the file
- Reports parse errors via the
ErrorSinktrait
This means the parser always produces a result, even for malformed files, it extracts as much structure as possible.
ParseOutcome
Individual parse functions return ParseOutcome<T>:
ParseOutcome::parsed(value): successfully parsedParseOutcome::rejected(): could not parse this node (error already reported)
This allows the parser to skip individual malformed elements while continuing to parse the rest of the file.
Parser Equivalence
The reference corpus is the primary correctness signal:
cargo test -p talkbank-parser-tests --tests reference_corpus_parses
Each .cha file is its own test, so failures are reported per file. The file
count is deliberately not stated here, because it grows. Ask the tree
(rg --files -g '*.cha' corpus/reference | wc -l).
TreeSitterParser API
TreeSitterParser is the concrete canonical parser handle. Reuse an instance
across calls. The shared talkbank_model::ChatParser trait also supports
generic callers and backend parity tests; both tree-sitter and re2c implement
it. Its sink methods are generic, so the trait is not a dyn trait object.
use talkbank_parser::TreeSitterParser;
let parser = TreeSitterParser::new()?;
// Full-file parsing (methods on TreeSitterParser).
// ParseProduct::Built retains the file and diagnostics together, even when
// recovery was necessary. Unbuildable has diagnostics without a model.
let product = parser.parse_chat_file(&source);
// parse_chat_file_streaming pushes diagnostics into an ErrorSink as it
// goes, useful for very large files or LSP-style incremental flows.
let chat_file = parser.parse_chat_file_streaming(&source, &errors);
// Fragment parsing (methods on TreeSitterParser), used when synthesizing
// CHAT from non-CHAT sources (ASR output, UD annotations).
let word = parser.parse_word_fragment(word_text, document_offset, &errors);
let main_tier = parser.parse_main_tier_fragment(tier_text, document_offset, &errors);
Diagnostic coordinates
Fragment parsing adds the caller’s offset to model spans and
diagnostic document locations. FragmentSource admits the complete input
range before parsing and owns this translation for both backends. Admission
checks origin + input.len() without overflowing; ranges beyond u32::MAX
are rejected with E310 and an unknown location, since no representable
location exists. Origins above i32::MAX remain supported: the existing
signed edit-shift interface receives bounded positive steps, never a wrapped
negative origin. A diagnostic’s ErrorContext owns its own
source text, so its highlight remains relative to that text. Wrapper removal
is a separate operation owned by WrappedFragment; it projects the synthetic
source before applying any document origin.
The same ownership rule applies to SpanShift on ParseError: document
locations and secondary labels shift, but the retained ErrorContext text and
its relative highlight do not. Rebasing a diagnostic cannot rewrite an
independent context snapshot. Corpus fragment contracts check this public
operation against parser-produced rebasing, including the inverse shift.
Word and main-tier fragments use the multi-root grammar directly, so there is
no synthetic prefix to subtract. MainTierFragment admits a typed main-tier
node from a ParsedFragment, deriving the original input from the same owner
that assembled and parsed the newline-extended source. Capacity admission
precedes allocation; callers cannot supply an independent original input. The
node is admitted only when it covers the complete parse source and the root has no extra
or unexpected content. Lowering consumes that proof together with the original
input, clipping the aggregate tier/content spans to exclude an appended line
terminator. LF and CRLF supplied by the caller remain part of those spans.
Trailing garbage or another tier cannot be silently ignored.
Non-colon separator decoding consumes the generated typed_or_placeholder
view: present and kind-admitted MISSING alternatives share one semantic mapping,
while unclassified placeholders retain their distinct recovery diagnostic.
The consumer does not reclassify raw MISSING nodes independently.
Marked-token and typed-text decoding share checked node-text admission. Its typed result distinguishes a range outside the supplied source from a range that cuts a UTF-8 code point. Marker stripping receives only the admitted text and still rejects the wrong marker. This boundary proves readable bytes, not identity with the source that produced the tree; producer-bound slices remain the stronger API.
Language-list recovery carries its offending node in a LanguageListFault
variant: an unexpected first/subsequent code or an unparsable repeated group.
Comma and post-comma whitespace faults have their own variants as well.
The consuming reporter owns the corresponding wording and shared diagnostic
location, rather than accepting arbitrary message text at separate call sites.
Participant recovery follows the same pattern with ParticipantFault: list
comma/whitespace/entry expectations retain E506 and list context, while malformed
entry structure retains E316 and entry context. Sharing the reporter does not
merge those policies or remove unobserved recovery states.
Sign-tier token decoding retains either a generated word or whole recovery-group
node in SinTokenSource, which selects its diagnostic context and shares checked
text admission. Whole-group fallback preserves the complete source text as one
token; an incompatible source reports a read failure. Its direct boundary test
does not claim that the fixture itself enters recovery during normal parsing.
Morphology feature-value decoding similarly accepts only its generated node
type, including kind-proven MISSING placeholders, and uses shared checked text
admission. Empty-value and unreadable-source refusals retain their existing
diagnostics; source compatibility is not mistaken for tree/source identity.
ReadableRecovery::fragment_diagnostic owns the pairing of a source-located
span with fragment-local display context. Word-error classifications supply
only their code and message; they cannot accidentally use another recovery
node’s location or another fragment’s context through this constructor.
The generic recovery classifier uses the same constructor for fragment-local
diagnostics. Classifications that need a subspan or full-source context remain
separate.
Main-tier conversion returns either the model or a private-constructor
ReportedMainTierError issued when the producer reports its rejection
diagnostic. Speaker admission retains that evidence through conversion, so the
fragment consumer needs no synthetic “failed to build” fallback. Recovery
diagnostics still precede rejection, and successful models may still carry
diagnostics; the result does not certify validity.
Header, utterance, participant-entry and dependent-tier adapters, including
the standalone parse_header and parse_tiers entry points, use an
owned WrappedFragment: its constructor records the actual input boundary as
it assembles the source, and both the model projection and diagnostic sink use
that boundary. Its diagnostic sink removes that prefix from both primary and
secondary spans. Context is projected only when its text exactly matches the
owned synthetic source; an independent context retains its own coordinates
regardless of length. cargo test -p talkbank-parser --lib api::fragment::tests
checks long inputs, related labels and independent context text. These adapters
do not use the legacy sink’s length heuristic. Header lowering consumes a
HeaderFragment that owns the located node together with its wrapped source.
Admission requires the node to account for all caller text: only surrounding
whitespace may lie outside it or extend into the synthetic line terminator.
A start-only check would accept the first of two headers and discard the
second. The public regression covers that refusal boundary, while controls
retain folded header content and caller-supplied LF/CRLF. Raw ordinal lookup
and document-root navigation remain separate traversal improvements.
Header lookup failures carry tree facts in HeaderNotFound, rather than
constructing a parse error with invented empty context. The fragment caller
attaches the real input and its full span; public fragment rebasing then adds
the document origin to the location while leaving that context local. The
context_public_api::unlocated_header_reports_the_callers_source_and_origin
regression exercises this failure through the public API at origins zero and
200.
Complete documents passed to the utterance adapter are recognized
through generated typed CST traversal and receive no extra document wrapper.
Main-tier recovery collection is a method of the admitted MainTierFragment.
The generated SourceBound::descendants() iterator owns its private cursor,
starts at the admitted main tier and cannot leave that subtree. Each descendant
retains its canonical source slice, including ERROR/MISSING nodes; a runtime
range refusal remains a diagnostic rather than a skipped node. Bound error
analysis consumes that slice without an independent source argument or another
UTF-8 admission. No caller supplies a separate recovery node, source or offset. Missing and error
nodes remain reported in source order. Diagnostic ranges are clipped to caller
input, excluding an appended newline, and error text uses checked UTF-8 slices.
cargo test -p talkbank-parser --test integration context_public_api reproduces
the fragment regression checks: UTF-8 word spans, caller offsets, rejected
fragments, and utterances with or without a trailing newline or document headers.
These guard against prefix subtraction that collapses valid word spans to
zero, subtraction of caller origins from errors, and unremoved synthetic
prefixes on rejected utterances and participant entries.
The E326 boundary test exercises both parsers with LF and CRLF, UTF-8 content,
and offsets zero and 200. Unsupported-line recovery must identify each skipped
line, retain following utterances, and preserve the diagnostic’s local source
highlight. The fragment_range_tests public-API controls exercise both
backends above 2 GiB, at the final representable byte, and with overflowing
ranges. They also verify that diagnostic context stays snippet-relative.
Synthetic terminators for morphology, phonology and grammatical-relation
fragments are parsed at local origin zero; only the extracted caller result
is moved into document coordinates. Wrapper allocation separately admits its
complete synthetic source size, and trimming a caller newline cannot bypass
admission of the original input range. The legacy
SpanShift edit API and Span::from_usize truncation remain for unrelated
callers; the admitted parser paths use no raw origin casts. The legacy
OffsetAdjustingErrorSink is exported by the model crate for compatibility;
the tree-sitter parser has no callers of it.
Missing encoding declaration
The document grammar admits an absent @UTF8 anchor as an explicit optional
slot. Lowering consumes that generated slot and retains the present headers
and utterances without inventing an encoding declaration. Shared validation
still rejects the file with E503. The canonical parser retains the present headers rather than
discarding the document and reporting them as missing. The authored E503 example and its declaration-present
control exercise this recovery; CLAN CHECK reports the corresponding CHECK (69).
The re2c file parser carries each header’s lexer extent and separator together
in HeaderProvenance. Lowering uses that extent instead of an unknown span,
so shared missing-header diagnostics derive a real EOF from the final header.
The cross-backend missing_encoding_keeps_the_document_and_locates_the_single_refusal
test checks retained utterances, exact diagnostics, and EOF locations at zero,
nonzero and maximum representable document origins.
Recovery-wrapper suppression requires a private DocumentRecoveryWrapper
proof at the document position: every direct child must be a generated document
construct or separately reported recovery. A recognizable header beside malformed
raw tokens, or a stray header beside a complete document, cannot suppress E316.
DocumentRoot retains the parse owner and the selected document state. It
derives the syntax root from that owner and a complete document’s raw node from
its generated typed wrapper; neither identity is stored a second time.
Lowering finds a complete document even after a recovery sibling, while the
diagnostic backstop covers the entire source. This prevents trailing text after
@End from validating clean and avoids missing-header cascades when a leading
error precedes an otherwise complete document. Private fields prevent callers
from combining a document with an unrelated diagnostic scope.
Lowering consumes source-ordered DocumentPart values, rather than extracting
document children and discarding outer recovery. ERROR siblings of a concrete
document pass through the same producer-bound diagnostic route as errors inside
it. A duplicate @End fixture therefore retains E316 when omission of the final newline
moves recovery outside the document node. Reconstructed ERROR wrappers do not
establish that outer-document boundary: their siblings remain covered by the
whole-input backstop, so a lone End after a malformed Begin is not mislabeled
as a duplicate. The E501 spec supplies both final-newline regression variants.
Its input is ParsedSource, created by
TreeSitterParser::parse_source_incremental. The generated owner retains the
exact input supplied to tree-sitter and has no independent tree/string
constructor. Document lowering consumes the classification and derives source
and capacity from it; there is no separate source parameter. into_tree
consumes the association for incremental editing, and reparsing produces a new
association. Incremental editing does not clone the whole-document tree.
Recovery text uses the generated SourceSlice boundary. Raw node binding
checks both tree membership and canonical coordinates: a copied tree-sitter
node can be edited without editing its tree. Foreign and independently edited
nodes cannot yield a source slice. Required-slot Absent remains a legitimate
reconstruction result for incomplete ERROR nodes; lack of a corpus witness
does not make it impossible. Standalone word and main-tier admission retain
checked source slices as well, with word selection driven by the generated
source-file union. Wrapped headers follow WrappedFragment -> ParsedFragment -> HeaderFragment: the middle phase parses the wrapper’s own source, and header
admission binds the selected node through that owner before checking complete
input coverage. Deeper token APIs still retain independent source parameters
and remain migration work.
The transitional raw-node text helper returns ParseOutcome<&str>, rejecting
out-of-bounds or non-boundary ranges without manufacturing empty text. Its
consumers only construct successful model values after a successful read.
This is a checked range boundary, not a substitute for source ownership.
Morphology feature and marked-token readers use that same boundary: unreadable
source associations are internal failures, not malformed features or encoding
errors attributed to the author. Empty-feature and missing-marker recovery
remain separate structural diagnostics.
Other-speaker event lowering uses generated required slots; the generator,
not a parallel list of child positions and kind strings, owns its CST shape.
Header fragments and full documents share the exhaustive pre-@Begin header
decoder. Its generated choice includes @PID, @Window, @Color words, and
@Font; routing only @PID through that decoder would let the other
three fall through to successful Unknown values. Finite reference-corpus
tests compare header, main-tier and utterance fragment models with full-file
parsing, carrying the file’s CA semantic context explicitly.
The shared decoder consumes a source-bound choice. Document lowering
projects it from the associated repeat; fragment lowering classifies its
already-bound source slice. Header spans derive from that same choice, not a
separate caller argument. PID lowering projects its generated free-text field
and admits its range before reading the value. A failed binding propagates as
an internal failure, distinct from missing/empty PID recovery. Window, font and
color-word decoding use the same generated source-associated leaf admission:
their shared reader takes no independent source string and cannot turn a
binding failure into a malformed-header fallback. Empty, missing and error
slots retain their existing recovery policy.
The @Languages decoder likewise retains source ownership through the contents
and repeated list fields. Present language codes become model values only after
source-bound leaf admission; missing codes and malformed separators retain
their existing recovery diagnostics. A producer fault is an internal failure,
not an empty language code or a successful partial validation.
Participant-header contents likewise distinguish a missing structural slot
from a failed source read. The latter propagates as an internal failure and
rejects header construction, without adding an empty-participants diagnostic.
The source-associated content reader for simple headers and participant metadata
carries distinct structural-recovery and producer-failure variants.
Only structural recovery can construct an Unknown header; a source-binding
failure is reported once at dispatch and rejects construction. Media-body
range admission propagates the producer failure before its infallible text reads,
instead of returning an unreadable body as a successful unknown header.
Missing/error/absent slot policies remain
unchanged.
Participant entries own their WriteChat implementation, which the enclosing
header writer also uses. Reference-corpus entries exercise the standalone
participant fragment API through this canonical wire boundary, without a
second formatting implementation or fabricated participant models.
Error-spec documents also supply retained main-tier, header and dependent-tier
source spans to the standalone fragment APIs. Recovery tests compare local
and rebased admission, models and diagnostics, requiring evidence for rejection
and caller-owned UTF-8 ranges for diagnostic spans and labels, while leaving
diagnostic context relative to the original snippet. Full-document spec tests
independently enforce the authored diagnostic claims.
Main-tier inputs are also exercised after removing their terminal newline,
so the coordinate contract covers the adapter’s synthetic newline boundary.
The legacy utterance adapter’s full-document input mode is checked over the
same spec corpus, with and without terminal newlines, including documents
that cannot supply an admissible utterance. These are coordinate and recovery
contracts, not new claims about which whole-file validation rules should fire.
Nested choices receive a generated FromNodeKind implementation when their leaf
kind sets are disjoint. The generator retains each leaf’s complete constructor
path and refuses ambiguous or composite alternatives. Content recovery uses
this classifier; it maintains no parallel list of alternatives.
Present content items retain their typed wrapper through dispatch, eliminating
the raw-node conversion and impossible second kind refusal.
Syntax completeness does not establish semantic validity; shared validation still owns required headers and other CHAT rules. The LSP owns source-bound analysis snapshots rather than a second parser-level cache-admission API.
At EOF, lowering binds a generated MainTierNode stranded outside its line
wrapper to the original parse owner before reusing the normal utterance builder
and parse-health transition.
For the flattened simple terminal sequence without a final newline,
TerminalMainTier pairs the generated grammar tokens with the original source
range. Lowering binds the retained speaker, contents and terminator nodes to
their original parse owner, then uses the ordinary contents and terminator
decoders directly. It does not reparse a source fragment. The diagnostic backstop
uses the same structural admission. The E502 example checks retained speech
and diagnostics in both newline forms, including maximum representable source
origins; leading and trailing recovery-region regressions remain separate.
AST Structure
The resulting ChatFile AST has a recursive content structure:
flowchart TD
cf["ChatFile"]
hdr["Headers\n@Languages, @Participants,\n@ID, @Options"]
utts["Utterances[]"]
mt["MainTier\nspeaker + content"]
dt["DependentTiers[]\n%mor, %gra, %pho, %sin, %wor"]
uc["UtteranceContent\n24 variants"]
leaf["Leaves\nWord | ReplacedWord | Separator"]
group["Groups\nGroup | AnnotatedGroup |\nRetrace | PhoGroup | SinGroup | Quotation"]
cf --> hdr & utts
utts --> mt & dt
mt --> uc
uc --> leaf & group
group -->|recurse| uc
Parser String Handling
The tree-sitter parser constructs owned model types (e.g., MorWord, GrammaticalRelation) directly from CST text. String-heavy types like PosCategory and MorStem use Arc<str> interning to avoid redundant allocations for repeated values. Short strings in model newtypes use SmolStr for inline storage up to 23 bytes.
Editor source revisions
The LSP stores one DocumentAnalysis owning exact source bytes, a tree-sitter
CST, the lowered model, and diagnostics. Its constructor is the only route to
those artifacts. Reusing the CST first applies an InputEdit computed from its
own prior source, so debounced intermediate edits cannot substitute the wrong
baseline. Both the edit boundaries and tree-sitter columns use UTF-8 bytes;
LSP wire positions remain UTF-16.
Each changed analysis lowers the model and calls ChatFile::validate_with_alignment
in full. Previous header errors or absolute AST spans are not copied into a new
revision. This removes the independent cache maps and custom validation sequence
that missed deleted headers and file-level checks. Tree-sitter incrementality
and whole-analysis reuse for identical source remain. More selective semantic
reuse needs an explicit dependency and span-identity design plus measurements.
Feature requests during debounce admit cached models/trees only for identical source, otherwise parsing the requested text transiently. Pull diagnostics use the same analysis constructor as pushed diagnostics. A replaced or closed revision cannot commit its analysis, and push results carry the editor version. No cache guard crosses asynchronous publication.
The existing stdio integration binary checks fresh-open/edit parity, skipped revisions and requests during debounce. Its process owner handles shutdown and cleanup; the message inbox preserves interleaved notifications while awaiting responses. This catches production orchestration errors that isolated tree splicing helpers could not.
For a local computation measurement, run the ignored measure_analysis_latency
library test with --ignored --nocapture. Optionally set
TALKBANK_LSP_BENCH_SOURCE to an existing transcript; it is read without changes.
Record build profile and distinguish computation from the 250 ms debounce.
The benchmark is intentionally excluded from CI and sets no timing threshold.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
CHAT Data Model
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
The talkbank-model crate defines the typed AST for CHAT files. Every
other crate, parser, transform, CLAN, CLI, LSP, and the entire batchalign
runtime, depends on it. This page describes the model itself, the
three-level content hierarchy, the content-walker primitives, and the
extract → infer → inject pattern that all NLP tasks follow.
Date admission
Construct ChatDate through from_text or new. Valid(CheckedChatDate)
contains private components admitted together with the original spelling;
read them with day(), month(), year() and as_str(); match Valid(date)
and use these accessors. There is no independent component constructor or mutable field access.
This checks ASCII DD-MMM-YYYY syntax, uppercase month abbreviations and days
01–31, not calendar validity: impossible month/day combinations and year
0000 retain their existing admission behavior. Unsupported text is preserved
for diagnostics. JSON remains a string and deserialization uses the same
admission boundary; no unchecked payload deserialization is exposed.
ChatFile
The root type is ChatFile, representing a complete CHAT transcript:
pub struct ChatFile {
pub lines: ChatFileLines,
pub participants: IndexMap<SpeakerCode, Participant>,
pub languages: LanguageCodes,
pub options: ChatOptionFlags,
pub media: Option<Box<MediaHeader>>,
pub line_map: Option<LineMap>,
}
validate_into consumes the mutable model and returns either an immutable
ValidChatFile with policy/name/diagnostics or a ValidationFailure retaining
the rejected model. into_unchecked consumes a proof before editing.
Constructing rather than parsing
Utterance::new starts with unknown provenance. Appending a dependent tier
withdraws previous provenance and derived alignment results. Read provenance
through parse_health(); callers cannot assign the field. Parser adapters
finish a complete utterance through their accumulated
ParseHealth::finish_utterance, after recording recovery across all its tiers.
That adapter is a trusted producer boundary, not a general-purpose validator.
ASR and resegmentation callers assemble typed structure, then consume the file
through validate_construction_with_policy(policy, errors, name). This uses
the existing model rules, including alignment when selected, without serializing
and reparsing CHAT. Success returns ValidChatFile; newly admitted utterances
carry ParseHealthState::Constructed, not parser-backed Clean. Recorded
parser recovery still rejects. Failure withdraws new construction admission.
This proves the supplied structure under the selected policy, not completeness
relative to any original source or fidelity to audio.
Construction admission never authorizes source-byte splicing, even when copied
components retain spans. JSON omits runtime provenance and restores Unknown
on import; the existing JSON representation is unchanged. Consuming a proof
with into_unchecked withdraws its construction admission before mutation.
Parser-backed mutable models still expose content fields: callers editing those
fields must withdraw stale provenance with forget_parse_provenance().
The private health field is not a claim that all mutable-model edits are sealed.
Each Line is either a Header or an Utterance. The full ownership
tree:
flowchart TD
chatfile["ChatFile\n(talkbank-model/src/model/file/chat_file/core.rs)"]
valid["ValidChatFile (immutable validation evidence)"] --> chatfile
chatfile --> lines["lines: ChatFileLines\n(ordered Line newtype)"]
chatfile --> participants["participants:\nIndexMap<SpeakerCode, Participant>"]
chatfile --> languages["languages: LanguageCodes"]
chatfile --> options["options: ChatOptionFlags"]
chatfile --> media["media: Option<MediaHeader>"]
chatfile --> line_map["line_map: Option<LineMap>\n(not serialized)"]
lines --> header_line["Line::Header (Header)"]
lines --> utt_line["Line::Utterance (Utterance)"]
utt_line --> preceding["preceding_headers:\nSmallVec<Header>"]
utt_line --> main["main: MainTier"]
utt_line --> deptiers["dependent_tiers:\nVec<DependentTier>"]
utt_line --> health["parse_health: ParseHealthState"]
main --> speaker["speaker: SpeakerCode"]
main --> tiercontent["content: TierContent"]
tiercontent --> linkers["linkers: Vec<Linker>"]
tiercontent --> uttcontent["utterance_content:\nVec<UtteranceContent>\n(28 variants)"]
tiercontent --> terminator["terminator: Option<Terminator>"]
tiercontent --> bullet["bullet: Option<Bullet>"]
The DependentTier enum has 32 variants: structured linguistic
(Mor/Gra/Pho/Mod/Sin/Act/Cod/Wor), with-inline-bullets
(Add/Com/Exp/Gpx/Int/Sit/Spa), text-only
(Alt/Coh/Def/Eng/Err/Fac/Flo/Gls/Ort/Par/Tim),
Phon-project (Modsyl/Phosyl/Phoaln/Xphoint), and UserDefined /
Unsupported.
Three-Level Content Hierarchy
CHAT main-tier content is a tree with three nesting levels. Every content traversal must understand all three.
ChatFile
└── Line::Utterance
└── MainTier
└── TierContent
├── content: Vec<UtteranceContent> ← Level 1
│ ├── Word(Box<Word>)
│ │ └── content: Vec<WordContent> ← Level 3
│ ├── OverlapPoint(OverlapPoint)
│ ├── Group(Group)
│ │ └── BracketedContent
│ │ └── Vec<BracketedItem> ← Level 2
│ ├── PhoGroup, SinGroup, Quotation
│ │ └── (same BracketedContent)
│ ├── Retrace(Box<Retrace>)
│ ├── Pause, Event, Separator, ...
│ └── AnnotatedWord, AnnotatedGroup, ...
├── bullet: Option<Bullet>
├── linkers: Linkers
└── terminator: Terminator
Level 1, UtteranceContent (28 variants)
What you iterate when walking utterance.main.content.content.0:
| Category | Variants |
|---|---|
| Words | Word, AnnotatedWord, ReplacedWord |
| Groups and quoted spans | Group, AnnotatedGroup, PhoGroup, SinGroup, Quotation, AnnotatedQuotation |
| Retraces | Retrace, AnnotatedRetrace |
| CA markers | OverlapPoint, Separator |
| Events | Event, AnnotatedEvent, OtherSpokenEvent |
| Actions | Action, AnnotatedAction |
| Timing | InternalBullet |
| Scope markers | LongFeatureBegin/End, NonvocalBegin/End/Simple, UnderlineBegin/End |
| Other | Freecode, Pause |
Critical rule: every match on UtteranceContent must explicitly
list all 28 variants. No _ => catch-alls. Project policy: silent data
loss when new variants are added is unacceptable.
Level 2, BracketedItem (22 variants)
Content inside groups (<...>, ‹...›, 〔...〕, "..."). Accessed via
group.content.content.0 (the double .content.content.0 is not a
typo, Group.content is BracketedContent, which has .content: BracketedItems, which has .0: Vec<BracketedItem>).
BracketedItem mirrors UtteranceContent closely. Retrace content
(<word word> [/], word [//]) is a dedicated Retrace variant at
both levels, not hidden inside AnnotatedGroup. Groups can nest
arbitrarily deep.
Level 3, WordContent (13 variants)
Content inside a single word token, read through word.content():
| Variant | Example |
|---|---|
Text | plain text segment |
Phonetic | @u phonetic transcription segment |
Shortening | (lo) omitted sound |
OverlapPoint | butt⌈er⌉, overlap inside a word |
CAElement | ↑ ↓ prosody markers |
CADelimiter | ° ∆ paired delimiters |
StressMarker | ˈ ˌ |
Lengthening | : |
SyllablePause | ^ |
CompoundMarker | + in ice+cream |
CliticBoundary | ~ in a cliticized form |
UnderlineBegin/End | scope delimiters |
Key insight: overlap markers can appear at all three levels, as
standalone UtteranceContent::OverlapPoint (space-separated:
⌈ word ⌉), as BracketedItem::OverlapPoint (inside groups), or as
WordContent::OverlapPoint (intra-word: butt⌈er⌉). Any traversal
looking for overlap markers must check all three levels.
Annotated Wrappers and Replaced Words
Annotated<T>
Adds scoped annotations ([/], [* m], [= explanation], etc.) to any
annotatable inner type:
pub struct Annotated<T> {
pub inner: T,
pub scoped_annotations: AnnotatedContentAnnotations, // NEVER empty
pub span: Span,
}
The annotations are never empty, and that is enforced by the type rather
than checked afterwards: AnnotatedContentAnnotations::new returns None for
an empty list, so an annotated wrapper cannot be built without an annotation.
A construct carrying none is the BARE variant instead, and that Option IS the
bare-versus-annotated decision at every construction site.
At Level 1: AnnotatedWord(Box<Annotated<Word>>),
AnnotatedGroup(Annotated<Group>),
AnnotatedEvent(Annotated<Event>),
AnnotatedAction(Annotated<Action>). The same variants exist at Level 2, and
so does every bare counterpart: see
Annotations for the full pairing.
ReplacedWord
Represents word [: replacement], a surface form with a replacement:
pub struct ReplacedWord {
pub word: Word,
pub replacement: Replacement,
}
pub struct Replacement {
pub words: Vec<Word>,
}
Convention when extracting words for NLP depends on the domain. Mor uses replacement words when present because morphology follows the correction. Wor uses the original surface word because timing follows what was spoken.
Tier Domains
Different NLP tasks need different views of the same content. The
TierDomain enum controls which words count for each tier and how
groups are traversed:
| Domain | Used by | Skips | Counts separators? |
|---|---|---|---|
Mor | %mor / %gra generation | Retrace groups | Yes, , „ ‡ carry mor items (cm|cm, end|end, beg|beg) |
Wor | %wor generation, FA | Nothing | No |
Pho | %pho alignment | PhoGroup | No |
Sin | %sin alignment | SinGroup | No |
The content walker takes Option<TierDomain>: Some(domain) for
domain-aware gating, None to recurse everything unconditionally.
Content Walkers
talkbank-model exports closure-based walkers. Two layers:
walk_content: generic, visits all content items (custom traversals).walk_words/walk_words_mut, filtered to words / replaced words / separators, with domain-aware gating. The primary primitive.
use talkbank_model::alignment::helpers::{
walk_words, walk_words_mut,
WordItem, WordItemMut,
TierDomain,
};
walk_words(content, Some(TierDomain::Mor), &mut |leaf| {
match leaf {
WordItem::Word(word) => { /* ... */ }
WordItem::ReplacedWord(replaced) => { /* ... */ }
WordItem::Separator(sep) => { /* ... */ }
}
});
flowchart TD
input["&[UtteranceContent]\n+ domain: Option<TierDomain>"]
dispatch["Match variant\n(24 UtteranceContent variants)"]
word["Word → emit WordItem::Word"]
rw["ReplacedWord → emit WordItem::ReplacedWord"]
sep["Separator → emit WordItem::Separator"]
group["Group / AnnotatedGroup /\nPhoGroup / SinGroup / Quotation"]
gate{"Domain\ngating"}
skip["Skip\n(atomic unit)"]
recurse["Recurse into\ngroup.content"]
input --> dispatch
dispatch --> word & rw & sep & group
group --> gate
gate -->|"Mor: skip retraces"| skip
gate -->|"Pho/Sin: skip groups"| skip
gate -->|"None: recurse all"| recurse
recurse -->|back| dispatch
What walk_words does NOT visit
Only words and separators. Not OverlapPoint (any level), not
CAElement within words, not events / pauses / actions, not internal
bullets. For these, walk with walk_content, which yields every item
kind; extract_overlap_info below is the worked example.
extract_overlap_info, overlap regions
Walks the content with walk_content at the %wor domain, so its word
positions are on the %wor projection’s scale, and pairs every
OverlapPoint at all three content levels (⌈ with ⌉ by index) into
OverlapRegion structs. Used by the alignment pipeline (onset estimation)
and the validator (pairing checks). For whole-file analysis,
analyze_file_overlaps() matches top regions (⌈) with bottom regions
(⌊) across utterances with 1:N support (used by E347 and
chatter debug overlap-audit).
Validation
Beyond what the grammar enforces, validate_with_alignment() checks
semantic constraints:
%moralignment: number of MOR items matches alignable main-tier words.%grastructure: sequential indices, ROOT checks, circular dependency.- Header consistency:
@IDcodes match@Participants. - Speaker references: all
*SPEAKER:codes declared.
Structural alignment and %wor timing state are computed from the same typed
main-tier model:
flowchart TD
main["MainTier content"]
walker["walk_words()\ncount alignable words"]
subgraph "Structural Alignment and Timing State"
mor["%mor\ncustom logic\n(clitic handling)"]
pho["%pho\npositional_align()\n(skip PhoGroup)"]
sin["%sin\npositional_align()\n(skip SinGroup)"]
wor["%wor timing sidecar\nMissing | Drifted | CountMatched\nmain lexical + %wor bullets"]
gra["%gra\nalign to %mor chunks\n(not main tier)"]
end
main --> walker
walker --> mor & pho & sin & wor
mor --> gra
For the alignment algorithms themselves, see Alignment.
Common Pitfalls
- “Consecutive” means in-order traversal, not adjacent array
indices. When CHAT tools speak of “consecutive” or “sequential”
items on the main tier, this always means document order via
recursive traversal, accounting for groups (
<...>), retrace groups (<...> [/]), quotations ("..."), and all other bracketed structures. Never check adjacency in the flatVec<UtteranceContent>, usewalk_wordsor equivalent in-order traversal. - Missing intra-word content. Overlap markers, CA elements, and
other markers can appear inside
Wordcontent. Checking onlyUtteranceContent::OverlapPointmissesWordContent::OverlapPoint(e.g.,butt⌈er⌉,a⌈nd). - Missing annotated variants.
UtteranceContent::AnnotatedWordandAnnotatedGroupwrap inner types inAnnotated<T>and are easy to forget. BracketedContentaccess.Group.content→BracketedContent, with.content: BracketedItems, with.0: Vec<BracketedItem>.- Separator counter sync (Mor domain). Tag-marker separators
(
,„‡) count as NLP words because they have %mor items. Any code counting words in the Mor domain must count these separators too.
Serialization
- CHAT:
WriteChattrait writes any model type back to CHAT format. - JSON: all model types implement
Serialize/Deserialize. Format per the JSON Schema. - JSON Schema: derived via
JsonSchema. Runjust schema-gento regenerateschema/chat-file.schema.json.
Memory and Interning
String-heavy types (PosCategory, MorStem, MorFeature) use
Arc<str> with a global interner, significant memory savings on large
corpora where the same POS tags and lemmas appear thousands of times.
Collections that are typically small use SmallVec for inline storage:
SmallVec<[MorFeature; 4]>: features per word (usually 0-4).SmallVec<[MorWord; 2]>: post-clitics (usually 0-1).
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Transform Pipeline
Status: Current Last updated: 2026-10-07 (commit 5e895791)
The talkbank-transform crate provides high-level pipelines that compose parsing, validation, and serialization into reusable workflows.
Core Pipelines
Transcript construction
build_chat assembles a typed TranscriptDescription into a mutable CHAT
document; validation remains a separate step. Media names accepted by
MediaFilename as HTTP/HTTPS references are preserved verbatim, including
quotes, extensions and Unicode spelling. They are not filesystem paths and
are never fetched. Complete input spelling must pass media representability
admission before local basename/extension reduction; malformed directory text
cannot be hidden by normalization. The resulting local stem is admitted again.
Parse + Validate
The most common pipeline: parse a CHAT file and validate it.
use talkbank_transform::parse_and_validate;
let result = parse_and_validate(source, &parser, &error_collector);
This:
- Parses the source text into a
ChatFileAST - Runs validation (alignment checks, header consistency, etc.)
- Collects all errors and warnings into the
ErrorSink
Source-bound preservation versus replacement
TreeSitterParser::admit_planned_tiers selects a tier-removal policy from the
headers of the same producing parse. It fully validates the retained document;
this does not certify original bytes when replacement was selected.
AdmittedReplacement::into_disposition consumes that result and distinguishes:
AdmittedDisposition::Preserved(AdmittedPreservation): no replacement was selected and no concrete tier was removed. The capability binds complete admission to the exact original source, including its formatting.AdmittedDisposition::Replaced(AdmittedReplacement): only the retained document is admitted. Even a selected replacement that matched no tier cannot manufacture original-source admission through this transition.
AdmittedSourceChat::from_preservation transfers the opaque preservation
capability directly to an unchanged-output proof without parsing again or
accepting independently supplied source/model arguments. Editing still consumes
admission and requires fresh checked construction before output. Recovery and
internal failures cannot produce either successful admission state.
TreeSitterParser::admit_word_timing_plan additionally supports
WordTimingPlan::PreferRetained. Its producer defers concrete %wor candidates
and their own diagnostics while lowering every other component normally.
Complete original admission retains valid timing, including partial timing and
legal stale sidecars; reusability is a separate downstream decision. If original
admission fails, the plan physically omits only the concrete word tiers the
evidence names (their own lowering failed, or a validation error lies inside
their source span), binds those diagnostics to each removal receipt, and
validates the reduced document again; errors located in no word tier remove
every remaining word tier. It requires complete retained admission before
returning a word-timing replacement receipt, never clears recovered-model taint
and never filters diagnostic codes. Internal producer failures cannot select
regeneration. book/src/architecture/wor-timing.md describes the loop.
Components parse and lower once. Word-tier entries move into the document
rather than being copied, and a rejected attempt returns the document with its
diagnostics (ValidationFailure::into_rejection), so no model copy is made;
removal finds each entry again by its own source span. WordTimingPlan::Preserve
gives no exemption and is appropriate for actual pass-through or preservation
policies.
The plan’s phases are types. Lowering consumes a PlanEntry, either already
decided or a header callback, and hands back the decided plan beside the
producing source, so nothing can route a tier or read a selection before the
decision. Each decided plan implements TierRouting: RetainAll (plain
parsing) and TierRemoval (nothing, or a fixed selection with its removal
receipts) never defer a tier, and their deferred type is uninhabited, so their
utterances carry no word candidates by type;
WordTimingDecision::PreferRetained defers word tiers and makes no removal
receipt during lowering. Removal receipts and deferred candidates therefore
never coexist in one plan.
stateDiagram-v2
[*] --> Decided: admit_replacing_tiers, plain parsing
[*] --> AfterHeaders: admit_planned_tiers, admit_word_timing_plan
AfterHeaders --> Decided: every header lowered, callback runs once
Decided --> RetainAll: plain parsing, every tier lowers
Decided --> TierRemoval: route removes selected domains
Decided --> WordTimingDecision: route defers word tiers
TierRemoval --> AdmittedReplacement: complete validation
WordTimingDecision --> AdmittedReplacement: Preserve, complete validation
WordTimingDecision --> WordTimingAdmission: PreferRetained, candidate loop
Source-bound utterance partitions
utterance_split::UtteranceSplitPlan selects morphology-domain or original-word
timing-domain assignments against one borrowed utterance. Execution accepts no
second model or assignment vector. It refuses count drift, disjoint child runs,
boundaries inside indivisible groups/replacements and separator-stranding
boundaries. Child-label magnitudes do not determine allocation sizes.
Execution returns SplitOutcome::Unchanged when every slot names one child, so
the caller keeps the utterance it holds and nothing is copied; only
SplitOutcome::Split rebuilds children.
The shared partition policy retains free-text dependent tiers on the first
child, preserves only count- and lexically corroborated %wor, and derives a
child’s main timing only from complete measured word timing. Original-turn
analysis and uncorroborated word timing are returned as explicit, source-bound
invalidation receipts, not merely logged. No full parent interval is assigned
to a partial child. Rebuilt children still need checked construction before
output; structural partition admission is not whole-file validity.
utterance_split::WordSpeakerSource admits complete source timing before
requesting acoustic inference and keeps its immutable source borrow through
that request. bind_timeline then establishes WordSpeakerSplitPlan, without
rematching timing or accepting another source. The convenience split-plan
constructor performs those same transitions when evidence already exists.
The split plan adds measured word-level speaker ownership to that same partition
owner. It requires complete, lexically corroborated %wor, uses
union-of-held-time attribution from rediarize, and refuses uncovered or tied
words. It never picks the nearest turn or carries a previous speaker through a
gap. Returning to a prior speaker creates a new contiguous child run rather than
reordering words. A single run is WordSpeakerPartition::Relabeled, which names
the owning track and leaves the source for the caller to relabel; only several
runs rebuild children. Execution retains ownership evidence and invalidation
receipts; participant identities, headers and final output admission remain the
caller’s responsibility. The existing whole-turn rediarization API and its
contestation reporting are unchanged.
CHAT → JSON
Convert a CHAT file to its JSON representation:
use talkbank_model::ParseValidateOptions;
use talkbank_transform::{JsonLayout, chat_to_json};
let json = chat_to_json(source, ParseValidateOptions::default().with_validation(), JsonLayout::Pretty)?;
The JSON follows the schema at schema/chat-file.schema.json, and is checked
against it before it is returned.
JsonLayout is Pretty (indented, the CLI’s default) or Compact (one line,
chatter to-json --compact, which chatter’s clap parsing turns into the
value). It was a pretty: bool, which every caller had to
spell as a bare true or false. chat_to_json_named adds the transcript’s
name (so E531 can run), chat_to_json_with_schema_policy adds a
JsonSchemaPolicy, and chat_to_json_unvalidated skips the schema check; all
four take the layout the same way, and the schema policy and the layout are
matched as one pair, so every combination has exactly one serializer.
JSON → CHAT
The JSON produced by chat_to_json is schema-conformant and
round-trips. Deserialize it back into a ChatFile with serde_json
(the model derives Deserialize), then serialize through WriteChat
to reproduce CHAT text:
let chat_file: talkbank_model::ChatFile = serde_json::from_str(json_str)?;
let chat_text = chat_file.to_chat_string();
The chatter from-json command wraps this path
(crates/chatter/src/commands/json.rs, json_to_chat).
CHAT → CHAT (Normalize)
Parse and reserialize to normalize formatting:
use talkbank_transform::normalize_chat;
let normalized = normalize_chat(source, &parser)?;
normalize_chat lives in
crates/talkbank-transform/src/pipeline/convert.rs.
Selective name pseudonymization: planning API under development
pseudonymize::NameMap admits a private caller-supplied mapping.
TranscriptNames::admit_document binds that mapping to parsed, validated source,
including alignment checks. PseudonymizationInput::plan_words creates sensitive
review previews without mutating the input. PseudonymizationInput::prepare_output
can produce an admitted in-memory PseudonymizedDocument, but this is
not yet a user-facing de-identification command. Private receipt persistence,
CLI integration and broader policy/corpus acceptance remain unfinished.
Output preparation refuses unsafe lexical, morphology, timing or pronunciation plans. It applies only their source-bound edit ranges, copying every intervening byte unchanged; applied private receipt entries are built during that same operation. The rewritten source must pass parsing and alignment-aware validation under the input’s filename and rule context. The output parse is reused for a follow-up plan, which must propose no further changes. A refusal retains original review findings but exposes no partial output text. Accepted output still is not a promise of complete de-identification: report-only findings remain visible in its private review and callers must protect both output and receipts.
Lexical reviews retain a producer-assigned WordLocation: zero-based utterance
and original-word indices, plus WordSpelling distinguishing spoken material
from an indexed editorial replacement target. Retraced words count; separators
and targets do not advance the original-word index. These are review coordinates,
not morphology or phonology alignment indices. Aligned morphology, corroborated
timing and pronunciation refusals carry the selecting word’s location through
their existing bindings. Applied lexical edits retain it in EditOrigin; their
EditKind is derived from that origin rather than stored independently. Multiple
source fields removed from one shortening retain the same word location.
Metadata findings and edits carry HeaderFieldLocation (header index plus typed
field); prose uses ProseLocation, which distinguishes a document header from
an utterance’s dependent tier. Header indices count all headers, including
structural ones; dependent-tier indices count all tier kinds. Independently
reported morphology uses LemmaLocation: utterance, morphology-item index and
LemmaPart::Main or a specific post-clitic. It never invents a main-tier match
for an unmatched lemma. Exact source ranges remain available alongside this
context. All indices are zero-based. Persistent CLI receipts are being integrated.
The CLI publication boundary is being implemented separately from the transform.
It stages only admitted PseudonymizedDocument bytes beside an explicit output
destination and uses no-clobber publication. Existing files, directories and
symlinks are not replaced; a collision after staging also refuses publication.
Abandoning preparation removes its temporary file, not the input or destination.
Unix staging files are owner-only; other systems require a protected destination
directory’s inherited ACL. This does not establish a cross-file transaction or
crash-durable directory update.
The receipt owner encloses the low-level publisher. Its PreparedPublication
owns both the staged file and a committed private SQLite receipt; callers cannot
separate that evidence from its output or call the enclosed publisher directly.
The receipt records source/output BLAKE3 identities, the exact applied edits,
typed locations and report-only findings in queryable tables, not JSON blobs.
It starts as prepared, becomes written only after publication and file sync,
or records publication_failed on a no-clobber failure. An interruption or failed
completion update can leave prepared; that state means uncertain publication,
not success or proof of absence. Existing receipt files are never overwritten.
SQL dependencies live in the CLI, not the transform core. The user command
remains unavailable pending command routing and end-to-end acceptance. Refused
plans have separate output_refused receipts: proposals are labeled proposed,
never applied, and output coordinates and identity are absent. Input admission
refusals and missing map entries are distinct states. A derived summary view
distinguishes a mapped document with no findings from an unmapped document.
The authored word-features/pseudonymizer-source.cha / pseudonymizer-expected.cha
pair tests exact output, every untouched gap, cross-tier changes, metadata,
possessive retention, deterministic application and idempotence. Existing
pronunciation, morphology and timing mismatch references also exercise whole-
document refusal. A placeholder that reintroduces a mapped word at a Unicode
boundary is refused during map admission, before document planning. Placeholders
must first be plain CHAT lexical tokens; invalid tokens are refused separately.
The output stability check remains an independent boundary safeguard.
Admission retains the producer-owned ParsedSource alongside the validated
model. TreeSitterParser::parse_chat_file_with_source returns both from one
parse; its diagnostics still require review and its model still requires
validation. The retained CST enables generated source-bound field traversal
without a second parse, raw-line searches, or assuming whole-header serialization
preserves untouched formatting. CST ownership alone is not a validity proof.
Admission refuses both errors and warnings. Canonical controls exercise empty
input, missing participant roles, forbidden controls inside free text, and a
warning-only non-NFC media name paired with its canonical spelling. Even rejected
empty input can retain a CST; its existence never authorizes transformation.
Header-field review covers whole Unicode-bounded names within participant names and
typed @ID group, education and custom fields, including repeated names within
longer metadata text. It uses the same matcher and report-only possessive
policy as prose, not a second whole-field replacement path. Generated field
projections supply exact source ranges; a binding failure is a refusal, not
permission to search raw lines. ParticipantWordRoles owns the name/role
partition for both parser lowering and source-bound review. Speaker codes,
participant roles, spacing and structural pipes are not selected. Each match
has exactly one metadata or prose owner; these are still review findings, not
complete writable pseudonymization.
The current planning policy uses exact-case typed lexical components, reports
case near misses, and derives dependent-tier changes from selected main-tier
words. Main-tier previews retain exact generated text/shortening edits;
each shortening owns its parentheses and inner segment, so it cannot create
overlapping edits. Marker bytes, suffixes, annotations and unselected compound
partners remain outside those ranges. Source association and lexical-piece
corroboration are required: failure creates a refusal instead of an unbound
proposal. The word-features/selective-name-fields.cha reference exercises
lengthening and names on either side of a compound boundary.
Free-text planning uses unicode-segmentation word boundaries on typed prose
payloads, with apostrophe-s possessives such as Rose’s and Rose's reported
rather than partially rewritten. A mapped name may span several segments, such
as Rose-Marie. At each boundary the longest recognized spelling takes
precedence, including report-only near misses and possessives: mapping both
Rose and Rose-Marie never partly rewrites Rose-marie as Rose.
The headers/unicode-name-boundaries.cha reference exercises this policy in
participant names, ID metadata and prose. LexicalPlan::free_text returns private
source-bound exact matches, case near misses and possessive review findings.
The CST supplies prose segments within descriptive headers and free-text
dependent tiers (including user-defined tiers); timing bullets, identifiers,
configuration headers and structured tiers are not treated as prose. This does
not tokenize raw CHAT structure. Header-field review above reuses this matcher
while retaining its own typed metadata classification.
The headers/free-text-names.cha reference covers accented names, punctuation,
continuation lines, timing, possessives and several dependent-tier kinds.
Dependent-tier proposals are driven by the selected main-tier
words. Morphology requires established lemma correspondence. Edit locations
are not inferred from text searches: each morphology proposal retains
the generated, source-bound main-lemma field from the same admitted parse.
Tier item counts and lemma values corroborate the association; failure refuses
the proposal. Post-clitic fields are not mistaken for subsequent main items.
The canonical tiers/mor-name-source-fields.cha example exercises repeated
names, both shortening forms, uneven spacing and a preceding post-clitic item.
This establishes exact edit locations, not writable document output.
Unchanged mapped lemmas, including post-clitics, are reported separately as
exact matches or case near misses; the lemma scan cannot authorize an
independent replacement. Compound
components without proven lemma-part correspondence produce a refusal.
Timing uses the %wor corroboration transition. Words within a
timing tier receive the selected main word’s existing component decisions, not
another name-map lookup. Equal cleaned spellings with different compound
structure refuse the change. Exact timing-word lexical edits preserve markers
and leave timing bullets outside their ranges; a failure refuses the tier’s
proposal rather than exposing a partial list. The selective-name reference
also covers a name that occurs only in the timing tier: it is not independently
replaced.
Words within a phonological group share the group’s alignment position.
Pronunciation findings retain the
original %pho/%mod item and refuse automatic output; structured Phon
companions require correspondence review. No tier is silently dropped and no
orthographic placeholder is treated as a pronunciation. These review objects
contain protected transcript content and must not be logged by default.
Validation + Roundtrip Cache Lifecycle
The following diagram shows the full validation and roundtrip pipeline, including the cache layer:
flowchart TD
file["CHAT file"]
cache{"Cache\nhit?"}
parse["Parse\n(tree-sitter → AST)"]
validate["Validate\n(per-file → per-utterance →\nmain tier → dependent tiers)"]
rt{"Roundtrip\nflag?"}
ser1["Serialize → CHAT text"]
reparse["Reparse CHAT text"]
ser2["Serialize again"]
cmp{"Two\nserializations\nmatch?"}
store["Store in cache\n(SQLite)"]
pass["Pass"]
fail["Fail"]
cached["Return cached result"]
file --> cache
cache -->|miss| parse --> validate --> rt
cache -->|hit| cached
rt -->|yes| ser1 --> reparse --> ser2 --> cmp
rt -->|no| store --> pass
cmp -->|yes| store
cmp -->|no| fail
Streaming Parse
For large files or interactive use, the transform crate supports streaming parse where utterances are processed incrementally rather than loading the entire AST into memory.
The shared validation runner (every frontend, one engine)
All bulk validation, whatever the frontend, flows through the
validation_runner module’s two streaming entry points in
crates/talkbank-transform/src/validation_runner/:
validate_directory_streamingwalks a directory and feeds every CHAT transcript to a worker pool;validate_files_streamingruns an explicit file list through the same worker pool.
The one worker pool and the one walk
The pool is talkbank_transform::worker_pool::fan_out, and it is the only
pool: chatter to-json fans its directory conversions out through it too.
flowchart LR
items["items (caller's iterator)"] -->|"calling thread feeds"| queue["bounded queue, 2 x width"]
queue --> w1["worker 1<br/>16 MiB stack"]
queue --> wn["worker n<br/>16 MiB stack"]
w1 --> join["join every worker"]
wn --> join
join --> run["PoolRun { results, outcome }"]
- Width is
--jobs, else the machine’s parallelism, never zero; serial work is a pool of width one, not a second code path. - Every worker runs on a
CHAT_THREAD_STACK_BYTES(16 MiB) stack, the size of the CLI’s program thread, because parsing and validation recurse with the data. - Each worker RETURNS its result (the validation runner’s per-worker tally,
to-json’s counts); the caller combines them after the join. Nothing reads
shared counters whose exactness rests on a comment. Workers borrow what
they share (the event sender, the cancellation latch, the cache and the
configuration) through a
WorkerContextof references, because the pool’s threads are scoped: noArcclones, no per-worker config copy. - A worker that unwinds is joined and counted as
PoolOutcome::SomeUnwound; a thread that cannot be started isPoolOutcome::CouldNotStart, with nothing fed. The caller measures what was lost against what it fed. - Stopping early is the caller’s iterator (the runner feeds
take_while(not cancelled)); workers stop by returning.
Discovery is talkbank_transform::paths. Every walk descends every level.
walk_files returns every file found WITH its path relative to the walked
root (a FoundFile), built as the walk descends, plus every entry that could
not be read. keep is given each entry’s file name, so a rejected entry
never has its path built. walk_transcripts returns each transcript as a
FoundTranscript: its relative path and its StoredTranscript, whose name
is the listing entry that found it, so nothing lists the directory again to
learn the stored name (a stem that is not UTF-8, which validation could not
name, is a failure of the walk). What happens at a symbolic link is the
caller’s choice, a Links value:
Links::Follow(every reading walk): a link is walked as what it points to, a directory reached twice is walked once, and a link whose target is gone is a failure whatever its name: nothing says whether it was a file or a directory, and a link to an unmounted volume’s subcorpus hides every transcript under it.Links::Skip(to-json --prune, the one walk whose results are deleted): links are neither walked nor kept. Following them would let--prunedelete JSON outside--output-dirthrough a linked directory, and then remove that directory’s emptied parents.
A directory the walk cannot list is always a failure, never skipped, even an
operating system’s own folder at a volume root (.Spotlight-V100,
System Volume Information, lost+found): nothing can say whether it held
transcripts. Walk the corpus directory, not the volume root.
expand_transcript_arguments is the one expansion of command-line paths:
each file argument is resolved to its stored name by one
StoredNameResolver (each parent directory listed once), and each directory
argument contributes its walk’s transcripts, so the result is
StoredTranscript values. The runner’s work queue carries those values,
and validate_files_streaming resolves a plain path list the same way
before any worker starts, so no worker resolves a name. Nothing is skipped
silently, and there is one policy for what cannot be
read: it is a FileStatus::ReadError in the run’s own results, counted in
its totals, so the run fails and says which paths. The directory entry
point does this for its walk, and validate_arguments_streaming for
command-line arguments (chatter validate), so the unreadable path reaches a
JSON consumer as a record. Commands that are not streaming runs (fix,
to-json, the debug commands) refuse an incomplete input before
processing anything.
Both share one worker loop, so every consumer gets identical rule
coverage (including the file-stem-dependent checks such as the @Media
filename match), identical stats accounting, and the same on-disk cache.
The chatter CLI, the TUI, and the desktop app all call these
entry points. The invariant to preserve: no frontend grows its own
validation orchestration; a file must validate identically whether
selected alone or reached by a directory walk.
What a run checks: typed, end to end
ValidationConfig says what each worker checks with typed values, never
booleans: alignment: AlignmentValidation (Structure or
IncludeTierAlignment) and roundtrip: RoundtripCheck (Skip, the default,
or Run). A frontend parses its flags into these once, at its boundary:
chatter’s clap arguments are Flag<AlignmentValidation> and
Flag<RoundtripCheck> (see “Flags parsed into modes” below), and the desktop
translates its request’s checkbox where it deserializes it. The CLI’s
ValidationRules, the TUI and the desktop runner then carry the same values
the worker matches on; the libraries have no bool-to-mode constructors, and
there is one roundtrip type from flag to worker.
The CLI’s side: one presentation, one renderer, one ending
chatter validate resolves its output flags once, in dispatch, into a
ValidationPresentation, then runs one event loop:
flowchart TD
flags["--format, --quiet, --audit, --tui-mode"] --> resolve["ValidationPresentation::resolve\n(conflicts are usage errors)"]
resolve -->|Tui| tui["TUI loop\n(same runner, same error limit)"]
resolve -->|"Streamed(Lines / Json / Audit)"| renderer["one ValidationRenderer"]
renderer --> loop["event loop: discovering, started,\none FileComplete per file (status + diagnostics)"]
loop --> end["the run's RunEnding\n(Aborted(NoEnding) if the stream closed without one)"]
tui --> end
end --> finish["renderer.finish(&RunEnding): exhaustive"]
end --> exit["ValidationOutcome::failed(): !RunEnding::passed()"]
The runtime holds no output channel of its own: every fact a run has to say
(a stop, a loss, an abort, an input with nothing in it) is a RunEnding the
renderer’s finish matches, so JSON mode can keep stderr empty by
construction rather than by remembering to. The TUI shows the same ending
and returns it with the session (InteractiveEnd::Closed(RunPhase)), so
the exit status is the run’s on every surface: only RunEnding::passed
exits 0, and a session closed before its run ended fails. The TUI lists
every file that failed (its diagnostics, or why it could not be read, why
its roundtrip failed, or the tool failure) and shows the run’s notices and
cache events, as the other surfaces print them.
Flags parsed into modes
A presence flag selects one of two values of a typed mode. chatter’s
cli/args/flag_modes.rs has one generic clap argument, Flag<M>, for any
M: FlagMode (the flag’s name, help, optional short form, and the value
when it is absent or present), so a command’s field is
#[command(flatten)] alignment: Flag<AlignmentValidation> and its handler
receives the mode. FlagMode is chatter’s own trait, so it can be
implemented for library types (AlignmentValidation, RoundtripCheck,
JsonLayout, JsonSchemaPolicy) as well as chatter’s (FixMode,
ClearMode, CacheRefreshMode, JsonRefresh, OrphanJson,
AlignmentView). The translation from flag to mode therefore lives in the
CLI, and the library crates expose no constructor from a bool.
Two flags that select one value have their own Args impls, which refuse
the combinations that would leave one flag without effect: ToJsonCheckArgs
(--skip-validation conflicts
with --skip-alignment) and NormalizeCheckArgs (--skip-alignment
requires --validate), each yielding one CheckLevel.
How a run ends, and who decides
While its receiver remains connected, every stream ends with exactly one
ValidationEvent::Finished(RunEnding), and the runner decides which:
Complete(stats): at least one file was discovered and every one was accounted for. The only basis for a claim about the whole input, though still not “all files valid”:RunEnding::passedadds that no file was invalid, unreadable or a tool failure. A run that was told to stop after its last file had already been taken also ends here, because nothing was left to stop.NothingFound: the input named no transcript and nothing unreadable. It has no counts. The runner’sNonZeroUsize::new(total_files)is the one place this is recognised, so every snapshot’stotal_filesis non-zero.Stopped { stats, reason }: the run was told to stop and leftstats.missing_files()(non-zero) files unvalidated.reasonis aCancelReason:ErrorLimit { limit }(the run’s own error limit) orRequested(the caller’sCanceller). ItsDisplayis the one wording every surface uses.Incomplete { stats, cause }: the run reached its end without covering everything it discovered, and no requested stop explains it.causeis aLossCause:WorkerFaults, every worker failure the run observed (the pool’s ownPoolFault,UnwoundorThreadRefused, andParserUnavailable, none hiding another), orUnexplained, a runner defect.statsdescribes only what was processed.Aborted(reason): no totals at all. A drop guard on the orchestrating thread sendsAborted(Panicked)during an unwind, so a panicking run terminates its stream instead of closing it in silence; a consumer whose stream closes with no ending anyway reportsAborted(NoEnding).
passed() is the single answer to “may this run be reported as a success”:
the CLI’s exit status, the TUI’s header color and the desktop’s all-valid
claim all read it. Counted endings require producer-admitted CompleteStats
or PartialStats, each exposing read-only snapshot() counts. A snapshot
cloned from a stopped run cannot construct Complete, even when every
processed file passed. The partial payload owns its derived, nonzero
missing_files() count; callers cannot supply a contradictory shortfall.
RunEnding::stats() retains the common read-only snapshot view. Text, JSON
and desktop event formats are unchanged; Rust consumers matching an ending
use the admitted payload’s accessors instead of the removed count fields.
The ending of a run that reached its end is decided from three facts, each from its owner:
flowchart TD
cov{"stats.coverage()"}
cov -->|Complete| complete["Complete(stats)"]
cov -->|"Shortfall(partial)"| faults{"WorkerFaults::observe"}
faults -->|"Some(faults)"| faulted["Incomplete { cause: WorkerFaults }"]
faults -->|None| stop{"latched stop?"}
stop -->|"Some(reason)"| stopped["Stopped { stats: partial, reason }"]
stop -->|None| unexplained["Incomplete { cause: Unexplained }"]
A worker fault outranks a stop: files a worker took and never finished are
lost, and a worker that could not create its parser took none, whatever else
happened. Each worker returns its tally or a WorkerSetupFailure, so a
worker that could not start is a value the runner sees rather than an empty
tally indistinguishable from an idle worker. WorkerFaults::observe, the
one constructor, reads the pool’s outcome and those setup failures and
answers None when every worker started and returned, so a WorkerFaults
is never empty.
Stopping a run: the Canceller and the error limit
There are two ways to stop a run early, and both end in the same latch:
- the caller’s
Canceller(returned by both streaming entry points; the CLI’s Ctrl-C handler, the TUI’sckey and the desktop’s Cancel button hold one), whose only operation iscancel(); - the run’s own
ValidationConfig::error_limit, anErrorLimit(UnlimitedorStopAfter(n)), which the runner counts.
The error limit is the runner’s, not a consumer’s. Each worker spends
FileStatus::errors_found() from a shared ErrorBudget before taking its
next file: the file’s Severity::Error diagnostics after suppression, or 1
for a failed roundtrip. Warnings never count. The addition that reaches the
limit latches CancelReason::ErrorLimit, so the worker that crossed it stops
at once and the others at their next file. With one worker the stop is
exact: a limit of 1 stops before the second file starts.
The count lives in the runner, not in a consumer reading rendered events,
so warnings cannot spend the limit, a run that covered every file is never
reported as stopped, and the text, JSON, audit and TUI presentations of
chatter validate all honour the same limit. The desktop app sets no limit.
ValidationEvent and RunEnding are deliberately NOT #[non_exhaustive].
Adding a variant breaks external consumers on purpose: a new ending that a
consumer silently ignores is precisely the defect these types exist to
prevent, so a downstream crate gets a non-exhaustive match error and decides
for itself what the new ending means. The compiler-checked
validate_directory_streaming rustdoc example shows an exhaustive event
loop. Keep the Canceller alive while consuming the stream (dropping it is
not a cancel), and treat channel closure without an ending as
Aborted(NoEnding), never as successful validation.
Dropping the result receiver is also a stop boundary. Once a worker fails to deliver a file-completion event, it must not start another queued file. This applies to fresh validation, cached results and file-read failures alike. The already completed attempt is accounted for before delivery; disconnection does not turn an unreadable input into a parser error or a cacheable result. With multiple workers, other in-flight attempts may finish before observing closure. The reference-backed lifecycle contract synchronizes at a cache lookup and checks subsequent work counts, rather than relying on sleeps or timing.
Two design points worth keeping:
- Every short ending is a VARIANT, not a field. A
lost: usizeor acancelled: boolbeside the totals would be something every consumer must remember to check, and forgetting yields a false clean bill of health: files abandoned by a crashed worker contribute to no counter, so partial totals look immaculate. A 500-file corpus could validate 480 and report “all valid”. - Loss is DERIVED, not counted.
ValidationStatsSnapshot::coveragereconcilestotal_filesagainst the per-file counters in one place, so there is no third counter free to drift from the two it reconciles. Whether a shortfall was requested is not a count: the snapshot reports onlyRunCoverage::{Complete(complete), Shortfall(partial)}, and the runner decides stop versus loss from the latch and the pool. A requested shortfall is not lost data, and reporting it as such would make the incompleteness report routine, and therefore ignored.
Caching
The transform layer integrates with a file-system cache. Validation results are keyed by content hash, so unchanged files skip re-validation. Cache location is platform-specific: ~/Library/Caches/talkbank-chat/ (macOS), ~/.cache/talkbank-chat/ (Linux), %LocalAppData%\talkbank-chat\ (Windows).
Use --force to bypass the cache for specific paths.
Error Collection
Pipelines use the ErrorSink trait for error reporting. Callers can provide:
- A collecting sink (gathers all diagnostics for batch output)
- A printing sink (writes diagnostics to stderr in real-time)
- A custom sink (for LSP diagnostics, JSON output, etc.)
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Merge Pipeline, Domain Types
chatter merge,chatter pipelineandchatter batchare not CLI commands. Where this page names them it describes the operation of the structural library intalkbank-transform, which is current. See the removal notice.
Status: Draft Last modified: 2026-10-07 (commit 5e895791)
This page specifies the typed Rust vocabulary shared by chatter merge,
chatter speaker-id, the override-file reader/writer, and the
adjudication tooling (CLI today; a VS Code or web UI would share the
same types). It is a design-first specification against the user
contract in chatter merge and
chatter speaker-id. The
implementation has shipped, and this page records the shipped form: where
the implementation departs from the design (the owning crate, several type
names, and the schema-v2 per-speaker role map), the affected section says so
explicitly.
The design follows the cross-cutting rules in this repo’s root
AGENTS.md:
newtypes over primitives at every stable boundary; no boolean
blindness; no tuple-packed seams; typed errors via thiserror;
deterministic BTreeMap/BTreeSet over hash maps for
serialized state.
Where the types live
The merge-pipeline types live in
crates/talkbank-transform/src/speaker_id/, co-located with the
algorithms (identify_mapping, apply_mapping) that produce and
consume them, and are re-exported at
talkbank_transform::speaker_id::* (see that module’s mod.rs). The
structural-merge error type (MergeError) lives beside the merge
algorithm in crates/talkbank-transform/src/transcript_merge.rs.
Design history. The original design placed the types in a new
talkbank-model::merge module, on co-location-with-CHAT-types and
lightweight-dependency grounds. That module was never created: the
implementation kept the types next to the algorithms whose invariants
they encode, in talkbank-transform. talkbank-model still owns the
CHAT-domain vocabulary the merge types reference (SpeakerCode,
ParticipantRole, ParticipantEntry, IDHeader, ChatFile); a
consumer that wants the merge types depends on talkbank-transform,
which the CLI, LSP, and desktop app already do.
Designed vs shipped (quick map)
The sections below preserve the original type specification, updated
in place for the types most central to the override-file contract.
This table maps each designed name to what actually shipped, so a
reader grepping the codebase finds the right symbol. All shipped
paths are relative to crates/talkbank-transform/src/.
| Designed (this page, 2026-05) | Shipped |
|---|---|
InsertedRole | InsertedRoleSpec (speaker_id/override_file.rs): on-disk code / tag strings plus optional specific_role |
MappingAction | SpeakerAction (speaker_id/override_file.rs): Rename / Drop |
DecisionMode | OverrideMode (speaker_id/override_file.rs): Auto / Explicit / Override |
SpeakerMapping (single shared inserted_role) | On disk: MergeOverride.mapping plus the per-speaker MergeOverride.adult_roles map (schema v2). In memory: MappingSpec = HashMap<SpeakerCode, SpeakerAssignment> (speaker_id/mapping.rs), each Rename carrying its own code / role / specific-role |
Margin enum (Finite / Unbounded) | ConfidenceMargin::{NoInformation, Finite, Unbounded} (speaker_id/types.rs); the stable JSON evidence report uses corresponding tagged states |
JaccardScore (fallible serde newtype) | JaccardScore(f64) (speaker_id/types.rs), privately constructed from admitted multiset counts; on-disk override scores remain bare f64 values |
ConfidenceThreshold (associated DEFAULT) | ConfidenceThreshold(f64) with checked new and FromStr boundaries (speaker_id/types.rs) plus DEFAULT_CONFIDENCE_THRESHOLD (speaker_id/identify.rs) |
RetainSet newtype | Not shipped; merge_chat_files takes retain: &[SpeakerCode] (transcript_merge.rs) |
MergeFlag enum | Not shipped; MergeOverride.flags is Vec<String> |
OperatorId / SessionId newtypes | MergeOverride.operator is String; override entries are keyed by String session IDs (a SessionId newtype exists in the speaker_id/judgment/ submodule for the LLM-judgment surface) |
OverrideFile::CURRENT_SCHEMA_VERSION = 1 | Module-level CURRENT_SCHEMA_VERSION: u32 = 2 (speaker_id/override_file.rs) |
SpeakerIdError / MergeError / OverrideFileError variant sets | Shipped with revised variants; see the updated Error types section below |
Existing types reused (not redefined)
The finite reference workflow exercises lexical evidence using unchanged basic conversation and phonological-group CHAT. Self-comparison establishes exact multiset support; raising the threshold refuses the same observation without discarding its evidence. Nested-group speech also distinguishes an unbounded margin (positive winner, empty runner-up) from no information (both scores zero). The stable report serializes these as tagged states, not a JSON infinity or a fabricated finite ratio. Speaker-removal replay exercises missing-reference and single-donor refusals. These checks establish lexical and wire contracts, not that lexical similarity alone proves speaker identity. The spec corpus also passes its actual parse/validation failures through the input-rejection report for both donor and reference roles. The report must preserve the failure phase and ordered diagnostic codes from required admission without inventing a lexical match report. These corpus controls witness parse and ordinary validation refusal; the incomplete-validation adapter remains a separate coverage obligation. Speaker sampling likewise bounds operator head/tail limits against actual turns. A borrowed head and a suffix of the remaining slice make overlap impossible without adding the two limits; maximum-sized budgets cannot overflow or duplicate turns. Reference controls cover anchor-first ordering, head-only, tail-only, overlapping and zero windows, and CJK character caps. Character caps count Unicode scalars rather than UTF-8 bytes. No external judgment provider is needed to verify these deterministic preparation contracts.
Judgment-context controls combine unchanged reference headers with explicitly
authored sidecar records. Missing records and omitted fields preserve unknown
labels; a configured age takes precedence over header fallback. Label admission
rejects blank strings while preserving nonblank spelling and role order. Invalid
labels cannot be constructed through public tuple fields: Rust callers and JSON
both use TryFrom<String>, with read-only as_str() access afterward. Corrected
sample-type verdicts use that same boundary after trimming their wire payload;
their existing blank-correction error remains unchanged. Invalid
age shapes and unknown sidecar fields are rejected, not silently treated as
missing metadata. These controls use references with at most one declared age;
they do not settle which participant should supply age in a multi-age session,
nor infer consent or participant facts from the authored labels.
Header fallback consumes the model’s typed age components, not a second parser
over retained text. Unsupported ages (including invalid day text) and missing
months stay unknown. Widening bounded year/month components makes total-month
arithmetic safe; an explicit sidecar age avoids fallback entirely. This projection
does not replace CHAT validation or select the intended child among several ages.
| Type | Defined in | Used as |
|---|---|---|
SpeakerCode | talkbank-model::model::header::codes::speaker | Identifier for *<CODE>: speakers, dictionary keys in mappings, --retain set elements |
ParticipantRole | talkbank-model::model::header::codes::participant | Role-tag in @Participants and @ID (Target_Child, Investigator, Mother, etc.) |
ParticipantName | talkbank-model::model::header::codes::participant | Optional participant name in @Participants |
ParticipantEntry | talkbank-model::model::header::codes::participant | Single @Participants row |
IDHeader | talkbank-model::model::header::id | Single @ID row |
ChatFile | talkbank-model::model::file::chat_file::core | The merge stages’ inputs and outputs (mutable, without a validity proof) |
None of these are redefined; the speaker_id and transcript_merge
modules import and reference them.
New types (specification)
The subsections below are the type specification. The ones central
to the override-file contract (InsertedRoleSpec, SpeakerAction,
the speaker-mapping pair, OverrideMode, MergeOverride,
OverrideFile, and the three error enums) describe the shipped form.
The remaining subsections
(JaccardScore, ConfidenceThreshold, Margin, RetainSet,
MergeFlag, OperatorId, SessionId) describe the design; where the shipped form differs, the designed-vs-shipped table
above is authoritative for the current symbol and shape.
LexicalMatchEvidence and the recorded report types carry absolute support so
it cannot be discarded.
JaccardScore
A multiset-Jaccard similarity value, by construction in the closed
range [0.0, 1.0].
/// Multiset Jaccard similarity between two bags of tokens.
///
/// By construction in [0.0, 1.0]. `JaccardScore::zero()` is the
/// no-overlap point; `JaccardScore::one()` is identical-bag.
///
/// Used by the speaker-id stage to score how well each donor
/// speaker matches a reference anchor's content.
#[derive(Clone, Copy, Debug, PartialEq, PartialOrd, Serialize, Deserialize, JsonSchema)]
#[serde(try_from = "f64", into = "f64")]
pub struct JaccardScore(f64);
impl JaccardScore {
pub fn new(v: f64) -> Result<Self, JaccardScoreError>;
pub fn zero() -> Self;
pub fn one() -> Self;
pub fn value(self) -> f64;
}
impl Display for JaccardScore { /* "0.735" three-digit */ }
impl TryFrom<f64> for JaccardScore { /* validates range */ }
impl From<JaccardScore> for f64 { /* infallible widen */ }
The shipped type has no public scalar constructor. It is born from admitted
multiset intersection and union counts and exposes only value(). This is
stronger than validating an arbitrary scalar after the relationship that
produced it has already been discarded.
ConfidenceThreshold
The minimum Jaccard margin (winner / loser) the speaker-id stage
will auto-accept. By construction in [1.0, ∞), a threshold of
< 1.0 makes no sense (means the loser scores higher than the
winner, which can’t happen). Default 2.0 per the empirical
calibration recorded in
chatter speaker-id.
#[derive(Clone, Copy, Debug, PartialEq, PartialOrd, Serialize, Deserialize, JsonSchema)]
#[serde(try_from = "f64", into = "f64")]
pub struct ConfidenceThreshold(f64);
impl ConfidenceThreshold {
pub const DEFAULT: Self = Self(2.0);
pub fn new(v: f64) -> Result<Self, ConfidenceThresholdError>;
pub fn value(self) -> f64;
}
impl Default for ConfidenceThreshold {
fn default() -> Self { Self::DEFAULT }
}
Margin
The decisive ratio between the highest-scoring speaker and the
runner-up. Distinguished from ConfidenceThreshold by intent
(this is observed; the threshold is configured) and from
JaccardScore by range (margin is ≥ 1.0; score is ≤ 1.0).
Uses an enum rather than a bare float to model both divide-by-zero and the
zero/zero no-information case. The shipped type is ConfidenceMargin, with
NoInformation, Finite(FiniteConfidenceMargin), and Unbounded variants.
/// Ratio of winning speaker's score to runner-up's score.
///
/// `Finite(r)` for `r >= 1.0`. `Unbounded` when the runner-up
/// has zero score (winner scored anything, runner-up scored
/// nothing). Compares meaningfully against `ConfidenceThreshold`
/// regardless of variant.
#[derive(Clone, Copy, Debug, PartialEq, Serialize, Deserialize, JsonSchema)]
#[serde(untagged)]
pub enum Margin {
Finite(f64),
/// Serialized as the JSON/TOML string "unbounded"; never as
/// f64::INFINITY (which round-trips inconsistently).
Unbounded,
}
impl Margin {
pub fn from_scores(winner: JaccardScore, loser: JaccardScore) -> Self;
pub fn meets(self, threshold: ConfidenceThreshold) -> bool;
}
impl Display for Margin { /* "3.81x" or "∞" */ }
RetainSet
The set of speaker codes specified by --retain on chatter merge.
A BTreeSet<SpeakerCode> wrapped in a newtype so the type
signatures of merge functions communicate intent. Empty is
allowed (means “no speakers come from File 1; File 1 contributes
only headers”, a degenerate but legal case).
/// Speakers whose utterances come from the first input to
/// `chatter merge`. All other speakers come from the second
/// input.
#[derive(Clone, Debug, Default, PartialEq, Eq, Serialize, Deserialize, JsonSchema)]
pub struct RetainSet(BTreeSet<SpeakerCode>);
impl RetainSet {
pub fn new() -> Self;
pub fn from_iter<I: IntoIterator<Item = SpeakerCode>>(it: I) -> Self;
pub fn contains(&self, code: &SpeakerCode) -> bool;
pub fn iter(&self) -> impl Iterator<Item = &SpeakerCode>;
pub fn is_empty(&self) -> bool;
}
impl FromStr for RetainSet {
type Err = RetainSetParseError;
/// Parses `"CHI,SI2"` → `{CHI, SI2}`. Empty entries rejected.
fn from_str(s: &str) -> Result<Self, Self::Err>;
}
InsertedRoleSpec (designed as InsertedRole)
The CHAT identity recorded for one renamed speaker: a speaker code, a
standard role tag, and (only when needed) a specific-role label. A
struct rather than separate function arguments because the triple is
meaningful as a unit (in TOML override files it serializes as an
inline table; on the CLI a CODE:TAG pair parses into one). Shipped
in speaker_id/override_file.rs under the name InsertedRoleSpec,
with on-disk String fields (this is the serialized form written
into override files) rather than the designed SpeakerCode /
ParticipantRole newtypes; MergeOverride::to_mapping_spec lifts
the strings back into the typed CHAT primitives at the read boundary.
/// Inline-table form of the inserted-role spec recorded in each
/// override entry.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct InsertedRoleSpec {
/// CHAT speaker code (e.g. `INV`, or `INV1` when disambiguated from
/// a same-role collision).
pub code: String,
/// CHAT standard role tag (e.g. `Investigator`).
pub tag: String,
/// Specific-role label for `@Participants`' name/specific-role slot
/// (e.g. `First_Investigator`), set only when two adults in the same
/// judgment share `tag` and need the CHAT manual's `CHI1`/`CHI2`-style
/// disambiguation. `None` for the ordinary single-adult-per-role case.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub specific_role: Option<String>,
}
The specific_role field is never operator-typed: it is filled by
the same-role auto-disambiguation described under the speaker-mapping
section below. On the CLI, --inserted-role INV:Investigator and
each OLD=CODE:ROLE assignment in --mapping supply the code / tag
pair; both halves are required.
SpeakerAction (designed as MappingAction)
What happens to a particular speaker in the input. Enum (not
boolean) to avoid blindness. Shipped in
speaker_id/override_file.rs under the name SpeakerAction.
/// Action applied to one speaker in the input file.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum SpeakerAction {
/// Rename the speaker per its own entry in `adult_roles`.
/// Rewrites speaker codes on every utterance and the
/// corresponding @Participants and @ID entries.
Rename,
/// Remove this speaker's utterances and its @Participants /
/// @ID rows entirely.
Drop,
}
The TOML serialization uses "drop" / "rename" lowercase
strings, matching the override-file format documented in
merge-overrides.md.
The design left room for a future RenameTo { code, tag } variant;
that never became necessary, because schema v2 instead resolves every
Rename through the per-speaker adult_roles map, which carries
each speaker’s own target identity (next section).
Speaker mapping: on-disk mapping + adult_roles, in-memory MappingSpec (designed as SpeakerMapping)
The decision record produced by the speaker-id stage and consumed by
the apply step. Carries enough information to apply deterministically
to a ChatFile. The original design was a SpeakerMapping struct
with a single shared inserted_role: InsertedRole field and the
constraint “all renamed speakers go to the same role in v1 of this
schema”. Schema v2 replaced that constraint with a per-speaker role
map, and the shipped code splits the concept into an on-disk shape
and an in-memory shape.
On disk, two sibling fields of MergeOverride
(speaker_id/override_file.rs):
/// Per-donor-speaker-code role assignment, for every speaker whose
/// `mapping` action is `Rename`. Invariant: every `Rename` key in
/// `mapping` has a matching entry here.
pub adult_roles: BTreeMap<String, InsertedRoleSpec>,
/// Map from input speaker codes to actions. Every speaker that
/// exists in the input must appear here.
pub mapping: BTreeMap<String, SpeakerAction>,
Every Rename resolves via that speaker’s own adult_roles
entry, so one entry can rename two speakers to two different roles
(PAR0 -> INV:Investigator, PAR1 -> FAT:Father). When two adults
in the same session are assigned the same role, the writer
auto-disambiguates per the CHAT manual’s CHI1/CHI2 convention:
numbered speaker codes (INV1, INV2), the shared standard role tag
unchanged, and ordinal specific-role labels (First_Investigator,
Second_Investigator, falling back to bare numerals past Fourth)
recorded in each spec’s specific_role field
(speaker_id/judgment/consume.rs, disambiguate_adult_roles). A
hand-edited file that records a Rename with no matching
adult_roles entry fails closed at replay time with
SpeakerIdError::OverrideRenameMissingRole; the sanctioned
constructors (MergeOverride::auto_decision,
MergeOverride::operator_decision) maintain the covering invariant.
In memory (speaker_id/mapping.rs), the apply step consumes a
typed per-speaker assignment map:
/// What to do with a speaker named in the input file.
pub enum SpeakerAssignment {
/// Drop the speaker entirely.
Drop,
/// Rename the speaker to `code` with role tag `role` (and an
/// optional specific-role label for `@Participants`).
Rename {
code: SpeakerCode,
role: ParticipantRole,
specific_role: Option<ParticipantName>,
},
}
/// Operator-supplied mapping from input speaker codes to
/// post-relabeling assignments.
pub type MappingSpec = HashMap<SpeakerCode, SpeakerAssignment>;
MergeOverride::to_mapping_spec converts the on-disk pair into a
MappingSpec for apply_mapping; parse_mapping_spec builds one
directly from the CLI --mapping string. The on-disk contract
requires every speaker that exists in the input to appear in
mapping (we want every decision to be explicit). Note a shipped
gap: apply_mapping currently passes through unchanged any speaker
absent from the in-memory MappingSpec; enforcing the
every-input-speaker precondition at apply time is a documented
follow-up (speaker_id/apply.rs).
OverrideMode (designed as DecisionMode)
How a MergeOverride entry came to exist. Three variants matching
the three speaker-id operation modes. Shipped in
speaker_id/override_file.rs under the name OverrideMode.
/// How a speaker-id decision was made.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum OverrideMode {
/// Reference-mode auto-decide above confidence threshold.
Auto,
/// Operator-supplied `--mapping` (typically after a low-confidence
/// reference-mode attempt).
Explicit,
/// Replay of a prior decision read from another override file.
Override,
}
MergeFlag
Extensible operator-supplied flags on an override entry. Closed
variants for known cases plus a Custom(String) escape hatch.
#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize, JsonSchema)]
#[serde(rename_all = "kebab-case")]
pub enum MergeFlag {
/// ASR diarization mixed multiple real-world roles into one
/// speaker label. The rename may still be the best available
/// approximation but the output is imperfect.
DiarizationMixed,
/// The operator could not confidently determine which speaker
/// is which; mapping is best-guess.
BestGuess,
/// Open variant for contributor-specific flag vocabulary.
/// Serializes as the inner string verbatim.
#[serde(untagged)]
Custom(String),
}
OperatorId
Who made the decision. String newtype.
string_newtype!(
/// Identifier of the operator who created an override entry.
/// Free-form; typically a username or initials. Recorded as
/// audit trail.
pub struct OperatorId;
);
SessionId
Identifies an entry within an override file. Typically the
basename stem of the input CHAT file, but the override-file
schema doesn’t constrain its shape, contributors may use any
stable identifier they like (<participant>-<timepoint>,
<recording-id>, etc.).
string_newtype!(
/// Identifies a session within an override file. Free-form
/// stable string; typically the CHAT-file basename stem.
pub struct SessionId;
);
MergeOverride
A single per-session decision record. The unit of operator
adjudication. As shipped (speaker_id/override_file.rs):
/// A single override-file entry: the operator decision for one
/// session. See `merge-overrides.md` for field semantics.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct MergeOverride {
/// How the decision was made.
pub mode: OverrideMode,
/// Per-donor-speaker-code role assignment, for every speaker
/// whose `mapping` action is `Rename` (schema v2; see the
/// speaker-mapping section above).
pub adult_roles: BTreeMap<String, InsertedRoleSpec>,
/// Map from input speaker codes to actions. Every speaker that
/// exists in the input must appear here.
pub mapping: BTreeMap<String, SpeakerAction>,
/// Per-speaker Jaccard scores recorded at decision time.
/// Present for `Auto` (and `Explicit` decisions that followed a
/// low-confidence reference-mode attempt).
#[serde(skip_serializing_if = "BTreeMap::is_empty", default)]
pub scores: BTreeMap<String, f64>,
/// Winner-score / runner-up-score margin. Serialized as a
/// number; the divide-by-zero case is `f64::INFINITY`.
#[serde(skip_serializing_if = "Option::is_none", default)]
pub margin: Option<f64>,
/// Free-form identifier of the operator who made the decision.
pub operator: String,
/// When the decision was made: RFC 3339 in New York time, whole
/// seconds (`talkbank_transform::recorded_time::RecordedTime`).
pub decided_at: RecordedTime,
/// Free-text operator note. Strongly recommended for `Explicit`
/// and `Override` modes.
#[serde(skip_serializing_if = "Option::is_none", default)]
pub note: Option<String>,
/// Operator-supplied audit flags (e.g. `"diarization-mixed"`,
/// `"best-guess"`).
#[serde(skip_serializing_if = "Vec::is_empty", default)]
pub flags: Vec<String>,
/// Which engine produced this decision. Absent in pre-provenance
/// files, which deserialize as `Deterministic`.
#[serde(default)]
pub engine: DecisionEngine,
/// LLM audit trail; present only for `engine = Llm` decisions.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub judgment: Option<JudgmentProvenance>,
}
The struct embeds the timestamp via chrono::DateTime<Utc>; serde
serializes to RFC 3339 (2026-05-27T08:41:00Z) by default. TOML
preserves this format faithfully. The engine / judgment
provenance fields record whether a decision was deterministic or LLM-made
(see speaker_id/provenance.rs); they need no schema bump because they are
backward compatible in both directions, as documented in
merge-overrides.md.
OverrideFile
The top-level container. Holds schema version + per-session entries. Read from / written to disk as TOML.
/// Current schema version supported by this binary (module-level
/// const in `speaker_id/override_file.rs`). Readers refuse files
/// with any other value; there is no implicit version, no fallback,
/// no auto-migration. Bumped from 1 to 2 for the per-speaker
/// `adult_roles` map (was `inserted_role`, a single shared field).
pub const CURRENT_SCHEMA_VERSION: u32 = 2;
/// The full override-file document.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct OverrideFile {
/// Schema version. Currently 2. Always `CURRENT_SCHEMA_VERSION`
/// when this binary writes; readers reject other values with a
/// typed error rather than guessing.
pub schema_version: u32,
/// Per-session entries, alphabetically ordered by session ID
/// via the `BTreeMap` default.
#[serde(flatten)]
pub entries: BTreeMap<String, MergeOverride>,
}
impl OverrideFile {
/// Read an override file from disk, or return an empty default
/// (with the current schema version) if the path does not
/// exist. Refuses any `schema_version != CURRENT_SCHEMA_VERSION`.
/// Used by the `--write-override` append flow.
pub fn read_or_default(path: &Path) -> Result<Self, OverrideFileError>;
/// Serialize to TOML and write via `.tmp` + rename, so a crash
/// mid-write leaves the prior file intact rather than truncated.
pub fn write(&self, path: &Path) -> Result<(), OverrideFileError>;
/// Insert (or replace) the entry for `session_id`.
pub fn upsert(&mut self, session_id: String, entry: MergeOverride);
pub fn get(&self, session_id: &str) -> Option<&MergeOverride>;
}
(The designed standalone read never shipped; read_or_default is
the single read path, and the designed insert shipped as upsert.
Iteration helpers session_ids, auto_entries, and llm_entries
serve diagnostics, the post-merge sanity scan, and LLM
audits respectively.)
The #[serde(flatten)] on entries means the on-disk TOML is
flat tables keyed by session ID (as shown in the
speaker-id.md schema):
schema_version = 2
[S01-1]
mode = "auto"
adult_roles = { PAR0 = { code = "INV", tag = "Investigator" } }
# ...
rather than nested under an [entries] table.
Error types
Two thiserror-based enums covering the merge pipeline’s failure
modes. Each variant carries enough information for the CLI to
produce a useful diagnostic and for callers to pattern-match
behavior.
SpeakerIdError
As shipped (speaker_id/error.rs; several designed variant names
changed, and the low-confidence payload became the full
DonorMatchReport rather than loose fields):
#[derive(Debug, thiserror::Error)]
pub enum SpeakerIdError {
/// The `--mapping` spec couldn't be parsed.
#[error("invalid --mapping spec: {0}")]
InvalidMappingSpec(String),
/// Reference mode: no utterances for the requested anchor
/// speaker in the reference transcript.
#[error("reference transcript has no utterances for anchor speaker {anchor}")]
ReferenceMissingAnchor { anchor: SpeakerCode },
/// Reference mode: fewer than two distinct donor speakers, so
/// there is nothing for multiset-Jaccard to choose between.
DonorTooFewSpeakers { speakers: Vec<SpeakerCode> },
/// Reference mode: winner-to-runner-up margin below the
/// confidence threshold; the auto-decision is refused.
LowConfidence {
/// Full match report: would-be winner, per-speaker scores,
/// margin. `--write-pending` records it for adjudication.
report: DonorMatchReport,
threshold: ConfidenceThreshold,
},
/// Override-file replay: the requested session ID is not in the
/// override file; the available IDs are surfaced.
SessionIdNotFound { session_id: String, available: Vec<String> },
/// Override-file replay: a `Rename` action with no matching
/// `adult_roles` entry (hand-corrupted file); fails closed.
OverrideRenameMissingRole { speaker: SpeakerCode },
/// Underlying parse error from the input file.
#[error("parse error: {0}")]
Parse(#[from] PipelineError),
}
The LowConfidence variant is the only “soft” failure: the caller
(CLI) maps it to exit code 4 and prints the scores. Parse maps to
exit 1 (invalid input); every other variant maps to exit 2
(precondition violation) per the user-guide contract. The mapping is
the CLI layer’s job; SpeakerIdError itself just classifies the
failure mode. (The designed SpeakerNotInMapping /
MappingSpeakerNotInInput variants did not ship: apply_mapping
currently passes through speakers absent from the mapping unchanged,
and enforcing the every-input-speaker precondition is a documented
follow-up in speaker_id/apply.rs. The designed OverrideIo
wrapping also did not ship; override-file I/O failures surface as
OverrideFileError directly.)
MergeError
transcript_merge.rs owns the exhaustive error enum and its payloads.
Parsing errors belong to the caller: the merge accepts already parsed models.
The CLI explicitly maps every merge refusal to exit 2.
The merge refuses missing retained content or timelines, conflicting speakers or languages, malformed donor metadata order, unpositioned selected utterances, and reversed source starts. It also refuses ambiguous section placement, inconsistent participant joins, and invalid assembled output. The merge contract specifies the ordering policy.
Internal admission and public reporting form distinct transitions:
flowchart LR
A[Parsed source ASTs] --> B[SourceAdmission]
B -->|finish: attach section bounds| C[OrderedSource]
C --> D[AdmittedMerge]
D -->|assemble, join participants, validate| E[Merged: ValidChatFile]
E -->|report| F[Reported: ValidChatFile]
F -->|immutable borrow| G[Serialization]
F -->|into_file: relinquish validity| H[Editable ChatFile]
Only admission constructs source cursors. Assembly can consume their frontiers;
it cannot sort them or detach dependent tiers from their owning utterances.
Section comparisons may still refuse: admission supplies the source bounds,
not a claim that every cross-source order is determined. Only successful
assembly and full validation construct Merged. Edits after into_file need
fresh validation.
Two shipped rules worth calling out because they refine the designed
“exact @Languages match” and “concatenate @Participants”
contracts:
@Languagesis donor-subset matching, not exact equality. File 2 (the donor, typically ASR output) may declare a subset of File 1’s languages (an ASR run in a fixed language mode under-claims; that is expected). Only donor over-claiming, a donor language absent from File 1, raisesLanguageMismatch, since it may signal a wrong-file pairing or a language the annotator missed.@Participantsinsertion dedupes. A donor entry whose speaker code File 1 already declares is silently skipped (not inserted twice) when File 1’s declaration is vestigial: zero utterances under that code, and role/name metadata matching the donor’s. If File 1 has real utterances under the code, or the two declarations disagree, the merge refuses withParticipantAlreadyDeclaredinstead. The same dedupe set filters the inserted@IDrows.
OverrideFileError
Independent enum because override-file I/O is also called by
non-speaker-id code paths (the adjudication tool, future UIs). As
shipped (speaker_id/override_file.rs), it is leaner than the
designed five-variant version: read/write/parse failures collapse
into Io and Toml, and found is an Option<u32> so a missing
schema_version field is reported distinctly from a wrong one:
#[derive(Debug, thiserror::Error)]
pub enum OverrideFileError {
/// The file's `schema_version` is missing or not equal to
/// `CURRENT_SCHEMA_VERSION` (currently 2). The binary refuses to
/// interpret unknown versions rather than risk silent misreads.
#[error("unsupported override-file schema_version {found:?}; this binary supports {supported}")]
UnsupportedSchemaVersion {
/// The schema version as read from the file (None if the
/// field was absent entirely).
found: Option<u32>,
/// The schema version this binary supports.
supported: u32,
},
/// I/O error reading or writing the file.
#[error("override-file I/O error: {0}")]
Io(#[from] std::io::Error),
/// TOML parse / serialize error.
#[error("override-file TOML error: {0}")]
Toml(String),
}
(The designed NotFound variant is unnecessary: read_or_default
treats a missing file as the empty-file default, and any other I/O
failure surfaces through Io.)
Module layout
As shipped. The design’s talkbank-model/src/merge/ layout was
never created (see “Where the types live”); the real layout is:
crates/talkbank-transform/src/speaker_id/
mod.rs pub re-exports (the crate-facing surface)
types.rs JaccardScore, ConfidenceMargin, ConfidenceThreshold
mapping.rs MappingSpec, SpeakerAssignment, parse_mapping_spec
identify.rs identify_mapping, DonorMatchReport,
DEFAULT_CONFIDENCE_THRESHOLD
apply.rs apply_mapping, apply_mapping_chat
override_file.rs CURRENT_SCHEMA_VERSION, OverrideMode,
SpeakerAction, InsertedRoleSpec, MergeOverride,
OverrideFile, OverrideFileError
provenance.rs DecisionEngine, JudgmentProvenance, ModelId, ...
error.rs SpeakerIdError
judgment/ LLM holistic-judgment surface (sampling, prompt
rendering, provider, consume; home of the
adult_roles same-role auto-disambiguation)
crates/talkbank-transform/src/transcript_merge.rs
merge_chat_files, MergeError, DEFAULT_STRIP_TIERS,
Merged, Reported, MergeOrigin, ReferenceFate, DonorFate,
ReferenceIdx, DonorIdx
Each file aims for the ≤400-line target; concerns that outgrew a
single file (the LLM judgment surface) became the judgment/
subdirectory, exactly the split-further move this section
anticipated.
Type design rules followed
A spot-check against the cross-cutting design rules in this repo’s
root AGENTS.md, restated against the shipped code:
- Newtypes over primitives. Every numeric domain value
(
JaccardScore,FiniteConfidenceMargin,ConfidenceThreshold) is wrapped;ConfidenceMarginis a sum type; CHAT-domain strings reuse the existingSpeakerCode/ParticipantRole/ParticipantNamewrappers. (The designedSessionId/OperatorIdnewtypes shipped as plainStringat the on-disk serialization boundary; see the designed-vs-shipped table.) ✓ - No tuple-packed seams.
InsertedRoleSpecis a struct, not(code, tag);SpeakerAssignment::Renamecarries named fields;MergeOverridelikewise. ✓ - No boolean blindness.
SpeakerAction,OverrideMode, andConfidenceMarginare enums. No-information, finite, and unbounded confidence states cannot be conflated. ✓ - Typed errors. Three
thiserrorenums (SpeakerIdError,MergeError,OverrideFileError) with named-field variants carrying full context. ✓ - Deterministic seams.
BTreeMapfor every serialized collection (adult_roles,mapping,scores,entries). The in-memoryMappingSpecis aHashMap; it is never serialized directly. ✓ - Module browseability. One file per concern in
speaker_id/, with the LLM judgment surface split into its ownjudgment/subdirectory. ✓ Defaultimpls present where meaningful.DEFAULT_CONFIDENCE_THRESHOLD(2.0);OverrideFile::default()for the empty-file case. ✓Displayimpls present where user-visible.JaccardScore,ConfidenceMargin,ConfidenceThreshold. ✓- Parse functions at the CLI boundary, not regex hacks in
command code.
parse_mapping_specfor--mapping; theCODE:ROLEpair parse for--inserted-role. ✓
Decisions on the seven open questions
Resolved 2026-05-27, captured here so implementers don’t re-litigate.
1. JaccardScore representation: f64
Multiset Jaccard J(A, B) = sum_w min(A[w], B[w]) / sum_w max(A[w], B[w])
is computed from u64 token counts, which fit in f64’s 53-bit
mantissa for any plausible CHAT bag-of-words. The division is
inexact in general but IEEE 754 makes it bit-deterministic given
the same inputs across every platform that implements 754 (all of
ours: Windows, macOS, Linux, x86_64, arm64).
The bit-deterministic reproducibility property is load-bearing
because the override-file audit trail records scores; a researcher
re-running speaker-id years later on the same inputs must compute
the same score to verify the decision. f64 arithmetic provides
this for free given workspace platform constraints. Document the
property in the type’s rustdoc.
A rational u64/u64 representation was considered for “true”
reproducibility but adds boilerplate and a comparison-against-
threshold operation that loses the same precision in the end (the
threshold is a ratio too). Reject.
2. DateTime<Utc> crate: chrono
The workspace already pins chrono = "0.4" at the root
Cargo.toml. The merge code (in talkbank-transform) uses the
workspace version verbatim via chrono = { workspace = true }. No
new datetime dep.
The “succession-aware” rule from the workspace-root AGENTS.md
contributor guide (outside the book) and the analogous
feedback_no_terraform_only_opentofu discipline from operator
memory says: do not fragment the ecosystem by introducing a
second tool when a workspace tool already does the job. jiff is
a fine library but adopting it for one new module would mean two
datetime crates in tree.
Override-file timestamps serialize as RFC 3339 UTC; chrono’s serde
feature handles this with #[serde(with = "chrono::serde::ts_rfc3339")]
or the default Serialize/Deserialize impl.
3. TOML library: toml (the workspace-pinned crate)
Workspace already pins toml = "^1.1.2". That crate reads AND
writes, no need to combine toml and toml_edit for the v1
override-file format.
toml_edit was considered for its formatting/comment preservation
across in-place edits. The case for it is hypothetical right now:
override files are primarily machine-written by chatter speaker-id --write-override; human edits exist but are not the dominant
workflow. The cost of toml_edit is the second TOML dep (workspace
churn, plus the friction every contributor pays parsing TOML
through one API and writing through another).
If a workflow emerges where operators heavily hand-edit override
files and lose formatting on each batch re-run, swap to toml_edit
then. Defer.
4. MergeOverride::flags: Vec<MergeFlag>
Operator-supplied flags are semantically set-like (each flag
present or absent), but Vec is the right representation because:
MergeFlagincludes aCustom(String)#[serde(untagged)]variant. DerivingOrdon this enum requires a manualOrdimpl that hashes the discriminator + the inner string. Doable but adds maintenance load.- The order of flags in the on-disk file isn’t load-bearing for correctness; deterministic single-source-write produces a deterministic Vec.
- Duplicates are noise but not corrupting. Document in the field’s rustdoc that consumers should treat as set semantics (deduplicate before comparing).
The writer (speaker-id --write-override path) inserts flags in a
deterministic order; on-disk Vec is fully reproducible. If a
hand-edited file has an out-of-order or duplicated flag list, that
shows up as a non-corrupting noise in subsequent diffs, acceptable.
5. SpeakerMapping::assignments: BTreeMap<SpeakerCode, MappingAction>
Confirmed. BTreeMap gives:
- One-action-per-speaker by construction (no duplicate keys).
- Deterministic serialization order (alphabetical by
SpeakerCode). - Cheap membership tests during apply.
The AGENTS.md “no tuple-packed seams” rule targets raw tuples as
struct fields or function arguments. A BTreeMap’s internal
key-value pairing is not a domain seam exposed to the API; it’s
the representation. Approved.
(As shipped, this decision holds for the serialized shape:
MergeOverride.mapping is BTreeMap<String, SpeakerAction> and
MergeOverride.adult_roles is BTreeMap<String, InsertedRoleSpec>.
The in-memory MappingSpec is a HashMap because it is never
serialized directly.)
6. Schema versioning policy: strict refuse-with-clear-error
The reader (OverrideFile::read_or_default, as shipped) refuses any
schema_version != CURRENT_SCHEMA_VERSION with a typed
OverrideFileError::UnsupportedSchemaVersion { found, supported }.
No automatic migration.
This is the conservative default. Reasons:
- We have no upgrade history yet; building a migration framework
for a problem that doesn’t exist is premature abstraction
(
AGENTS.md“Always Fix Root Causes” + the general “no premature abstraction” instinct). - The override file is fundamentally a record of operator decisions. If the schema breaks, operators re-adjudicate; the prior file becomes a historical artifact that can be read by scripts with old binaries.
- When a real schema change lands and there is real upgrade
friction, that’s the moment to write a one-shot migration
(
chatter merge migrate-overrides --from <path> --to <path>). Until that happens, premature migration code is dead weight.
Document this in the reader’s rustdoc so the policy is explicit to
callers. The policy has since been exercised for real: the 2026-07
v1 -> v2 bump (the per-speaker adult_roles map) was a breaking,
non-migrating change exactly as designed here; v1 files are refused
and their sessions re-adjudicated. The version-to-version diff and
migration instructions live in
merge-overrides.md.
7. Where the --mapping parser lives: beside the mapping type
parse_mapping_spec("PAR0=drop,PAR1=INV:Investigator") -> Result<MappingSpec, SpeakerIdError>
lives alongside the MappingSpec type it returns, in
talkbank_transform::speaker_id::mapping as shipped (the design
said talkbank-model::merge::mapping; the parser moved with the
types when they landed in talkbank-transform, see “Where the types
live”).
Why:
- The spec format is part of the type’s contract. A reader looking
for “how do I construct a
MappingSpecfrom a string?” should find the answer where the type is defined, not in the consumer CLI crate. - A future non-CLI consumer (HTTP API, library wrapper, scripting
binding) wants the same parser without re-implementing or
depending on
chatter. talkbank-transformhas no CLI-framework dependency (noclap), but a free function returningResult<MappingSpec, _>doesn’t need one. Theclapvalue-parser inchatterbecomes a thin shim overparse_mapping_spec.
If at some point a SECOND mapping syntax becomes useful (e.g.,
JSON-inline, or a TOML fragment), add a parse_mapping_json
sibling rather than reshaping parse_mapping_spec. The existing
parser stays the lingua franca.
Each source speaker may occur only once in this string syntax. Repeated keys,
including identical repeats, return InvalidMappingSpec; a vacant map entry
is the parser’s only insertion capability, so no earlier decision is silently
replaced. This boundary does not make the public MappingSpec alias a validated
speaker/role type, nor does it prove target-code uniqueness or CHAT validity.
These decisions are the design baseline going into spec authoring and implementation. Future revisions to any of them require an explicit doc update plus a deprecation/migration plan, not a silent change in the implementation.
Relationship to specs and tests
The design intended a spec entry in spec/constructs/merge-types/
per type/invariant pair, regenerated into Rust tests via the
spec/tools generators. That directory was never created: as
shipped, the behavioral invariants are pinned directly by the Rust
test suites instead, per the layered scheme in the
Test Plan: transform-level tests
(crates/talkbank-transform/tests/speaker_id_tests.rs,
transcript_merge_tests.rs, adjudication_tests.rs), CLI
subprocess tests (crates/chatter/tests/merge_tests.rs,
speaker_id_tests.rs, adjudication_tests.rs), and per-module
#[cfg(test)] unit tests beside the types themselves (e.g. the
round-trip and per-speaker-role tests in
speaker_id/override_file.rs). Folding the fragment-level cases
(token cleaning, Jaccard goldens) into spec/constructs/ remains an
open option, not a shipped mechanism.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Merge Pipeline, Test Plan
chatter merge,chatter pipelineandchatter batchare not CLI commands. Where this page names them it describes the operation of the structural library intalkbank-transform, which is current. See the removal notice.
Status: Draft Last modified: 2026-10-02 (commit 2d7e886b)
This page is the test-coverage roadmap for the new merge pipeline
(chatter speaker-id + chatter merge + chatter adjudicate +
the override-file format + the underlying
talkbank-transform::speaker_id types). It exists because, per this
repo’s root AGENTS.md
red/green TDD rule, every new feature starts with failing tests
at the highest level the feature lives at, and we want to
enumerate those tests before writing the implementation, so
coverage is designed, not discovered.
The cycle plan below is design context, not proof of current coverage. The
implemented merge is an ordered AST merge, not a global sort. Current regression coverage in
transcript_merge_tests.rs checks complete reference line order, donor body
comments, direct model validity, refusal of missing or reversed timing,
determined versus ambiguous section placement, and malformed donor metadata
admission. See the current merge contract and
domain transitions for the implemented rules.
TDD discipline, what “strict red/green” means here
Every cycle of impl-phase work is:
- RED. Write ONE failing test at the highest layer the feature lives at. The test exercises a real user-observable behavior, not an internal helper. Commit the failing test alone (or stage it before any code change), verify it fails for the right reason (the missing behavior), not for a compile error or a typo.
- GREEN. Write the smallest code change that makes the test pass. No anticipating future tests, no scaffolding for tests that don’t yet exist. The codebase should compile and pass tests at this point.
- REFACTOR. With the green test as the safety net, tighten the implementation: extract helpers, rename for clarity, replace primitives with newtypes, document tricky parts. Tests stay green throughout.
- DRILL DOWN if needed. If the L3 (or L2) test passes but pinned the behavior less precisely than the contract requires (e.g., the L3 test asserts “exit 2 with some error” but the contract says “the specific MergeError variant must match”), add an L2 (or L1) test next that drills into the precise path. The drilled test FAILS at first against the green-but-imprecise impl, motivating the tighter impl.
Cycles must be atomic: one RED → one GREEN → optional REFACTOR → optional drill-down. Do not stack multiple tests on top of a single impl change; do not write impl ahead of tests. The discipline matters because the bug bar of this pipeline is high (CHAT-data byte-stable preservation, audit-trail reproducibility) and TDD is the cheapest way to catch regressions before they ship.
Three test layers + the adjudication layer
The merge pipeline’s behavior spans four substrates with different testing mechanisms.
| Layer | Substrate | Why tests live here |
|---|---|---|
| L1, Spec / fragment | spec/constructs/speaker-id/ → current spec/tools generators | Token-cleaner behavior on CHAT fragments (markup strip for Jaccard scoring). Same mechanism that pins parser/grammar tests; regenerated regression. |
| L2, Transform / AST | crates/talkbank-transform/tests/ | Pure-Rust tests over parsed ChatFile values. identify_mapping, apply_mapping, merge, run_adjudication semantics on hand-built or parsed CHAT inputs. No process boundary. |
| L3, CLI / subprocess | crates/chatter/tests/merge_tests.rs (new) | End-to-end behavior of chatter speaker-id, chatter merge, and chatter adjudicate invoked as subprocesses (assert_cmd + predicates). Exit codes, flag parsing, file I/O, stderr formats. |
| L4, Scripted adjudication | crates/talkbank-transform/tests/adjudication_tests.rs + scripted prompter | Operator-decision paths in chatter adjudicate. Uses ScriptedPrompter injecting synthetic operator choices. See Adjudication Workflow for the prompter abstraction. |
L1 ⊂ L2 ⊂ L3 in terms of failure-mode coverage: a failing L1 test implies a failing L2 test which implies a failing L3 test. So when the same invariant could be tested at multiple layers, the starter test is the highest layer and lower-layer tests are supplements that pin the precise internal path. L4 sits beside L2/L3, same crate/file conventions but a dedicated layer because the prompter-injection pattern is specific to adjudication.
L1, Spec / fragment tests
Lives in spec/constructs/speaker-id/. Three subdirectories:
token-cleaner/: what the Jaccard tokenizer strips and keepsjaccard-scoring/: fixed-input → fixed-score golden testsmapping-application/: header rewrite rules on real fragments
L1.1, Token cleaner
Each spec is a CHAT main-tier fragment + the expected token list
after cleaning. Behavior pinned: bracket markup stripped,
angle-bracket retracing unwrapped, terminator variants
discarded, &-... / &+... discarded, xxx/yyy/www
discarded, 0 discarded, @l / @n / @c suffix dropped,
_-compound split to spaces, punctuation stripped, lowercased,
≥2-char alpha filter, NAK bullets stripped.
| Spec | Input fragment | Expected tokens |
|---|---|---|
clean-plain-utterance | *CHI:\thello world . | ["hello", "world"] |
clean-strip-bracket-codes | *CHI:\thello [*] [/] world [//] . | ["hello", "world"] |
clean-unwrap-angle-retrace | *CHI:\t<two of the> [//] three of the presents . | ["two", "of", "the", "three", "of", "the", "presents"] |
clean-strip-fillers | *CHI:\t&-um &+pre something &-uh . | ["something"] |
clean-strip-zero-and-paralinguistic | *CHI:\t0 [=! nodding] . | [] |
clean-strip-unintelligible | *CHI:\txxx and yyy and www . | ["and", "and"] |
clean-strip-bullets | *CHI:\thello world . \x150_1234\x15 | ["hello", "world"] |
clean-special-form-suffix | *CHI:\tnaming l@l u@l l@l u@l . | ["naming"] |
clean-compound-underscore | *CHI:\tValentine's_Day and Fruit_Loops . | ["valentine", "day", "and", "fruit", "loops"] |
clean-terminator-variants | *CHI:\thello +//. world +... again +/. last ! | ["hello", "world", "again", "last"] |
clean-overlap-markers | *CHI:\t↫here↫ and there . | ["here", "and", "there"] |
clean-lowercase-filter | *CHI:\tHello World A I am . | ["hello", "world", "am"] |
Each spec file in spec/constructs/speaker-id/token-cleaner/ has
the standard # name, ## Input, ## Expected tokens, and
## Metadata sections per the spec authoring template at
spec/AGENTS.md in the workspace root (outside the book).
L1.2, Jaccard scoring
Fixed bag-of-tokens pairs with known multiset Jaccard. These
guard against off-by-one errors in the sum_w min / sum_w max
implementation and against any future “optimizations” that
silently change scoring.
| Spec | Bag A | Bag B | Expected J(A,B) |
|---|---|---|---|
jaccard-identical | {hello:2, world:1} | {hello:2, world:1} | 1.0 |
jaccard-disjoint | {hello:1} | {world:1} | 0.0 |
jaccard-empty-empty | {} | {} | 0.0 |
jaccard-empty-nonempty | {} | {x:1} | 0.0 |
jaccard-multiset-counts | {a:3, b:1} | {a:1, b:1} | 2/4 = 0.5 |
jaccard-partial-overlap | {a:1, b:1, c:1} | {b:1, c:1, d:1} | 2/4 = 0.5 |
L1.3, Mapping application on fragments
Header-rewrite micro-tests. Each spec gives an input
@Participants: or @ID: row and a small mapping; the expected
output row is the rewritten form.
| Spec | Input row | Mapping | Expected output row |
|---|---|---|---|
participants-rewrite-rename | @Participants:\tPAR0 Participant, PAR1 Participant | PAR0→INV:Investigator, PAR1→drop | @Participants:\tINV Investigator |
participants-preserve-name-token | @Participants:\tCHI Alex Target_Child, PAR0 Participant | PAR0→INV:Investigator | @Participants:\tCHI Alex Target_Child, INV Investigator |
id-rewrite-rename | @ID:\teng|corpus_name|PAR0|||||Participant||| | PAR0→INV:Investigator | @ID:\teng|corpus_name|INV|||||Investigator||| |
id-drop-removes-row | @ID:\teng|...|PAR1|||||Participant||| | PAR1→drop | (row removed) |
id-preserves-other-fields | @ID:\teng|2|CHI|6;01.|female|NF||Target_Child||| | (no-op for CHI) | identical to input |
L2, Transform / AST tests
Lives in crates/talkbank-transform/tests/. Three test files:
speaker_id_tests.rstranscript_merge_tests.rsoverride_file_tests.rs
Each tests behavior over parsed talkbank-model::ChatFile values,
using inline synthetic CHAT strings parsed via
talkbank_parser::parse_chat_file (no subprocess overhead).
L2.1, identify_mapping (reference mode)
| Test | Scenario | Assertion |
|---|---|---|
identify_mapping_clean_winner | Reference has CHI saying content X; donor has PAR0 saying X verbatim and PAR1 saying unrelated content | Returns SpeakerMapping { drop: {PAR0}, rename: {PAR1: INV} }, margin >> 2.0 |
identify_mapping_borderline_refuses | Reference and both donor speakers share substantial vocabulary (margin < 2.0) | Returns Err(SpeakerIdError::LowConfidence { scores, threshold, margin }) |
identify_mapping_anchor_missing | Reference has no utterances tagged with anchor speaker | Returns Err(SpeakerIdError::AnchorMissingInReference { anchor: CHI }) |
identify_mapping_single_speaker_donor | Donor has only one speaker | Returns Err(SpeakerIdError::InsufficientSpeakers { n: 1 }) |
identify_mapping_threshold_at_exact_value | Constructed donor where margin = 2.0 exactly with threshold 2.0 | Returns Ok(_) (≥ comparison, not strict >) |
identify_mapping_threshold_below_exact_value | Margin = 1.9999 with threshold 2.0 | Returns Err(SpeakerIdError::LowConfidence) |
identify_mapping_unbounded_margin | Donor PAR1 has Jaccard 0 against reference; PAR0 > 0 | Returns Ok(_) with margin = Margin::Unbounded |
identify_mapping_deterministic | Same inputs, repeated call | Identical SpeakerMapping byte-for-byte (BTreeMap ordering) |
L2.2, apply_mapping
| Test | Scenario | Assertion |
|---|---|---|
apply_mapping_renames_main_tier | Donor has *PAR0:\t... and *PAR1:\t...; mapping renames PAR0→INV, drops PAR1 | Output has *INV:\t... for original PAR0 utts; PAR1 utts absent |
apply_mapping_byte_stable_except_prefix | Donor has rich CHAT markup, %wor, %com on every utt | Every retained utt is byte-identical except the *CODE:\t prefix; dependent tiers preserved exactly |
apply_mapping_rewrites_participants | Donor @Participants: has PAR0+PAR1 entries | Output has only INV entry (after PAR1 drop) |
apply_mapping_rewrites_id | Donor @ID: rows for PAR0+PAR1 | PAR0 row rewritten to INV with role tag; PAR1 row removed |
apply_mapping_speaker_not_in_input | Mapping references PAR9 which isn’t in donor | Returns Err(SpeakerIdError::MappingSpeakerNotInInput { speaker: PAR9 }) |
apply_mapping_speaker_not_in_mapping | Donor has PAR0+PAR1+PAR2 but mapping only covers PAR0+PAR1 | Returns Err(SpeakerIdError::SpeakerNotInMapping { speaker: PAR2 }) |
apply_mapping_preserves_other_headers | Donor has @Languages, @Media, @Comment | All non-Participants/non-ID headers pass through verbatim |
apply_mapping_idempotent_on_rerun | Apply mapping, parse output, apply identity mapping | Output unchanged (byte-stable) |
L2.3, merge (core invariants)
These mirror the user-guide’s “What the merged output guarantees” section directly. Each invariant from that section maps to one or more L2 tests; the L3 tests then re-exercise the same invariant through the CLI.
| Test | Invariant from user-guide | Assertion |
|---|---|---|
merge_retained_speakers_byte_stable | “Retained speakers are byte-stable” | Every *CHI: block from File 1 (main tier + all dependent tiers, including %com) appears in the output byte-identical, in original order |
merge_strips_default_derived_tiers | “Inserted speakers’ downstream-generated tiers are stripped” | Output has no %wor, %mor, %gra, %pho on inserted-speaker utts; other dependent tiers preserved |
merge_strip_tiers_configurable | “configurable via --strip-tiers” | Custom strip_tiers=[com] removes %com instead of the defaults |
merge_strip_tiers_empty_preserves_all | empty strip set | Inserted utts retain %wor, %mor, %gra, %pho from File 2 verbatim |
merge_utterance_order_by_start_time | “Utterance order is timeline order” | Output utterances sorted by start_ms ascending |
merge_stable_tiebreak_file1_first | “first-file utterance comes first” | When File 1 and File 2 each have an utterance starting at exactly t, the File 1 one appears first in the output |
merge_bullets_pass_through | “Time bullets are pass-through” | Every bullet in the output is exactly the bullet from its source utterance, merge does not recompute, smooth, or refresh |
merge_bullet_lift_from_wor | “If main tier lacks bullet, lift from %wor” | Donor utt with no end-of-line bullet but a %wor row gets a derived \x15<first>_<last>\x15 appended; original %wor then stripped per the tier policy |
merge_no_overlap_markers_injected | “Overlap markup is NOT injected” | Even when inserted utt’s bullet overlaps a retained utt’s bullet by 500ms, no [>]/[<] tokens appear anywhere in the output that weren’t in the original retained file |
merge_preserves_existing_overlap_markers | retained file already has [>] somewhere | The original [>] is preserved byte-stable on the retained utt |
merge_header_languages_passthrough | Header reconciliation rule | Output @Languages matches File 1’s |
merge_header_media_file1_wins | Header reconciliation rule | File 1 says video, File 2 says audio → output says video (no warning emitted for modality only) |
merge_header_participants_concatenates | Header reconciliation rule | Output @Participants: is File 1’s entries + File 2’s non-retained entries, in that order, with dedupe-on-insert: a File 2 entry whose speaker code File 1 already declares is skipped rather than inserted twice (legal only when File 1’s declaration is vestigial: zero utterances, matching role/name metadata; otherwise the merge refuses with MergeError::ParticipantAlreadyDeclared) |
merge_header_id_concatenates | Header reconciliation rule | Output @ID: rows are File 1’s + File 2’s non-retained, original order within each file; the same dedupe-on-insert set also filters File 2’s @ID rows, so a deduped participant contributes no duplicate @ID row |
merge_header_comments_concatenate | Header reconciliation rule | Output @Comment rows are File 1’s + File 2’s, in original order (ASR provenance preserved) |
merge_preconditions_retain_missing | exit code 2 precondition | File 1 declares no CHI; merge with retain={CHI} returns Err(MergeError::RetainSpeakersMissing) |
merge_preconditions_no_timeline | exit code 2 precondition | File 1 has no utterances with bullets → Err(MergeError::NoTimelineInFile1) |
merge_preconditions_language_mismatch | exit code 2 precondition | File 1 @Languages: eng, File 2 @Languages: yue → Err(MergeError::LanguageMismatch) |
merge_preconditions_ambiguous_speaker | exit code 2 precondition | Both files have INV utterances and retain={CHI} (INV not in retain) → Err(MergeError::AmbiguousSpeaker { speaker: INV }) |
merge_warns_on_backward_bullet_drift | “small backward-time bullets … proceeds” | File with utt1: 100_200, utt2: 190_300, succeeds, emits a warning |
L2.4, Override file I/O
| Test | Scenario | Assertion |
|---|---|---|
override_file_round_trip | Construct OverrideFile with one entry, write, read back | Re-read value == original |
override_file_refuses_missing_schema_version | TOML with no schema_version | Err(OverrideFileError::UnsupportedSchemaVersion { found: None, supported: 2 }) (the shipped found is an Option<u32>, so absence reports as None, not a sentinel value) |
override_file_refuses_wrong_schema_version | schema_version = 99 (any value other than the current 2; a pre-bump schema_version = 1 file is refused the same way, per the v1-to-v2 migration note in merge-overrides.md) | Err(UnsupportedSchemaVersion { found: Some(99), supported: 2 }) |
override_file_rejects_unknown_field | Entry has an extraneous field extra = "x" | Err(OverrideFileError::Parse) |
override_file_rejects_malformed_mode | mode = "guess" | Err(Parse) (only auto/explicit/override accepted) |
override_file_atomic_write | Write to a path that already exists | Original file is replaced atomically; no <path>.tmp left behind |
override_file_deterministic_serialization | Same struct, write twice | Bytes on disk are byte-identical between writes |
override_file_omits_empty_optionals | Entry has empty scores, no margin, empty flags | TOML output does not contain those keys |
override_file_preserves_margin_unbounded | Entry has margin = Margin::Unbounded | TOML on disk has margin = "unbounded"; reads back as Unbounded |
override_file_preserves_margin_finite | Entry has margin = Margin::Finite(3.81) | TOML on disk has margin = 3.81; reads back equal |
override_file_read_or_default_missing | Path does not exist | Returns empty OverrideFile with current schema version |
override_file_get_returns_entry | File has one entry under SessionId X | get(X) returns Some; get(Y) returns None |
L2.5, Domain-type unit tests
Smaller per-type tests. Each in its module’s #[cfg(test)] mod tests section.
Note on type names: several of the designed types referenced below
shipped under different names or shapes (InsertedRole is
InsertedRoleSpec; Margin is the explicit
ConfidenceMargin::{NoInformation, Finite, Unbounded} sum type;
RetainSet and MergeFlag never shipped as newtypes). See the
designed-vs-shipped table in
Domain Types
before writing any still-pending test from this table against the
current code.
| Test | Type | Assertion |
|---|---|---|
lexical_match_retains_the_counts_that_produce_its_score | LexicalMatchEvidence / JaccardScore | Reference, donor, and intersection counts derive union and score; no public scalar score constructor can detach them |
recorded_report_preserves_support_and_margin_state | recorded match report | Stable JSON retains every count and the typed margin state |
confidence_threshold_default_is_2_0 | ConfidenceThreshold | Default::default().value() == 2.0 |
confidence_threshold_rejects_below_1 | ConfidenceThreshold | new(0.5) → Err |
margin_from_scores_zero_loser | Margin | Positive evidence-derived score versus a zero evidence-derived score produces Margin::Unbounded |
margin_from_scores_zero_zero | Margin | Two zero evidence-derived scores produce Margin::NoInformation |
margin_meets_threshold | Margin | Finite(3.81).meets(threshold=2.0) == true; Finite(1.5).meets(2.0) == false; Unbounded.meets(threshold) == true for any threshold |
retain_set_parse | RetainSet | "CHI".parse() == Ok({CHI}); "CHI,SI2".parse() == Ok({CHI, SI2}); "".parse() == Err; "CHI,,SI2".parse() == Err |
inserted_role_parse | InsertedRole | "INV:Investigator".parse() == Ok(_); "INV".parse() == Err; ":Investigator".parse() == Err |
mapping_spec_parse_simple | parse_mapping_spec | "PAR0=drop,PAR1=INV:Investigator" parses to a complete MappingSpec with Drop for PAR0 and a Rename carrying PAR1’s own code + role |
mapping_spec_parse_drop_only | parse_mapping_spec | "PAR0=drop" parses; a drop-only mapping is legal in isolation, since roles are per-speaker and a mapping with no Rename needs no role at all |
mapping_spec_parse_multiple_roles | parse_mapping_spec | "PAR0=INV:Investigator,PAR1=MOT:Mother" parses, with each speaker’s Rename carrying its own role. (The original plan named this mapping_spec_parse_conflicting_roles and expected an error because the designed v1 schema allowed only one shared inserted role; the shipped schema-v2 per-speaker adult_roles map makes multiple distinct roles a supported case, and two adults assigned the same role auto-disambiguate to numbered codes INV1/INV2 with First_/Second_ specific-role labels.) |
merge_flag_serde_known_variants | MergeFlag | DiarizationMixed serializes as "diarization-mixed" (kebab-case); deserializes the same |
merge_flag_serde_custom | MergeFlag | Unknown string deserializes as Custom("unknown-flag"); serializes verbatim |
L3, CLI / subprocess tests
Lives in crates/chatter/tests/merge_tests.rs (new file).
Uses the same assert_cmd + predicates + tempfile pattern
as the existing integration_tests.rs. Each test invokes
chatter speaker-id or chatter merge as a subprocess against
files written to a tempdir().
L3.1, chatter merge, success paths
| Test | Invariants exercised |
|---|---|
merge_basic_clinician_pattern | E2E happy path: small hand-coded child-only file + small ASR-labeled file → exit 0, output exists, retained CHI byte-stable, inserted INV present with derived tiers stripped. Single-invocation smoke test. |
merge_writes_to_stdout_by_default | No -o flag → output goes to stdout, exit 0 |
merge_writes_to_output_path | -o merged.cha → file created with correct content; nothing on stdout |
merge_retain_multi_speaker | --retain CHI,SI2 keeps both CHI and SI2 byte-stable; everything else from File 2 |
merge_strip_tiers_custom | --strip-tiers com,act removes %com and %act instead of default set |
merge_strip_tiers_empty | --strip-tiers '' preserves %wor from File 2 in output |
L3.2, chatter merge, error paths
| Test | Asserted exit code | Asserted stderr |
|---|---|---|
merge_missing_file1 | 1 | “No such file” or equivalent typed message |
merge_unparseable_file1 | 1 | parser diagnostic |
merge_missing_retain_flag | 2 (clap) | clap usage message |
merge_retain_empty_value | 2 | typed error from RetainSet::from_str |
merge_no_retain_speakers_in_file1 | 2 | RetainSpeakersMissing rendered |
merge_no_timeline_in_file1 | 2 | NoTimelineInFile1 rendered |
merge_language_mismatch | 2 | LanguageMismatch { file1: eng, file2: yue } rendered |
merge_ambiguous_speaker | 2 | AmbiguousSpeaker { speaker: ... } rendered with hint to use –retain |
L3.3, chatter speaker-id, reference mode
| Test | Scenario | Assertion |
|---|---|---|
speaker_id_reference_auto_clean_winner | Reference + donor where margin >> 2.0 | Exit 0; output has expected renamed/dropped speakers |
speaker_id_reference_writes_override | With --write-override path.toml | File created; entry has mode = "auto", scores, margin, decided_at, operator |
speaker_id_reference_appends_to_existing_override | --write-override path.toml where file already has another session | New session added; existing session preserved |
speaker_id_reference_low_confidence_exits_4 | Margin < threshold | Exit 4; stderr contains per-speaker scores |
speaker_id_reference_anchor_missing_exits_2 | Reference has no anchor speaker utterances | Exit 2; typed error in stderr |
speaker_id_reference_threshold_override | --confidence-threshold 1.5 on a margin-1.7 case | Exit 0 (would have refused at default 2.0) |
speaker_id_reference_anchor_required | --reference without --anchor | Exit 2 (clap or our own); usage error |
L3.4, chatter speaker-id, explicit-mapping mode
| Test | Scenario | Assertion |
|---|---|---|
speaker_id_explicit_basic | --mapping "PAR0=drop,PAR1=INV:Investigator" | Exit 0; output renames PAR1→INV, drops PAR0 |
speaker_id_explicit_mapping_speaker_not_in_input | --mapping references PAR9 not in input | Exit 2; typed error |
speaker_id_explicit_speaker_missing_from_mapping | Input has PAR0+PAR1+PAR2; mapping only covers PAR0+PAR1 | Exit 2; typed error naming PAR2 |
speaker_id_explicit_with_note_records_in_override | --mapping + --write-override + --note "verified by listening" | TOML entry has note = "verified by listening" and mode = "explicit" |
L3.5, chatter speaker-id, override-file mode
| Test | Scenario | Assertion |
|---|---|---|
speaker_id_override_file_replay | Override file has entry for session-X | Reading override + applying produces same output as the original auto/explicit run |
speaker_id_override_file_missing_entry | Override file has no entry for the requested session | Exit 2; OverrideEntryMissing in stderr |
speaker_id_override_file_missing_file | --override-file path.toml where file doesn’t exist | Exit 1; NotFound in stderr |
speaker_id_override_file_wrong_schema_version | File has schema_version = 99 | Exit 1; UnsupportedSchemaVersion in stderr |
speaker_id_override_file_mutually_exclusive_modes | --reference AND --mapping both set | Exit 2 (clap or our own); only one operation mode allowed |
L3.6, Pipeline composition
These exercise chatter speaker-id → chatter merge composed
end-to-end through the file system, simulating the orchestrator
workflow.
| Test | Scenario | Assertion |
|---|---|---|
pipeline_speaker_id_then_merge | Run speaker-id on anonymous ASR file; run merge on the result + hand-coded file | Final merged file passes all merge invariants (retained byte-stable, etc.) |
pipeline_replay_via_override_file | Run once with auto; capture override file; delete intermediates; replay via --override-file; merge again | Final merged file is byte-identical to the original run (audit-trail-reproducibility property) |
pipeline_low_confidence_then_explicit | Run speaker-id; gets exit 4; capture scores from stderr; run again with --mapping matching what the operator would decide; record via --write-override; merge | All steps succeed; override file has mode = "explicit" with prior scores recorded |
L4, Scripted adjudication tests
Lives in crates/talkbank-transform/tests/adjudication_tests.rs.
Uses the Prompter trait and ScriptedPrompter documented in
Adjudication Workflow §The prompter abstraction.
Each test constructs a pending-adjudications input, scripts the
operator’s decisions, runs run_adjudication, and asserts on
the resulting override file plus the residual pending file.
L4.1, Speaker-id adjudication paths
| Test | Scripted decision | Assertion |
|---|---|---|
adjudicate_speaker_id_accepts_suggested | AcceptSuggested { note: None } for one pending entry | Override file entry has mode = "explicit", mapping matches suggested, pending file emptied |
adjudicate_speaker_id_override_mapping | OverrideMapping { mapping: { PAR0=rename, PAR1=drop }, note: Some("verified by listening") } (opposite of suggested) | Override file mapping matches operator’s choice; note recorded |
adjudicate_speaker_id_defer | Defer { reason: "need to listen to audio" } | Pending entry untouched; override file unchanged; tool exits 4 (deferred) |
adjudicate_speaker_id_block | Block { reason: "reference file missing bullets" } | Pending entry tagged as blocked; override file unchanged |
adjudicate_speaker_id_kind_mismatch_rejected | OverrideInsertedRole { ... } against a speaker-id-low-confidence entry | Returns Err(AdjudicationError::DecisionKindMismatch); nothing written |
L4.2, Parent-role-lookup adjudication paths
| Test | Scripted decision | Assertion |
|---|---|---|
adjudicate_parent_role_accepts_default_inv | AcceptSuggested | Override entry uses INV:Investigator (the safe default) |
adjudicate_parent_role_overrides_to_mother | OverrideInsertedRole { code: "MOT", tag: "Mother" } | Override entry uses MOT; note recorded |
adjudicate_parent_role_overrides_to_father | OverrideInsertedRole { code: "FAT", tag: "Father" } | Override entry uses FAT |
adjudicate_parent_role_invalid_code_rejected | OverrideInsertedRole { code: "", tag: "Mother" } | Returns Err; with --skip-on-error, logs and proceeds |
L4.3, Diarization-mix and sanity-scan paths
| Test | Scripted decision | Assertion |
|---|---|---|
adjudicate_diarization_mix_flag_only | Flag { flags: [DiarizationMixed], note: "PAR0 mixes clinician+parent" } | Existing override entry gets flag added; mapping unchanged |
adjudicate_sanity_scan_swap_mapping | OverrideMapping { ... } reversing original speaker-id | Override entry updated; mode = "explicit"; original mapping preserved in history |
adjudicate_sanity_scan_confirms_real_overlap | Flag { flags: [Custom("real-overlap-confirmed")] } | Override entry gets custom flag; mapping unchanged |
L4.4, Workflow plumbing
| Test | Scenario | Assertion |
|---|---|---|
adjudicate_empty_pending_file_noop | Pending file has empty entries array | Exit 0; nothing changes |
adjudicate_resumption_skips_decided_entries | Pending file has 3 entries; first 2 already decided in override; only 3rd has no override entry | Prompter is called exactly once, for the 3rd entry |
adjudicate_re_adjudicate_preserves_history | Existing override entry; --re-adjudicate with new decision | New decision saved; prior decision preserved in history array |
adjudicate_kind_filter_processes_only_matching | Pending file has mixed kinds; --kind parent-role-lookup flag set | Prompter only called for parent-role-lookup entries; other kinds untouched |
adjudicate_dry_run_writes_nothing | Any pending input + any decision; --dry-run set | Override file unchanged; pending file unchanged |
adjudicate_scripted_mode_unknown_session_aborts | Scripted decisions reference session-X but pending has only session-Y | Returns Err(AdjudicationError::ScriptedDecisionWithoutPendingEntry); tool exits 2 |
adjudicate_scripted_mode_extra_pending_aborts | Pending has session-X and session-Y; scripted decisions cover only session-X | Returns Err(AdjudicationError::PendingEntryWithoutScriptedDecision); tool exits 2 |
adjudicate_mutually_exclusive_modes | --interactive + --scripted both set | Returns Err; tool exits 2 (clap or our own validator) |
L4.5, Prompter contract conformance
These tests pin the contract that any Prompter impl must
satisfy, so future UI backends (VS Code, web) can be developed
against the same invariants.
| Test | Scenario | Assertion |
|---|---|---|
prompter_terminal_round_trip_decision | TerminalPrompter reading a scripted stdin | Returns the expected OperatorDecision parsed from the operator’s typed input |
prompter_scripted_returns_decisions_in_order | ScriptedPrompter::from_decisions([d1, d2, d3]) | Three consecutive ask() calls return d1, d2, d3 in order |
prompter_scripted_panics_on_unscripted_session | ScriptedPrompter has decisions for session A; tool asks for session B | ask() returns Err(PrompterError::NoDecisionFor(SessionId)) |
prompter_scripted_toml_round_trips | Write a scripted-decisions TOML, read with ScriptedTomlPrompter, run | Same OperatorDecision sequence as a ScriptedPrompter::from_decisions with equivalent contents |
Fixture catalog
These are the synthetic CHAT pairs that the tests above consume. Each is small (≤20 utterances), exercises a precise invariant, and is fully fictional (no real corpus content).
The fixtures live as inline const FIX_*: &str blocks in the
respective test modules, following the precedent in
chatter/tests/integration_tests.rs (which has
const VALID_CHAT: &str = r#"..."# etc.).
FIX_REF_TWO_UTT_NO_MARKUP
The smallest possible valid CHAT pair input. Two *CHI:
utterances, no markup beyond a simple terminator, time bullets
on both. Used by cycle 1’s smoke test where the impl must
work without yet handling any markup edge cases.
FIX_ASR_LABELED_TWO_UTT
The matching donor for FIX_REF_TWO_UTT_NO_MARKUP: two
*INV: utterances at different time positions. Used by
cycle 1.
FIX_REF_CHILD_ONLY_SIMPLE
A 6-utterance child-only hand transcript with rich CHAT markup (error code, retracing, filled pause, special-form letter, zero realization with paralinguistic). Used by every L2/L3 merge test from cycle 2 onward as the canonical “File 1”, the reference / authoritative file. Has time bullets on every utterance.
FIX_ASR_ANON_2SPEAKER_SIMPLE
The matching ASR-output file with anonymous PAR0 (clinician,
asks questions) and PAR1 (child, says what FIX_REF_* shows
plus some extra). Has %wor on every utterance. Used by every
speaker-id test where auto-mode is expected to succeed cleanly
(margin >> 2.0).
FIX_ASR_LABELED_INV_SIMPLE
FIX_ASR_ANON_2SPEAKER_SIMPLE after speaker-id has run with
PAR1→drop, PAR0→INV:Investigator. Used by merge tests where
we want to skip the speaker-id step and test merge alone.
FIX_ASR_BORDERLINE_VOCABULARY
ASR file where both speakers describe the same picture-book content (margin 1.6-1.9 against reference). Used by low-confidence tests.
FIX_REF_NO_BULLETS
A reference file with no time bullets at all. Used to test
NoTimelineInFile1 precondition.
FIX_REF_LANG_ENG / FIX_ASR_LANG_YUE
Two files with conflicting @Languages. Used to test
LanguageMismatch.
FIX_AMBIGUOUS_INV
Two files both containing *INV: utterances, with
--retain CHI (INV not in retain set). Used to test
AmbiguousSpeaker.
FIX_REF_MULTI_RETAIN
Reference file containing *CHI: and *SI2: utterances (sibling
target). Used to test --retain CHI,SI2.
FIX_ASR_NO_MAIN_BULLET
Donor file where some utterances have no main-tier bullet, only
%wor. Used to test bullet-lift behavior in normalization.
FIX_OVERRIDE_VALID / FIX_OVERRIDE_WRONG_SCHEMA / FIX_OVERRIDE_MALFORMED
Override files in valid, schema-rejected, and parse-rejected shapes. Used by override-file I/O tests.
FIX_PENDING_SPEAKER_ID / FIX_PENDING_PARENT_ROLE / FIX_PENDING_MIXED_KINDS
Pending-adjudications files exercising one kind, another kind, and a mix. Used by L4 adjudication tests.
FIX_SCRIPTED_ACCEPT_ALL / FIX_SCRIPTED_OVERRIDE_FIRST_DEFER_SECOND
Scripted-decisions TOML files for ScriptedTomlPrompter.
Cover the canonical accept-suggested case and a mixed
override+defer case.
The exact bytes of each fixture are pinned in their respective test modules when the implementation lands; this plan doesn’t freeze them yet, only their purpose. Drafting the actual bytes is the first step of impl-phase work.
Coverage matrix
Cross-checking that every behavioral invariant from the four design docs has at least one test:
| Invariant source | Invariant | First-failing layer | Test name |
|---|---|---|---|
| merge user-guide | Retained byte-stable | L3 → L2 | merge_basic_clinician_pattern + merge_retained_speakers_byte_stable |
| merge user-guide | Derived tiers stripped | L3 → L2 | merge_strip_tiers_custom + merge_strips_default_derived_tiers |
| merge user-guide | Order by start_ms | L2 | merge_utterance_order_by_start_time |
| merge user-guide | Tiebreak File1 first | L2 | merge_stable_tiebreak_file1_first |
| merge user-guide | Bullets pass-through | L2 | merge_bullets_pass_through |
| merge user-guide | Bullet lift from %wor | L2 | merge_bullet_lift_from_wor |
| merge user-guide | Header reconciliation (all rows, including @Participants / @ID dedupe-on-insert of donor codes File 1 already vestigially declares) | L2 | merge_header_* series |
| merge user-guide + memory | No overlap markers injected | L2 | merge_no_overlap_markers_injected + merge_preserves_existing_overlap_markers |
| merge user-guide | Each precondition → exit 2 | L3 | merge_*_exits_2 series in L3.2 |
| merge user-guide | Warns on bullet drift | L2 | merge_warns_on_backward_bullet_drift |
| speaker-id user-guide | Reference mode auto | L3 | speaker_id_reference_auto_clean_winner |
| speaker-id user-guide | Explicit mode | L3 | speaker_id_explicit_basic |
| speaker-id user-guide | Override-file mode | L3 | speaker_id_override_file_replay |
| speaker-id user-guide | Confidence threshold (exit 4) | L3 → L2 | speaker_id_reference_low_confidence_exits_4 + identify_mapping_borderline_refuses |
| speaker-id user-guide | Byte-stable except prefix | L2 | apply_mapping_byte_stable_except_prefix |
| speaker-id user-guide | Header rewrites | L2 + L1 | apply_mapping_rewrites_* + participants-rewrite-* specs |
| speaker-id user-guide | Provenance captured | L3 | speaker_id_reference_writes_override |
| speaker-id user-guide | Each precondition → typed error | L3 → L2 | various *_exits_2 and apply_mapping_* tests |
| speaker-id user-guide | Token cleaner spec | L1 | clean-* specs |
| speaker-id user-guide | Multiset Jaccard formula | L1 | jaccard-* specs |
| override-file ref | Schema-version refusal | L2 | override_file_refuses_* tests |
| override-file ref | Round-trip fidelity | L2 | override_file_round_trip |
| override-file ref | Deterministic serialization | L2 | override_file_deterministic_serialization |
| override-file ref | Atomic write | L2 | override_file_atomic_write |
| override-file ref | margin "unbounded" form | L2 | override_file_preserves_margin_unbounded |
| domain types | JaccardScore range | L2 | jaccard_score_new_in_range |
| domain types | ConfidenceThreshold ≥ 1 | L2 | confidence_threshold_* |
| domain types | Margin semantics | L2 | margin_* |
| domain types | RetainSet::from_str | L2 | retain_set_parse |
| domain types | InsertedRole::from_str | L2 | inserted_role_parse |
| domain types | parse_mapping_spec | L2 | mapping_spec_parse_* |
| domain types | MergeFlag serde | L2 | merge_flag_serde_* |
| domain types | Pipeline reproducibility | L3 | pipeline_replay_via_override_file |
Every invariant has at least one named test; many have multiple across layers. When the impl phase begins, the first commit should produce the fixtures, the second commit the highest-layer failing test for the simplest invariant, then drill down per the standard TDD progression.
What this plan does NOT cover
- Performance / scaling tests. Until the pipeline shows up on a measured workload, no targeted perf assertions. The reference corpus’s existing round-trip benchmarks remain the baseline.
- Fuzz testing. This repository has a local
fuzz/workspace for parser/validation fuzzing. If the merge crate stabilizes enough to justify dedicated fuzzing, adding a merge-specific target for random parseable CHAT-pair inputs is a follow-up, not a v1 blocker. - Cross-platform CI checks. Windows / Linux / macOS each build the workspace; the merge module rides the existing CI. No platform-specific tests needed (the merge operates on parsed AST and writes UTF-8; no path-or-line-ending quirks).
- Real-corpus regression sweeps. Once impl lands, running
chatter mergeover a curated subset of the reference corpus and snapshotting outputs is a smart follow-up. Lives in a separatetests/golden/style mechanism if added; not designed here.
TDD authoring sequence
Each numbered item is one full RED → GREEN → REFACTOR cycle. Cycles must run in order; do not start cycle N+1 until cycle N is green and committed. Numbers are designed so the first working pipeline (cycle 8) emerges from the absolute minimum set of types + algorithms, then each later cycle extends.
The starter test for cycle 1 is intentionally tiny: a 2-utterance fixture pair with no markup, one retain speaker. The smoke test exercises every layer (parser, transform, CLI) but with the simplest possible CHAT bytes, so the first impl is small enough to land in one cycle.
Phase A, minimal end-to-end pipeline (cycles 1-8)
These cycles produce the simplest possible chatter merge
working end-to-end with synthetic fixtures.
| # | RED (failing test) | GREEN (smallest impl that passes) |
|---|---|---|
| 1 | merge_basic_smoke, L3 subprocess test against the tiniest fixture pair (FIX_REF_TWO_UTT_NO_MARKUP + FIX_ASR_LABELED_TWO_UTT), retain={CHI}, asserts exit 0 and “merged file exists” | Stub chatter merge subcommand wiring; introduce minimal talkbank-transform::transcript_merge::merge that interleaves utterances by start_ms and emits parser→serializer round-trip. No tier-stripping, no header-reconcile, no validation. Just: parse, sort, serialize. |
| 2 | merge_retained_speakers_byte_stable, L2 over the smoke fixture, asserts every CHI block byte-identical | Implement byte-stable handling for retained utterances (preserve main_raw_lines + dependent tiers exactly). |
| 3 | merge_strips_default_derived_tiers, L2 against a fixture where the donor has %wor rows | Implement tier_strip per the per-tier policy; drop %wor/%mor/%gra/%pho from inserted-speaker utts. |
| 4 | merge_utterance_order_by_start_time, L2 with a fixture where File 1 and File 2 utterances interleave | Implement timeline sort key (start_ms primary; source-order tiebreak). |
| 5 | merge_header_participants_concatenates, L2 | Implement header_reconcile::participants_merge. |
| 6 | merge_header_id_concatenates, L2 | Extend header_reconcile for @ID rows. |
| 7 | merge_header_languages_passthrough + merge_header_media_file1_wins + merge_header_comments_concatenate, L2 | Extend header_reconcile for remaining headers per the contract table. |
| 8 | merge_preconditions_retain_missing + merge_preconditions_no_timeline + merge_preconditions_language_mismatch + merge_preconditions_ambiguous_speaker, L3, each asserting exit code 2 with a specific stderr message | Implement preconditions module + map MergeError to exit codes in the CLI. |
Phase A, actual cycle log
The four-precondition cycle 8 was deliberately split into four
single-variant cycles (9a / 9b / 9c / 9d) so each MergeError
variant lands with its own RED→GREEN cycle and L2 + L3 sibling
tests. The numbering here is therefore finer-grained than the
plan table above; the table records the shape of Phase A, the
log records what was actually committed.
| # | Test(s) | Layer | Status |
|---|---|---|---|
| 1 | merge_basic_smoke | L3 | done |
| 2 | merge_retained_speakers_byte_stable | L2 | done |
| 3 | merge_strips_default_derived_tiers | L2 | done |
| 4 | merge_strip_tiers_configurable | L2 | done |
| 5 | merge_strip_tiers_empty_preserves_all | L2 | done |
| 6 | merge_header_participants_concatenates | L2 | done |
| 7 | merge_header_id_concatenates | L2 | done |
| 8a | merge_header_comments_concatenate | L2 | done |
| 8b | merge_header_languages_passthrough + merge_header_media_file1_wins | L2 | done |
| 9a | merge_no_retain_speakers_in_file1 + _returns_err | L3 + L2 | done (L2 sibling backfilled in 9c) |
| 9b | merge_no_timeline_in_file1 + _returns_err | L3 + L2 | done |
| 9c | merge_language_mismatch + _returns_err | L3 + L2 | done |
| 9d | merge_ambiguous_speaker + _returns_err | L3 + L2 | done |
End of Phase A: chatter merge works on simple fixtures with
all four preconditions (retain / timeline / language / ambiguous
speaker) enforced. The pipeline is publishable as v0.
Phase B, actual cycle log
Phase B picks up at cycle 10 in the cycle log (Phase A used 9a-9d for the precondition split).
| # | Test(s) | Layer | Status |
|---|---|---|---|
| 10 | speaker_id_explicit_basic | L3 | done |
| 11 | apply_mapping_byte_stable_except_prefix + apply_mapping_rewrites_participants + apply_mapping_rewrites_id | L2 | done (regression-guards) |
| 12 | identify_mapping_clean_winner | L2 | done |
| 13 | identify_mapping_borderline_refuses | L2 | done |
| 14 | speaker_id_reference_low_confidence_exits_4 | L3 | done |
| 15 | speaker_id_reference_writes_override (+ OverrideFile data model) | L3 | done |
| 16 | speaker_id_override_file_replay (+ OverrideFile::get) | L3 | done |
| 17 | adjudicate_speaker_id_accepts_suggested (+ adjudication core) | L4 | done |
| 18 | adjudicate_scripted_accepts_suggested (+ chatter adjudicate CLI + scripted-TOML I/O) | L3 | done |
| 19 | speaker_id_reference_writes_pending_on_low_confidence (+ --write-pending flag + LowConfidence carries DonorMatchReport) | L3 | done |
| 20 | adjudicate_speaker_id_override_mapping (+ OperatorDecision::OverrideMapping variant + scripted-TOML override-mapping shape) | L4 | done |
| 21 | adjudicate_interactive_accepts_suggested (+ TerminalPrompter + --interactive flag) | L3 | done |
| 22 | adjudicate_parent_role_lookup_chooses_role (+ PendingKindData promotion + ParentRoleLookup kind + ChooseRole decision) | L4 | done |
| 23 | adjudicate_interactive_chooses_role (+ parse_operator_response + kind-aware prompt hint) | L3 | done |
| 24 | adjudicate_interactive_override_mapping (+ parse_override_mapping + parse_speaker_assignment) | L3 | done |
| 25 | pipeline_clean_winner_end_to_end (+ chatter pipeline subcommand) | L3 | done |
| 26 | batch_pass1_single_session (+ chatter batch subcommand, subprocess driver) | L3 | done |
| 27 | batch_mixed_outcomes (regression-guard: clean+borderline aggregation) | L3 | done |
| 28 | batch_pass2_replay (+ --override-file on pipeline + batch; per-session auto-detection) | L3 | done |
| 29 | batch_skip_existing (+ --skip-existing flag on batch for idempotent re-runs) | L3 | done |
| 30 | refactor, PipelineArgs + BatchArgs structs retire three #[allow(clippy::too_many_arguments)] markers | , | done (true-no-op refactor; covered by cycles 25-29 regression suite) |
| 31 | refactor, split commands/speaker_id.rs (472 lines) into speaker_id/{mod,modes,writes,support}.rs (158 + 196 + 103 + 86 lines); retire 4 stale #[allow(dead_code)] markers on ReferenceModeOutcome (fields are read by write_override_entry) | , | done (true-no-op refactor; covered by cycles 10-29 regression suite) |
| 32 | adjudicate_sanity_scan_accept_suggested (+ AdjudicationKind::SanityScanMisclassification variant, PendingKindData::SanityScanMisclassification { suggested, reason } variant, two apply-decision arms mirroring SpeakerIdLowConfidence, terminal prompter render + prompt-hint arm) | L4 | done, adjudication kind end-to-end; the post-merge scan detector itself (heuristic + auto-pending-write) is a separate cycle 33 |
| 33 | sanity_scan_flags_inverted_mlu (+ talkbank_transform::sanity_scan::scan_session + chatter sanity-scan subcommand; mean-utterance-word-count asymmetry heuristic, default 1.5×, binary-mapping only) | L3 | done, detector + CLI end-to-end; multi-rename support, batch integration, and alternative heuristics deferred |
| 34 | batch_writes_override_for_auto_decisions (+ --write-override on both chatter pipeline and chatter batch; threaded through PipelineArgs.write_override_path + BatchArgs.write_override_path; reference-mode auto-decisions audit-trailed for sanity-scan + future re-runs) | L3 | done |
| 35 | batch_with_sanity_scan_flag_flags_inverted_mlu (+ --sanity-scan + --sanity-scan-threshold on chatter batch; post-loop subprocess driver for chatter sanity-scan; precondition validation requiring --write-override + --write-pending) | L3 | done |
| 36 | refactor, split cli/args/core.rs (984 → 747 lines): extract DebugCommands → debug_commands.rs, CacheCommands → cache_commands.rs, config enums (LogFormat, TuiMode, OutputFormat, ParserBackend, AlignmentTier) → cli_types.rs, unit-test module → core_tests.rs (via #[path]); satisfies the 800-line hard limit | , | done (true-no-op refactor; covered by full regression suite + 110 bin/integration tests) |
| 37+ | sanity-scan multi-rename support; diarization-mix-review kind (operator workflow design needed); newtype threading at struct seams (deferred simplify finding); apply_decision arm dedup + per-kind OperatorDecision sub-enums | L3 + L4 | pending |
Phase B, speaker-id pipeline (cycles 9-16)
These cycles add chatter speaker-id and its three modes.
| # | RED | GREEN |
|---|---|---|
| 9 | speaker_id_explicit_basic, L3 against an anonymous-2-speaker donor with --mapping "PAR0=drop,PAR1=INV:Investigator", asserts output has only INV utts | Stub chatter speaker-id subcommand. Implement parse_mapping_spec + apply_mapping. Reference mode and override-file mode return unimplemented!() for now. |
| 10 | apply_mapping_byte_stable_except_prefix + apply_mapping_rewrites_participants + apply_mapping_rewrites_id, L2 | Tighten apply_mapping per header rewrite rules. |
| 11 | identify_mapping_clean_winner, L2 with a fixture where one donor speaker overwhelmingly matches the reference | Implement text_cleaner + jaccard modules. Implement identify_mapping using them. Reference mode in CLI now works. |
| 12 | identify_mapping_borderline_refuses, L2 with a borderline fixture | Add ConfidenceThreshold check + LowConfidence error path. |
| 13 | speaker_id_reference_low_confidence_exits_4, L3 against borderline fixture | Map LowConfidence to exit code 4 in the CLI; print scores to stderr. |
| 14 | speaker_id_reference_writes_override, L3 with --write-override | Implement OverrideFile::read_or_default + OverrideFile::write. |
| 15 | speaker_id_override_file_replay, L3 with --override-file + --session-id | Implement override-file mode in CLI (OverrideFile::get + apply). |
| 16 | Token-cleaner L1 specs (a handful of representative clean-* specs from L1.1) + current spec/tools generators | Move the regex-and-string cleaner into a spec-test-covered implementation. Specs become the regression net. |
End of Phase B: full chatter speaker-id + chatter merge
pipeline works auto + explicit + override modes.
Phase C, adjudication (cycles 17-22)
These cycles add the chatter adjudicate tool and its
prompter-injection testability.
| # | RED | GREEN |
|---|---|---|
| 17 | adjudicate_empty_pending_file_noop, L4 against an empty pending file, asserts exit 0 + no changes | Stub chatter adjudicate subcommand. Implement PendingAdjudications::read + run_adjudication core skeleton with a no-op Prompter trait. |
| 18 | prompter_scripted_returns_decisions_in_order, L4 | Implement ScriptedPrompter::from_decisions (in-memory) per the Prompter trait. |
| 19 | adjudicate_speaker_id_accepts_suggested, L4 against FIX_PENDING_SPEAKER_ID with one AcceptSuggested decision | Implement apply_decision for the speaker-id-low-confidence kind. Override file now gets the decision; pending entry removed. |
| 20 | adjudicate_speaker_id_override_mapping, L4 with OverrideMapping decision | Extend apply_decision for the override-mapping variant. |
| 21 | adjudicate_speaker_id_kind_mismatch_rejected, L4 with a OverrideInsertedRole against a speaker-id pending entry | Implement kind→variants validation in apply_decision. |
| 22 | adjudicate_scripted_mode_unknown_session_aborts + adjudicate_scripted_mode_extra_pending_aborts, L4 | Tighten scripted-mode validation; assert 1:1 mapping between pending entries and scripted decisions. |
End of Phase C: scripted adjudication tested end-to-end with synthetic operator inputs. Interactive terminal UX still unimplemented (next phase).
Phase D, interactive UX (cycles 23-25)
| # | RED | GREEN |
|---|---|---|
| 23 | prompter_terminal_round_trip_decision, L4 with mocked stdin/stdout | Implement TerminalPrompter parsing [a]/[o]/[f]/... keys + optional follow-up prompts. |
| 24 | adjudicate_resumption_skips_decided_entries, L4 with a partially-decided override file + full pending list | Implement skip-already-decided logic in run_adjudication. |
| 25 | Manual smoke test (NOT automated), run chatter adjudicate --interactive against the test fixtures; visually confirm the operator UX matches the doc’s mock-up | Polish terminal output: ANSI formatting, fixed-width alignment, the [m] Show more context action, the [p] Play media action. |
End of Phase D: full v1 pipeline complete.
Phase E, non-speaker-id adjudication kinds (cycles 26-29)
Each adjudication kind gets its own RED→GREEN cycle.
| # | RED | GREEN |
|---|---|---|
| 26 | adjudicate_parent_role_overrides_to_mother + adjudicate_parent_role_overrides_to_father, L4 | Implement parent-role-lookup kind end-to-end (pending schema, prompter context, decision application). |
| 27 | adjudicate_diarization_mix_flag_only, L4 | Implement diarization-mix-review kind end-to-end. |
| 28 | adjudicate_sanity_scan_swap_mapping, L4 | Implement sanity-scan-misclassification kind end-to-end. |
| 29 | adjudicate_re_adjudicate_preserves_history, L4 | Implement --re-adjudicate flag; add history field to MergeOverride. |
Phase F, breadth pass (cycles 30+)
Fill in every remaining test from L1-L4 that hasn’t been written yet. These are coverage-deepening tests, not behavior adders. The impl from Phases A-E should pass them with at most minor refactoring; if a test fails meaningfully, that’s a gap in the impl that this cycle closes.
The breadth pass is the only phase where multiple cycles can proceed in parallel (different contributors take different test groups). Phases A-E are strictly serial.
Hard rules during impl phase
- No test stubs. Every test in this plan, when written,
must FAIL before its impl exists and PASS after. Skipped or
#[ignore]-marked tests are not allowed in the regression net (use#[ignore]only for genuinely slow or environment-dependent tests, not for “not implemented yet”). - No test deletion to make CI green. If a test that was passing starts failing after a refactor, the refactor is wrong. Investigate; do not delete the test.
- Three cycle archetypes, distinguish them. A cycle is one
of:
- bug-fix: RED motivates new impl code (cycle N-1’s impl truly cannot satisfy the new test).
- regression-guard: RED pins an invariant the impl
inherits from upstream infrastructure (e.g. parse→serialize
byte-stability inherited from
talkbank-parser). The test passes against cycle N-1’s impl, but the cycle is valuable because it locks in the invariant against future “optimizations” that might break it. Verbose-output the actual behavior on first run to confirm the invariant holds for the right reasons, not by accident. - true no-op: RED tests something already pinned elsewhere. These ARE unnecessary; drop the cycle or sharpen the test. The difference between regression-guard and true no-op is whether the invariant is named explicitly anywhere else. If yes (e.g., the parser crate already has a roundtrip test that covers it), the cycle is true-no-op. If no, the cycle is a regression-guard and worth keeping.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Merge Pipeline, Crate Architecture
chatter merge,chatter pipelineandchatter batchare not CLI commands. Where this page names them it describes the operation of the structural library intalkbank-transform, which is current. See the removal notice.
Status: Draft Last modified: 2026-10-02 (commit 2d7e886b)
This page explains where the new merge-pipeline code lives in the
chatter workspace, which crates gain modules, what
depends on what, and which boundary each piece sits inside. The
goal is succession-readability: a contributor coming to this
work for the first time should be able to map a behavior they
read about in
chatter merge or
chatter speaker-id to
the precise crate + module that implements it.
Companion documents:
- Domain Types: the typed vocabulary in
talkbank-transform::speaker_id(andMergeErrorbeside the merge algorithm intalkbank-transform::transcript_merge). - Test Plan: what tests live where.
- Override File Format, the on-disk format.
Boundary decisions
Two boundary decisions govern where every new piece of code lives.
Both reference rules already documented in this repo’s root AGENTS.md
(workspace-root contributor guide, outside the book).
Decision 1: talkbank-* crates, not batchalign-* crates
The merge pipeline is pure CHAT-AST structural manipulation, no ML, no audio I/O, no network, no model loading, no fleet runtime. Per the crate-boundary decision test in the workspace AGENTS.md:
If code fundamentally needs ML models, audio processing, network services, or fleet runtime →
batchalign-*crate. Otherwise →talkbank-*crate.
chatter merge and chatter speaker-id answer “no” to each
ML/audio/network/runtime question. They consume parsed
ChatFile values, manipulate them, and emit parsed-and-serialized
output. Even the speaker-id text-similarity scoring is a
deterministic function over CHAT content tokens, no ML model,
no embedding, no inference. All new merge code lives in
talkbank-* crates.
The batchalign-* crates remain the home for batchalign3 transcribe (ASR), batchalign3 align (forced alignment), and
batchalign3 morphotag (Stanza-based morphological tagging),
the ML-bearing stages that surround the merge in the pipeline.
Decision 2: types and algorithms in talkbank-transform, CHAT vocabulary in talkbank-model, CLI in chatter
The merge pipeline’s code splits across the same talkbank-* crates that already host the parse/validate/normalize/JSON pipelines:
talkbank-modelowns the CHAT-domain vocabulary the merge code references (SpeakerCode,ParticipantRole,ParticipantEntry,IDHeader,ChatFile). It gained no new merge module.talkbank-transformowns both the merge-specific domain types (MappingSpec,InsertedRoleSpec,SpeakerAction,OverrideMode,MergeOverride,OverrideFile, the error enums) and the algorithms (token cleaning, Jaccard scoring, mapping application, structural merge, adjudication core). No CLI parsing, no clap.chatterowns the subcommands (chatter speaker-id,chatter merge,chatter adjudicate, plus the composingpipeline/batch/sanity-scandrivers). Thin shim layer that parses arguments and drives the transform layer.
Design history. The original design gave the domain types their
own talkbank-model::merge module (“types in the model crate,
algorithms in the transform crate”). As shipped, the types live with
the algorithms in talkbank-transform::speaker_id instead; the
talkbank-model::merge module was never created. See
Domain Types §Where the types live.
This mirrors how chatter validate, chatter normalize,
chatter to-json are wired today and keeps the crate boundaries
honest: a caller wanting the algorithms and types without CLI
machinery (e.g., a library binding, an HTTP service, an external
tool reading override files) depends on talkbank-transform
without pulling in clap.
Crate dependency graph
The new code does not introduce any new crate-level dependencies, every edge below already exists in the workspace today. The merge work adds modules to existing crates.
flowchart TD
derive["talkbank-derive\n(proc macros, unchanged)"]
model["talkbank-model\n(CHAT vocabulary, unchanged)"]
parser["talkbank-parser\n(unchanged)"]
transform["talkbank-transform\n(+ speaker_id, transcript_merge,\nadjudication, sanity_scan modules)"]
cli["chatter\n(+ speaker-id, merge, adjudicate,\npipeline, batch, sanity-scan subcommands)"]
cli_tests["chatter/tests/\n(+ merge_tests, speaker_id_tests,\nadjudication_tests, pipeline_tests, batch_tests)"]
transform_tests["talkbank-transform/tests/\n(+ transcript_merge_tests, speaker_id_tests,\nadjudication_tests)"]
derive --> model
model --> parser
model --> transform
parser --> transform
transform --> cli
model --> cli
transform --> transform_tests
transform --> cli_tests
cli --> cli_tests
Module layout per affected crate
talkbank-model: unchanged
talkbank-model gained no merge module. (The original design added
a crates/talkbank-model/src/merge/ module with scoring / role
/ mapping / retain / override_file / errors files and
pub use role::{InsertedRole, MappingAction}-style re-exports; none
of that was created. The domain types shipped inside
talkbank-transform::speaker_id instead, with revised names; see
Domain Types.) The merge code consumes
talkbank-model’s existing CHAT vocabulary (SpeakerCode,
ParticipantRole, ParticipantEntry, IDHeader, ChatFile)
unmodified.
talkbank-transform, speaker_id/ module + transcript_merge.rs
Sibling top-level modules, mirroring the user-facing distinction
between the two subcommands. speaker_id/ holds both the domain
types and the algorithms; transcript_merge fits in a single file:
crates/talkbank-transform/src/speaker_id/
mod.rs pub re-exports (the crate-facing surface)
types.rs JaccardScore, ConfidenceMargin, ConfidenceThreshold
mapping.rs MappingSpec, SpeakerAssignment, parse_mapping_spec
identify.rs identify_mapping (token cleaning + multiset
Jaccard), DonorMatchReport,
DEFAULT_CONFIDENCE_THRESHOLD
apply.rs apply_mapping, apply_mapping_chat
(@Participants / @ID rewriting per mapping)
override_file.rs CURRENT_SCHEMA_VERSION, OverrideMode,
SpeakerAction, InsertedRoleSpec, MergeOverride,
OverrideFile, OverrideFileError
provenance.rs DecisionEngine, JudgmentProvenance, ...
error.rs SpeakerIdError
judgment/ LLM holistic-judgment surface (sampling,
prompt rendering, provider, consume)
crates/talkbank-transform/src/transcript_merge.rs
merge_chat_files (preconditions, header reconciliation, timeline
interleave, tier strip) -> Merged, MergeError, DEFAULT_STRIP_TIERS
crates/talkbank-transform/src/adjudication.rs
run_adjudication core, Prompter trait, ScriptedPrompter,
PendingAdjudications
crates/talkbank-transform/src/sanity_scan.rs
post-merge misclassification heuristic (scan_session)
All of these land alongside the existing CHAT-core transform modules
(parse, serialize, validate, normalize) in talkbank-transform.
Exposed via crates/talkbank-transform/src/lib.rs:
pub mod adjudication;
pub mod sanity_scan;
pub mod speaker_id;
pub mod transcript_merge;
chatter, new command modules
The CLI dispatch pattern in this crate uses one directory per
multi-file command (e.g. commands/validate/) or one file for
single-file commands (commands/normalize.rs, commands/clean.rs).
Speaker-id warranted a directory (it has reference / explicit /
override-file operation modes plus override/pending write paths);
merge and the other pipeline commands fit in single files:
crates/chatter/src/commands/speaker_id/
mod.rs SpeakerIdArgs + run_speaker_id entry point
modes.rs reference / explicit / override-file / holistic-LLM
mode drivers
writes.rs --write-override / --write-pending output paths
support.rs shared helpers (CODE:ROLE parsing, session-ID
derivation, typed error-to-exit-code mapping)
crates/chatter/src/commands/transcript_merge.rs
run_merge: parses both inputs, drives merge_chat_files, reports
MergeNotice values, maps MergeError to exit codes
MergeNotice / report_merge_notices: the operator-facing warnings,
shared with commands::pipeline so both merge paths say the same thing
crates/chatter/src/commands/adjudicate.rs chatter adjudicate
crates/chatter/src/commands/pipeline.rs chatter pipeline (speaker-id
then merge, one session)
crates/chatter/src/commands/batch.rs chatter batch (many sessions)
crates/chatter/src/commands/sanity_scan.rs chatter sanity-scan
crates/chatter/src/commands/merge_preflight.rs merge preflight checks
The CLI argument surface extends the top-level Commands enum in
crates/chatter/src/cli/args/core.rs, which carries Merge,
SpeakerId, Adjudicate, Pipeline, Batch, and SanityScan
variants with inline field definitions (not separate *Args
structs in the command modules). Subcommand dispatch in
crates/chatter/src/commands/dispatch.rs matches on the enum and
wires each arm to the respective commands::*::run_* entry point.
Test crates
Per the Test Plan:
crates/talkbank-transform/tests/
speaker_id_tests.rs L2 tests for identify_mapping /
apply_mapping / override-file I/O
transcript_merge_tests.rs L2 tests for merge invariants
adjudication_tests.rs L4 scripted-prompter tests
crates/chatter/tests/
merge_tests.rs L3 subprocess tests for chatter merge
speaker_id_tests.rs L3 subprocess tests for chatter speaker-id
adjudication_tests.rs L3 subprocess tests for chatter adjudicate
pipeline_tests.rs L3 composition tests (speaker-id + merge)
batch_tests.rs L3 batch-driver tests
sanity_scan_tests.rs L3 sanity-scan tests
(The test plan’s L1 layer, spec/constructs/speaker-id/ fragment
specs regenerated via spec/tools, was not created; the
token-cleaner and Jaccard behaviors are pinned by the L2 tests
instead.)
Data flow for chatter merge
The full call graph when an operator runs
chatter merge file1.cha file2.cha --retain CHI -o out.cha:
sequenceDiagram
actor Operator
participant CLI as chatter<br/>(cli/args/core.rs, Commands::Merge)
participant Runner as commands::transcript_merge<br/>(run_merge)
participant Merge as talkbank-transform::transcript_merge<br/>(merge_chat_files)
Operator->>CLI: chatter merge file1 file2 --retain CHI
CLI->>Runner: run_merge(file1, file2, retain, output)
Runner->>Runner: read and parse_and_validate both inputs
Runner->>Merge: merge_chat_files(f1, f2, retain, strip_tiers)
Merge->>Merge: preconditions (retain / timeline /<br/>languages / ambiguous / already-declared)
Merge->>Merge: header reconcile (@Participants concat<br/>with dedupe-on-insert; @ID / @Comment injection)
Merge->>Merge: tier strip on inserted utts · timeline sort
Merge-->>Runner: Merged (file + origins + per-input fates) or MergeError
alt Ok(merged)
Runner->>Operator: stderr: MergeNotice sentences<br/>(e.g. File 1 speakers dropped by --retain)
Runner->>Runner: serialize via into_file, write to -o path (or stdout)
Runner-->>Operator: exit 0
else Err(MergeError)
Runner-->>Operator: formatted stderr + exit code 2 (Parse: exit 1)
end
The CLI layer is thin, but not a pass-through: clap parses
arguments into the Commands::Merge variant, run_merge reads and parses
both inputs, calls the transform layer’s merge_chat_files, and translates
the Result<Merged, MergeError> into stdout/stderr/exit-code output. All
algorithm logic lives in talkbank-transform.
It parses because the merge returns a Merged, not text: the provenance is
what lets it report what the merge DROPPED. A
File 1 speaker outside --retain loses every utterance while keeping its
@Participants row, and that was invisible to every operator until the CLI
moved onto the typed API. commands::pipeline does the same and calls the
same report_merge_notices, because a warning written at one call site is a
warning the other command silently lacks.
Data flow for chatter speaker-id
The reference-mode call path:
sequenceDiagram
actor Operator
participant CLI as chatter<br/>(cli/args/core.rs, Commands::SpeakerId)
participant Runner as commands::speaker_id::modes<br/>(run_reference_mode)
participant SpkId as talkbank-transform::speaker_id<br/>(identify.rs / apply.rs)
participant Override as talkbank-transform::speaker_id<br/>(override_file.rs)
Operator->>CLI: chatter speaker-id input --reference ref --anchor CHI<br/>--inserted-role INV:Investigator
CLI->>Runner: run_speaker_id(args) → run_reference_mode
Runner->>SpkId: parse donor + reference (parse_and_validate)
Runner->>SpkId: identify_mapping(reference, anchor, donor, threshold)
SpkId-->>Runner: DonorMatchReport or Err(LowConfidence { report, threshold })
alt Ok(report)
Runner->>Runner: build MappingSpec (winner → drop,<br/>others → inserted role)
Runner->>SpkId: apply_mapping_chat(donor, mapping)
SpkId-->>Runner: relabeled CHAT String
opt --write-override
Runner->>Override: OverrideFile::read_or_default(path)
Override-->>Runner: OverrideFile
Runner->>Override: upsert(session_id, MergeOverride::auto_decision), write
end
Runner-->>Operator: relabeled output, exit 0
else Err(LowConfidence)
opt --write-pending
Runner->>Runner: record pending-adjudication entry
end
Runner-->>Operator: scores to stderr, exit 4
end
The explicit-mapping and override-file modes use the same
apply_mapping and --write-override paths but skip
identify_mapping: the mapping comes from parse_mapping_spec or
from OverrideFile::get + MergeOverride::to_mapping_spec
respectively. A fourth mode (holistic LLM judgment, via the
judgment/ submodule) produces pending-adjudication entries for
chatter adjudicate rather than deciding directly; see
Adjudication Workflow.
How this composes with the post-merge ML stages
The end-to-end pipeline batchalign3 transcribe → chatter speaker-id → chatter merge → batchalign3 align → batchalign3 morphotag crosses the talkbank-* / batchalign-* boundary
twice:
flowchart LR
subgraph BA[Batchalign, ML / audio / network]
Trans["batchalign3 transcribe"]
Align["batchalign3 align"]
Morph["batchalign3 morphotag"]
end
subgraph TB[talkbank, pure CHAT-AST]
SpkId["chatter speaker-id"]
Merge["chatter merge"]
end
Media["mp4 / wav media"] --> Trans
Trans -->|ASR.cha| SpkId
Hand["hand transcript.cha"] -->|reference| SpkId
Hand --> Merge
SpkId -->|labeled.cha| Merge
Merge -->|merged.cha| Align
Align -->|+ bullets + %wor| Morph
Morph -->|+ %mor + %gra| Final["final.cha"]
Each crossing is CHAT-file-to-CHAT-file at a stable
serialization boundary: Batchalign emits a CHAT file, talkbank
consumes it; talkbank emits a CHAT file, Batchalign consumes
it. Neither side has a runtime dependency on the other; they
exchange data through the file system (or piped stdin/stdout)
exactly as the user-facing CLI commands do. This keeps the
boundary honest: a contributor working on the merge pipeline
never needs to load a Stanza model, and a contributor working
on batchalign3 align never needs to parse a speaker-id
override file.
Public surface impact
Cumulative public API additions (the surface a downstream library consumer would see):
| Crate | New pub items | Stability |
|---|---|---|
talkbank-model | None; the merge work reuses the existing CHAT vocabulary (SpeakerCode, ParticipantRole, ParticipantEntry, IDHeader, ChatFile) unmodified | Unchanged |
talkbank-transform | speaker_id::{identify_mapping, apply_mapping, apply_mapping_chat, parse_mapping_spec, MappingSpec, SpeakerAssignment, DonorMatchReport, SpeakerIdError, CURRENT_SCHEMA_VERSION, OverrideFile, MergeOverride, OverrideMode, SpeakerAction, InsertedRoleSpec, OverrideFileError, ...} (plus the judgment and provenance surfaces); `transcript_merge::{merge_chat_files, MergeError, DEFAULT_STRIP_TIERS, | |
| Merged, Reported, MergeOrigin, ReferenceFate, DonorFate, ReferenceIdx, | ||
DonorIdx}; adjudication::; sanity_scan::` | Stable, algorithms behind these are pinned by the test plan’s L2 tests | |
chatter | New Commands enum variants (Merge, SpeakerId, Adjudicate, Pipeline, Batch, SanityScan) | Internal to the binary, not a library surface |
No existing public surface is modified or removed; this is a
purely-additive change. Existing consumers (the VS Code
extension, talkbank-lsp, chatter-desktop, batchalign)
continue to depend on the existing surface and can ignore the
additions until a workflow uses them.
Where to look for things (newcomer guide)
| Question | File |
|---|---|
“What does chatter merge do?” | book/src/chatter/user-guide/merge.md |
“What does chatter speaker-id do?” | book/src/chatter/user-guide/speaker-id.md |
| “What’s in an override file?” | book/src/chatter/integrating/merge-overrides.md |
“What types are in talkbank-transform::speaker_id?” | book/src/architecture/merge-domain-types.md |
| “Where are the tests?” | book/src/architecture/merge-test-plan.md |
| “Which crate is this code in and why?” | This page |
| “Where does the merge code live in source?” | crates/talkbank-transform/src/speaker_id/ + crates/talkbank-transform/src/transcript_merge.rs + crates/chatter/src/commands/speaker_id/ + crates/chatter/src/commands/transcript_merge.rs |
“What’s in an utterance / ChatFile / %mor tier?” | talkbank-model crate rustdoc; book/src/architecture/chat-model/chat-model.md |
| “What’s the parser do?” | book/src/architecture/parsing.md; book/src/architecture/parser-model-contracts.md |
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Adjudication Workflow
Last modified: 2026-10-02 (commit 2d7e886b)
chatter merge,chatter pipelineandchatter batchare not CLI commands. Where this page names them it describes the operation of the structural library intalkbank-transform, which is current. See the removal notice.
Status: Draft
This page specifies how human-in-the-loop adjudication fits into the merge pipeline. Several pipeline stages have decision points where the algorithm cannot or should not auto-decide; this document specifies how those refusals reach an operator, how the operator’s decision is recorded, and how the pipeline resumes with the decision applied.
The design satisfies two constraints set explicitly upstream:
- Test the interaction. Every operator-decision path must be exercisable in automated tests by providing synthetic operator choices. No hardcoded stdin reads in the decision core; a pluggable prompter abstraction is mandatory.
- Batch-then-review is the default workflow. No mid-batch
interactive pauses in the main pipeline. The optional
--interactiveflag exists on the adjudication tool only, for small-batch debugging, and rides on the same data contract.
Companion documents:
- Merge Override File Format, the on-disk record of decisions.
- Domain Types:
SpeakerMapping,MergeOverride, etc. - Test Plan: where the adjudication tests live.
- Crate Architecture: where the adjudication code lives.
Why batch-then-review, and not real-time
The finite reference workflow checks the sanity-scan handoff using unchanged basic-conversation, regular-group and phonological-group CHAT. The ordinary dialogue and regular groups do not trigger the default threshold; the nested phonological example has five anchor words and three inserted-speaker words and does trigger an advisory suggestion. A higher threshold suppresses it. This is a counting and workflow contract, not evidence of an identity swap. The test persists the pending suggestion, explicitly accepts it through the operator interface, admits the recorded mapping, and reparses its CHAT replay. It verifies retained content and the requested speaker change. Incomplete roles and unsupported mapping shapes yield no suggestion. No scan result silently changes the original override or replaces operator review. The same reference controls explicitly remove each speaker through the mapping wire API and CHAT replay. Retained turns and dependent tiers remain semantically unchanged, but the scan must refuse to infer a swap when either observed speaker population is absent. Parseable transformed CHAT alone does not establish that the heuristic has enough evidence.
Every adjudication point in the pipeline is per-session local: the operator’s decision affects this session’s output and no other session in the same batch. There is no case where an operator decision propagates forward to influence how other sessions get processed.
The cases that might appear to want real-time interaction are better served by sampling:
| Case | Real-time approach | Better approach |
|---|---|---|
| Systematic pipeline failure (everything refuses) | Watch each refusal, abort batch | Run a 5-10-session canary first; examine; abort or proceed |
| Confidence-threshold calibration on a new corpus | Adjust threshold mid-batch | Run canary; pick threshold; full batch |
| Cross-session pattern (one contributor always has PAR0 = clinician) | Notice during interactive review | Run canary; observe pattern; add per-contributor explicit mapping to orchestrator config |
| Operator wants per-session progress visibility | Watch each step | chatter adjudicate --interactive after a batch run, walking the same pending queue |
TalkBank’s operational reality makes batch-then-review strictly better:
- Batches are research-scale (hundreds of sessions per donor). Forcing operator presence during the batch run = forcing hours of babysitting.
- Overnight and batch runs are routine; interactive doesn’t work for those.
- Focused operator review of all refusals together is more efficient than scattered per-batch decisions (less context-switching; easier to spot patterns across sessions).
- Aligns with the project’s “academic research, accuracy is the standard, take however long it takes” rule: operator efficiency dominates wall-clock latency.
The --interactive flag is preserved for the small-batch
debugging case but is explicitly NOT the dominant workflow.
The known adjudication points
The pipeline has at least five points where adjudication may be needed. Each is recorded as one or more entries in the override file via the same schema.
| # | Adjudication point | Trigger | Operator’s decision | Affects |
|---|---|---|---|---|
| 1 | Speaker-id low confidence | chatter speaker-id Jaccard margin < threshold | Per-speaker mapping (drop/rename) and per-donor-code adult_roles | Speaker labeling, drop set, downstream merge |
| 2 | Parent role lookup | Parent-sample session needs MOT vs FAT decision | adult_roles[donor_speaker].code and .tag for this session | The merged file’s headers + main-tier prefixes |
| 3 | Diarization-mix flag | Operator observes Batchalign collapsed multiple real-world speakers into one label | flags = ["diarization-mixed"] plus a note | Downstream consumers know output is imperfect; might gate publication |
| 4 | Post-merge sanity scan | Auto-scan flags retained-speaker utterances with high-text-similarity inserted-speaker utterances nearby (suggesting speaker-id misclassification) | Confirm or override the original speaker-id mapping | Triggers re-run of speaker-id + merge for the session |
| 5 | Unbulleted reference file | Reference CHAT file has no time bullets; merge can’t proceed | Either bullet the reference upstream, or request fresh authoritative data | Pipeline blocked for this session pending external fix |
Points 1-4 are handled by the unified chatter adjudicate tool
specified below. Point 5 is an out-of-scope failure mode: the
adjudication tool records that the session is blocked, but the
fix lives outside this pipeline (operator contacts the
contributor or runs forced-alignment first).
Data flow
flowchart TD
Inputs["Input CHAT files +<br/>reference files"]
Orch["Orchestrator<br/>(future: tb subcommand;<br/>now: shell/script)"]
SpkId["chatter speaker-id<br/>(per session)"]
Merge["chatter merge<br/>(per session)"]
Pending["pending-adjudications.toml<br/>(workflow queue)"]
Override["overrides.toml<br/>(durable decisions)"]
Adj["chatter adjudicate"]
Operator((Operator))
Final["merged/*.cha"]
Inputs --> Orch
Orch -->|pass 1: speaker-id| SpkId
SpkId -->|exit 0 → auto entry| Override
SpkId -->|exit 4 → pending entry| Pending
Orch -->|pass 1: merge for ok sessions| Merge
Merge --> Final
Pending --> Adj
Override --> Adj
Adj <-->|prompter| Operator
Adj -->|writes decision| Override
Adj -->|removes resolved| Pending
Override -->|pass 2| Orch
Orch -.->|loop until pending empty| SpkId
The orchestrator runs two passes:
Pass 1: for every input session, run chatter speaker-id in
reference mode. Successful auto-decides write to the override
file with mode = "auto" and immediately proceed to chatter merge. Refusals (exit code 4) and other adjudication-requiring
states write a pending entry to pending-adjudications.toml
and the session is skipped for the rest of pass 1.
Pass 2 (after operator runs chatter adjudicate): the
orchestrator re-runs chatter speaker-id for the previously
skipped sessions, finding decisions in the override file
(mode = "override"). Sessions complete; pending entries are
removed.
The pipeline is idempotent: re-running pass 1 on a partially adjudicated batch produces no spurious work, sessions with already-recorded decisions skip to merge directly.
The pending-adjudications artifact
Separate from the override file, a pending-adjudications.toml
file holds in-flight workflow state. Its purpose is to carry
the evidence the operator needs (per-speaker scores, opening
utterance previews) from the orchestrator’s pass 1 to the
adjudication tool, without polluting the override file with
“to-do” entries.
Schema
schema_version = 2
[[entries]]
session_id = "session-102-t1"
kind = "speaker-id-low-confidence"
created_at = "2026-05-27T11:00:00-04:00"
# Inputs the adjudication tool needs:
input_path = "asr/session-102-t1.cha"
reference_path = "chi-only/session-102-t1.cha"
anchor_speaker = "CHI"
# Evidence for the operator:
scores = { PAR0 = 0.6286, PAR1 = 0.3457 }
margin = 1.82
threshold_used = 2.0
# Opening turns (first N utterances per speaker) for context:
preview = """
*CHI: they start to bite . [0_1708]
*PAR0: They start to bite . [75_1165]
*PAR1: They do what . [1515_2245]
... (further preview)
"""
# Suggested defaults the operator can accept-as-is:
suggested = { mapping = { PAR0 = "drop", PAR1 = "rename" }, adult_roles = { PAR1 = { code = "INV", tag = "Investigator" } } }
[[entries]]
session_id = "session-103-t1-parent"
kind = "parent-role-lookup"
# ... different evidence for the MOT-vs-FAT case ...
Schema characteristics
kinddiscriminates the adjudication type (one ofspeaker-id-low-confidence,parent-role-lookup,diarization-mix-review,sanity-scan-misclassification). Each kind has its own required field set; the adjudication tool dispatches onkindto choose the right prompt template and the right validator for the operator’s response.suggestedcarries what the algorithm WOULD have chosen had the threshold been lower (for speaker-id) or a parsed default (for parent-role). The operator can accept-as-is or override.- Entries are a
[[entries]]array of tables (not a session-keyed[<session_id>]map) because the same session could conceivably have multiple pending decisions (e.g., a speaker-id refusal AND a parent-role lookup), each a separate array entry.
Lifecycle
- Written by: the orchestrator’s pass 1, when
chatter speaker-idexits with code 4 or when other adjudication triggers fire. - Consumed by:
chatter adjudicate, which reads it, prompts the operator entry-by-entry, writes decisions to the override file, and removes resolved entries. - Cleaned up: an empty
entriesarray is the “all clear” state; pass 2 of the orchestrator can proceed.
chatter adjudicate, CLI surface
A new chatter subcommand in chatter. Its job is to walk
a pending-adjudications file and write decisions to an override
file.
chatter adjudicate <PENDING_FILE> --override-file <OVERRIDE_FILE> [OPTIONS]
ARGUMENTS:
<PENDING_FILE> Path to pending-adjudications.toml.
REQUIRED OPTIONS:
--override-file <PATH>
Path to the override file (created if missing, appended if
existing). Decisions go here.
OPTIONS:
--interactive
(default) Prompt the operator for each pending entry via
a terminal UI. This is the only mode for v1; later UI
backends may add e.g. --backend=web for web-served prompts.
--scripted <PATH>
Read pre-canned decisions from a TOML file. Used in tests
and in automated bulk-decision workflows (e.g., the
operator has prepared a decision sheet in advance).
Mutually exclusive with --interactive.
--kind <KIND>
Process only pending entries whose `kind` matches. Useful
when the operator wants to batch through one class of
decision at a time (e.g., do all parent-role lookups
first, then all speaker-id refusals).
--skip-on-error
If the operator's response cannot be applied (e.g., they
typed an invalid speaker code), log and skip rather than
abort. Default: abort on first invalid response.
--operator <NAME>
Operator identifier recorded in override entries.
Default: $USER.
--dry-run
Read pending and prompt the operator, but do NOT write to
the override file. Useful for previewing what decisions
look like before committing.
Exit codes:
| Code | Meaning |
|---|---|
| 0 | All pending entries decided; pending file updated |
| 1 | I/O error (missing file, unparseable, write failure) |
| 2 | Operator-supplied decision rejected as invalid (when --skip-on-error not set) |
| 3 | Internal error |
| 4 | Operator deferred at least one entry (used :skip in the prompt); pending file still has entries |
The --scripted mode is the testability seam. A scripted
decision file looks like:
schema_version = 2
[[decisions]]
session_id = "session-102-t1"
kind = "speaker-id-low-confidence"
choice = { kind = "accept-suggested", note = "verified by listening" }
[[decisions]]
session_id = "session-103-t1-parent"
kind = "parent-role-lookup"
choice = { kind = "choose-role", adult_roles = { PAR0 = { code = "FAT", tag = "Father" } }, note = "per contributor data sheet" }
The current library reads decisions in file order and requires each consumed
decision’s session_id to match the next pending entry. The pending entry’s
typed kind determines which decision shapes are admissible; the scripted
entry’s kind is descriptive metadata, not an additional matching key.
A missing decision, mismatched next session, or incompatible decision shape
refuses without consuming that pending request. Do not assume that unconsumed
trailing scripted decisions are rejected.
The prompter abstraction (testability)
The adjudication tool’s core flow is:
// pseudocode, actual signatures live in talkbank-transform
pub fn run_adjudication(
pending: PendingAdjudications,
override_file: &mut OverrideFile,
prompter: &mut dyn Prompter,
operator: OperatorId,
) -> Result<AdjudicationOutcome, AdjudicationError> {
for entry in pending.entries() {
let context = build_context(entry);
let decision = prompter.ask(&context)?;
apply_decision(override_file, entry, decision, &operator);
}
Ok(...)
}
pub trait Prompter {
fn ask(&mut self, context: &AdjudicationContext)
-> Result<OperatorDecision, PrompterError>;
}
Production implementations:
TerminalPrompter: printscontextto stdout, reads operator response from stdin. Used by--interactive.
Test implementations:
ScriptedPrompter::from_decisions(Vec<(SessionId, OperatorDecision)>), returns each decision in turn, errors if asked for an unprovided session. Used by L2 transform tests.ScriptedTomlPrompter::read(path): reads the same TOML format as--scripted. Used by L3 CLI tests so subprocess tests and library-level tests share fixture format.
This means:
- Every adjudication test path is automated. No subprocess
PTY hackery, no expect-script DSL. Tests construct
ScriptedPrompter, run the adjudication core, assert on the resultingOverrideFile. - The terminal UI is dumb. All it does is
Display-format the context and parse the operator’s response into anOperatorDecision. No business logic in the UI layer. - Future UI backends (VS Code, web) implement
Prompterand drop in. The adjudication core is unchanged.
The OperatorDecision type
pub enum OperatorDecision {
/// Accept the algorithm's suggested mapping verbatim.
AcceptSuggested { note: Option<String> },
/// Override with an operator-supplied mapping (speaker-id).
OverrideMapping {
mapping: SpeakerMapping,
note: Option<String>,
},
/// Override the inserted role(s) only (parent-role lookup).
OverrideInsertedRole {
adult_roles: BTreeMap<String, InsertedRoleSpec>,
note: Option<String>,
},
/// Add or update flags on an existing entry.
Flag { flags: Vec<MergeFlag>, note: Option<String> },
/// Defer this entry; leave it in pending for later review.
Defer { reason: String },
/// Mark the session as blocked (e.g., unbulleted reference);
/// requires upstream action before pipeline can resume.
Block { reason: String },
}
Each variant maps cleanly to one or more adjudication kinds:
| Kind | Allowed OperatorDecision variants |
|---|---|
speaker-id-low-confidence | AcceptSuggested, OverrideMapping, Defer |
parent-role-lookup | AcceptSuggested, OverrideInsertedRole, Defer |
diarization-mix-review | Flag, Defer |
sanity-scan-misclassification | OverrideMapping, Flag, Defer |
| (any) | Block is always available |
The kind → allowed-variants mapping is enforced by the
adjudication tool: a kind = "parent-role-lookup" entry that
gets an OverrideMapping decision is rejected with a clear
error (AdjudicationError::DecisionKindMismatch).
Operator terminal UX (interactive mode)
What the operator sees when running chatter adjudicate pending.toml --override-file overrides.toml --interactive:
═══════════════════════════════════════════════════════════════
ADJUDICATION [1 / 14] session-102-t1 kind = speaker-id-low-confidence
═══════════════════════════════════════════════════════════════
Reference file: chi-only/session-102-t1.cha
Donor file: asr/session-102-t1.cha
Anchor speaker: CHI
Per-speaker Jaccard scores against reference's CHI:
PAR0 = 0.6286 ◄── higher
PAR1 = 0.3457
margin = 1.82× (threshold was 2.00×)
Opening turns side-by-side:
*CHI [0_1708] they start to bite .
*PAR0 [75_1165] They start to bite .
*PAR1 [1515_2245] They do what .
*CHI [1708_5966] they put up their shields at some point .
*PAR0 [2755_4405] They put up those heels .
*PAR1 [4865_6045] At some point oh .
(3 more turns shown; press 'm' for more)
Algorithm-suggested mapping:
PAR0 → drop (winner, matches CHI content)
PAR1 → rename to INV:Investigator
Your decision?
[a] Accept suggested
[o] Override mapping
[f] Flag and defer
[d] Defer (review later)
[b] Block (needs upstream fix)
[m] Show more context
[p] Play media (uses $TB_MEDIA_PLAYER)
[q] Quit (save progress and exit)
>
When the operator types a and then is prompted for an
optional note, the tool writes the decision to the override
file and advances to the next pending entry.
The [p] Play media action is just a wrapper around
Command::new($TB_MEDIA_PLAYER).arg(media_path).spawn(), the
adjudication tool doesn’t bundle an audio player. The operator
configures their preferred player via the environment.
Adjudication contexts beyond speaker-id
The same chatter adjudicate tool handles all five adjudication
points by dispatching on kind. For each, the displayed
context and the allowed decisions differ:
parent-role-lookup
Shown context: the session is a parent sample (basename
contains parent-suffix conventionally, or contributor data
sheet says so). The merged output needs an inserted-role code
of MOT, FAT, or PAR. The operator picks.
Session: session-103-t1-parent
Kind: parent-role-lookup
This is a parent-sample session. The merged file's inserted
speaker (currently labeled PAR0 → ???) needs a CHAT role.
Contributor data sheet (if attached): not available
Audio preview duration: 8m 14s
Algorithm-suggested: INV : Investigator (default for ambiguity)
Your decision?
[a] Accept suggested (INV : Investigator)
[m] MOT : Mother
[f] FAT : Father
[p] PAR : Adult (gender unknown)
[c] Custom role
[d] Defer
[b] Block (needs upstream metadata)
>
diarization-mix-review
Triggered by the operator (or a post-merge auto-scan) observing
that an ASR speaker’s content mixes real-world speakers. The
adjudication is to add the "diarization-mixed" flag plus a
note explaining the mix.
sanity-scan-misclassification
Triggered by the post-merge sanity scan when a retained-speaker utterance has high text similarity with a temporally-adjacent inserted-speaker utterance. The operator either confirms (“the original speaker-id was wrong, swap the mapping”) or overrides (“the duplication is real, both speakers said the same thing at the same time”).
Resumption and re-adjudication
The pending-adjudications file is the source of truth for
“what still needs deciding.” If the operator quits mid-review
(via [q] or process-kill), the next chatter adjudicate
invocation picks up where they left off, already-decided
entries have already been removed from pending and written to
the override file.
Re-adjudication of an already-decided entry is a planned
extension, not yet implemented. The proposed interface would
load the existing override entry, present it as the “current
decision,” and ask the operator whether to keep or replace it;
the operator’s decision would overwrite the entry, and the prior
decision would be preserved in a history array on the entry
(recording the prior mode, mapping, operator, decided_at,
and note). The proposed invocation shape (not a working command
today) is:
# Proposed, not yet implemented:
chatter adjudicate --re-adjudicate <SESSION_ID> --override-file overrides.toml
It needs a small override-file schema extension, a per-entry
optional history: Vec<MergeOverride> field. This is a minor,
additive schema change, comparable to the 2026-06 engine/judgment
addition (no version bump needed either way), not a breaking one;
schema_version is already 2 as of the adult_roles map (see
Merge Override File Format §Future schema
changes),
so a future breaking change to this schema would need schema_version = 3, not 2.
Composition with the orchestrator
The orchestrator (proposed tb merge or similar) drives the
pipeline. Its high-level flow:
// pseudocode for the orchestrator's main loop
let inputs = discover_input_sessions(input_dir);
let override_file = OverrideFile::read_or_default(override_path);
let mut pending = PendingAdjudications::default();
for session in inputs {
if let Some(decision) = override_file.get(&session.id) {
// Already adjudicated; apply directly.
let labeled = apply_mapping(&session.donor, &decision.mapping)?;
let merged = merge(&session.reference, &labeled, &session.retain)?;
write_merged(merged, &session.output_path)?;
} else {
// Try auto-decide.
match identify_mapping(&session.donor, &session.reference, ...) {
Ok(mapping) => {
let labeled = apply_mapping(&session.donor, &mapping)?;
let merged = merge(...)?;
write_merged(merged, &session.output_path)?;
override_file.insert(session.id.clone(), record_auto_decision(&mapping));
}
Err(SpeakerIdError::LowConfidence { scores, margin, threshold }) => {
pending.push(PendingEntry::speaker_id_low_confidence(
session.id.clone(),
scores, margin, threshold,
/* preview */ build_preview(&session),
));
}
Err(other) => return Err(other),
}
}
}
pending.write(pending_path)?;
override_file.write(override_path)?;
if !pending.is_empty() {
eprintln!(
"Pipeline complete for {} sessions; {} sessions need adjudication.\n\
Run: chatter adjudicate {} --override-file {}",
decided_count, pending.len(), pending_path, override_path
);
return Ok(ExitCode::NeedsAdjudication);
}
The orchestrator is the layer that hasn’t been designed yet at
the type level. It’s likely a tb subcommand (since tb is
the workflow tool for multi-repo / multi-step ops), with a
fallback shell-script form for the v0 pipeline.
What this design does NOT cover
- The orchestrator binary itself. That’s a separate design pass; this doc only specifies the contract between the pipeline stages and the adjudication tool.
- GUI/web adjudication backends. v1 is terminal-only. The
Promptertrait is the extension point; future backends implement it. The data contract (pending.toml,overrides.toml) does not change. - Audio playback / waveform display. v1 launches the
operator’s
$TB_MEDIA_PLAYERand gets out of the way. A future TUI with inline audio scrubbing is conceivable but is a major UI project, not v1. - ML-suggested decisions. A future version could feed
pending entries to a classifier that pre-fills “suggested”
with model output. Out of scope; the
suggestedfield exists today as a hook.
Test coverage
The library’s current failure contract is covered by a canonical-reference workflow: exhausted input, an out-of-order session, or an incompatible decision retains the failing request and all later requests field-for-field in their wire representation and in their original order. Accepted overrides remain in the in-memory document, and a retry resolves only the remaining requests. A private queue owner borrows the pending destination for the run; its only advancement operation first obtains and successfully applies a decision. Dropping the owner restores the uncommitted suffix on either success or error.
Before an override can be committed, a private prepared-decision state admits it
through the existing to_mapping_spec conversion. A rename without its required
adult role yields InvalidDecisionMapping, retaining the pending suffix. This
proves conversion readiness, not complete CHAT validity, role vocabulary, or
collision freedom. Raw public override records still require conversion at
their consumers. The CLI reports this refusal with exit status 2 before writing
either file.
This is an in-memory contract, not a transactional filesystem guarantee. The CLI currently persists the override and pending files only after a successful run. The corpus supplies real parsed speaker identities; scripted role choices are authored test inputs, not inferred facts about those recordings.
The canonical sanity-scan reference also exercises the real TOML boundary: pending suggestions survive file write/read field-for-field, both readers refuse an unsupported schema version rather than silently defaulting, and an explicit scripted TOML acceptance is required before mapping replay. This checks individual file operations, not atomicity across the pending and override files or crash recovery.
The list below is the design’s target inventory, not a claim that all listed behaviors are implemented or tested. Current tests exercise supported accept, override, and role-choice paths plus the in-memory failure/resume contract above. Defer and block decisions remain planned. See the Test Plan for the broader intended inventory:
- Each adjudication kind’s happy path (operator accepts suggested, decision written to override file)
- Each adjudication kind’s override path (operator types an alternative, decision validated and recorded)
- Each adjudication kind’s defer path (entry stays in pending)
- Each adjudication kind’s block path (entry marked blocked; pipeline reports blocker)
- Re-adjudication path (operator changes their mind; prior
decision preserved in
history) - Mutually-exclusive flag enforcement (
--interactive+--scriptedrejected) - Invalid operator response handling (with and without
--skip-on-error) - Schema-version refusal on the pending file
- Empty pending file (no-op, exit 0)
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Errors, CHAT core
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
The error infrastructure used across all CHAT-core crates
(talkbank-model, talkbank-parser, talkbank-transform,
chatter, talkbank-lsp). Defined in the
errors module of talkbank-model.
External runtime/application errors that live outside this repo’s CHAT core are documented separately in their owning projects. For the diagnostic UX standard that applies within this workspace, see error-diagnostics-ux.
Core Types
ParseError
Every diagnostic is a ParseError:
pub struct ParseError {
pub code: ErrorCode,
pub severity: Severity,
pub location: SourceLocation,
pub context: Option<ErrorContext>,
pub message: String,
}
ParseError::build(code) returns a required-field typestate builder. Supply a
message and a location in either order, then call finish() to obtain the
diagnostic directly. Optional severity, context, suggestion and labels survive
these transitions. Typestate guarantees field presence, not source validity;
use the source-aware SourceLocation constructors when admitting byte ranges.
#![allow(unused)]
fn main() {
use talkbank_model::{ErrorCode, ParseError};
let diagnostic = ParseError::build(ErrorCode::ParseFailed)
.message("Input cannot be parsed")
.at(0, 1)
.finish();
}
finish() returns the diagnostic directly: there is no try_finish() and no
ParseErrorBuilderError, because an incomplete builder cannot finish. Streaming pipeline source admission
reports the producer’s original diagnostic once and returns a parse failure;
it neither invents a source location nor fabricates an empty recovered document.
ErrorCode
Error codes follow a structured numbering scheme:
| Range | Category |
|---|---|
| E1xx | Encoding |
| E2xx | Words and content |
| E3xx | Main tier (speakers, terminators, content, retraces) |
| E4xx | Dependent tier structure |
| E5xx | Headers |
| E6xx | Dependent tier validation |
| E7xx | Alignment (%mor, %gra, %pho, %wor) |
| W1xx-Wxxx | Warnings (same categories) |
Codes are grouped by range as above. The numbering is a navigational aid, not
the authority on where a code is caught: most codes are emitted at the layer
suggested below, but a few main-tier checks (for example undeclared-speaker and
retrace structure) are validation-layer despite their E3xx number. The
per-code Layer in spec/errors/ is authoritative.
flowchart LR
subgraph "Parser layer\n(parser.parse_chat_file())"
E1["E1xx\nEncoding\n(BOM, charset)"]
E2["E2xx\nWords and content\n(word syntax, events,\noverlap markers)"]
E3["E3xx\nMain tier\n(speaker, content,\nterminator, retraces)"]
E4["E4xx\nDependent tier structure\n(tier presence, format)"]
E5["E5xx\nHeaders\n(format, required fields,\nparticipant resolution)"]
end
subgraph "Validation layer\n(validate_with_alignment)"
E6["E6xx\nDependent tier validation\n(tier name/format)"]
E7["E7xx\nAlignment\n(%mor/%gra/%pho/%wor counts,\nGRA indices, orphaned tiers)"]
end
W["Wxxx\nWarnings\n(same categories,\nnon-fatal)"]
E1 ~~~ E2 ~~~ E3 ~~~ E4 ~~~ E5
E6 ~~~ E7
The source of truth for error-code details is spec/errors/. Maintainers can
generate a local markdown reference set under docs/errors/ with
just spec-gen when they need a browsable error catalog while working on
diagnostics.
Severity
Error: must be fixed; indicates invalid CHAT.Warning: should be fixed; indicates questionable but parseable CHAT.
SourceLocation and Span
Byte offsets into the source text:
#![allow(unused)]
fn main() {
pub struct SourceLocation { pub start: usize, pub end: usize }
pub struct Span { pub start: usize, pub end: usize }
}
ErrorContext
Carries the source fragment around the error location:
pub struct ErrorContext {
pub source_text: String,
pub span: Span, // Relative to source_text, not the document location.
pub expected: SmallVec<[String; 2]>,
pub found: String,
pub line_offset: Option<usize>,
}
ErrorSink Trait
The central abstraction for error reporting:
flowchart LR
val["Validator / Parser"]
pe["ParseError\ncode + severity +\nlocation + message"]
sink["ErrorSink trait\n.report()"]
vec["ErrorCollector\ncollect to Vec"]
chan["ChannelErrorSink\ncrossbeam channel\n(feature = channels)"]
asyncchan["AsyncChannelErrorSink\ntokio mpsc"]
cfg["ConfigurableErrorSink\n(talkbank-transform)\npresentation policy"]
null["NullErrorSink\nno-op"]
val --> pe --> sink
sink --> vec & chan & asyncchan & cfg & null
pub trait ErrorSink {
fn report(&self, error: ParseError);
}
All parsing and validation functions accept &impl ErrorSink rather
than returning errors directly. This allows:
- Collecting all errors (for batch processing).
- Printing errors in real-time (for interactive use).
- Filtering by severity or code.
- Counting errors without storing them.
The trait uses &self (not &mut self) so it can be shared across
threads. Implementations typically use interior mutability
(Mutex<Vec<ParseError>>).
ErrorCollector is the in-memory collector in
errors/collectors.rs. The stored-diagnostics role is explicit in
both code and docs.
Module layout in talkbank-model:
errors/error_sink.rs: trait and lightweight forwarding sinks.errors/collectors.rs: in-memory collectors and counters.errors/async_channel_sink.rs: Tokio-channel streaming.errors/offset_adjusting_sink.rs: remove synthetic wrapper offsets.errors/rebased_sink.rs: translate raw document locations and labels while retaining the diagnostic’s self-contained source context. Use before display enhancement converts secondary labels into snippet-relative coordinates.errors/tee_sink.rs: forward diagnostics to both sinks.
ConfigurableErrorSink is the one adapter that does NOT live here: it
applies a PresentationPolicy (what a reader is shown), which belongs to
talkbank-transform so that talkbank-cache cannot reach it and fold a
display preference into the validation cache key. See the leniency-policy
chapter.
ChannelErrorSink is opt-in behind the channels feature so the
default talkbank-model dependency does not pull in crossbeam just
to own the core error trait and in-memory collectors.
Two Error Layers
Errors are detected at two layers. This distinction matters for spec testing.
-
Parser layer: structural errors caught during
parser.parse_chat_file(). These prevent the file from being fully parsed (missing@Begin, invalid syntax). Parser-layer specs test thatparser.parse_chat_file()returnsErr. -
Validation layer: semantic errors caught by
validate_with_alignment()after a successful parse. The file parsed correctly but violates constraints (%moralignment mismatch, undeclared speakers). Validation-layer specs test that validation reports specific error codes.
Adding a New Error Code
- Add the variant to
ErrorCodeincrates/talkbank-model/src/errors/codes/error_code.rswith a#[code("Exxx")]attribute. - Create a spec file in
spec/errors/Exxx-description.mdfollowing the existing template. - Construct
ParseError::new(ErrorCode::YourVariant, ...)at the detection site in the parser or validator. - Regenerate the affected spec artifacts with the current
spec/toolsgenerators (just spec-gen, and optionallyjust spec-gen). - Run the concrete verification commands from
book/src/contributing/dev-checks.md.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Validation
Status: Current Last modified: 2026-10-07 (commit 5e895791)
Validation levels and the pre/post gates a pipeline can build on. For the error-code infrastructure (codes, sinks, severities, layers) see chat-core-errors; for the diagnostic UX standard see error-diagnostics-ux.
All validation logic is Rust. talkbank-model::validation owns CHAT-core
validation; talkbank_transform::validate owns the gate functions
validate_to_level and validate_output.
Timing presence is not admission
ChatFile::timing_evidence() observes main-tier bullets (including recursively
nested internal bullets) and actual %wor timing through the same owner as
E544. Its recorded witness borrows both the document and its actual bullet;
callers cannot construct a witness from unrelated values. The observation does
not validate intervals, recovery, headers or alignments, and grants no write
permission. Regeneration consumers can distinguish restored timing from an
outstanding linkage obligation while still requiring complete output admission.
Legacy workflow validity levels
These levels are partial workflow checks, not complete CHAT-validity or output certificates. Complete source admission and checked construction remain the boundaries for retained input and writable output.
ValidityLevel (in talkbank-model::pipeline) is cumulative: each level
includes every check below it.
| Level | Name | Checks |
|---|---|---|
| L0 | Parseable | no parse errors |
| L1 | StructurallyComplete | @Participants and @Languages present, all speaker codes declared, every utterance has a terminator |
| L2 | MainTierValid | well-formed words, valid timing bullets if present |
The levels exist so a consumer can state the minimum quality its work needs and reject bad input BEFORE spending compute on it, rather than discovering the problem in the output.
use talkbank_transform::validate::validate_to_level;
// parse_errors come from the parser (typically parse_lenient).
validate_to_level(&file, &parse_errors, ValidityLevel::MainTierValid)?;
validate_to_level returns EVERY failure found up to the requested level, not
just the first. The L0 gate surfaces the first parse error’s code, source
excerpt and byte span in its message, so a user can locate the problem without
reading logs.
flowchart TD
cmd["a pipeline stage"]
gate["validate_to_level(file, parse_errors, required_level)"]
check{"meets the required\nValidityLevel?"}
reject["reject early with diagnostics;\nno compute spent"]
proceed["run the stage"]
cmd --> gate --> check
check -->|"no"| reject
check -->|"yes"| proceed
A selected level cannot waive invalid retained input. A command may tolerate defective generated tiers only through an admitted replacement plan that actually discards and regenerates them. Forced alignment does not acquire complete admission merely because a file is parseable.
Post-serialization validation
validate_output answers a narrower question: did a transformation DEGRADE the
file? It checks that every utterance still has a terminator (CA transcripts are
exempt, since terminators are optional under @Options: CA) and then applies
whatever command-specific checks it knows.
Known defect, recorded here rather than left for the next reader to
rediscover. validate_output takes the command as a &str and dispatches
with match command { "morphotag" => ..., "align" => ..., _ => {} }. Two
things are wrong with that and neither is cosmetic:
- The catch-all silently skips every command-specific check. A caller passing
a typo, or any command the match does not list, gets the terminator check
and nothing else, with no error and no warning. It type-checks perfectly.
clippy::wildcard_enum_match_armcannot see this one, because the match is over an open set of strings rather than a closed enum. - The strings name commands belonging to a downstream ML pipeline, which is workflow-specific knowledge embedded in a general-purpose CHAT library.
The fix is a closed enum owned by this crate, so an unhandled command is a compile error and the general library stops naming a particular consumer’s verbs. It is left undone here only because the signature is public API with an out-of-repo caller, so changing it is a coordinated change rather than a drive-by.
Severity posture
- Errors block output. Nothing writes CHAT that has error-level failures.
- Warnings are reported and do not block, because legacy corpora contain widespread minor violations and must remain processable.
The distinction is sharpest for %gra: pre-existing broken %gra in old
corpora is warned about rather than blocked, so files that already shipped that
way still round-trip, while newly GENERATED %gra is validated strictly before
writeback. The asymmetry is deliberate. Data we are responsible for producing is
held to a higher standard than data we merely have to keep readable.
Verification
The commands are in Developer Verification Checks
and Testing and Quality Gates; this page
does not duplicate them. Labels like G0-G14 come from a predecessor workspace
and name nothing here.
The reference corpus is a synthesized regression signal, not a validity authority; treating it as one leads someone to weaken a validator so a fixture stays green. When a change makes a reference file fail, adjudicate the FILE.
Known limitations
- Validation is deliberately permissive on legacy data. Some checks warn rather than error so legacy corpora remain processable while the issue is still surfaced.
%worword counts are not validated against the main tier.%woris a timing-annotation sidecar, so legacy files may carryxxx, fragments or nonwords in%worwithout producing alignment errors. Timing consumers can request a typed binding.Driftedfails closed without making the legacy file invalid.CountMatchedpermits a canonical display-token comparison but exposes no timing slots. Only the laterCorroboratedstate exposes timing, so a detectable same-count lexical edit also fails closed.- Cross-utterance quotation validation is off by default
(
enable_quotation_validation): the walker exists but is not wired into the standard gate. - Some error specs have no validator yet.
just spec-statusis the authority on which, and on how many; it derives the answer from the specs rather than from a count written in prose.
Consumers outside this repository
chatter contains no ML-pipeline code. Downstream consumers embed these crates and add their own gates, bug reporting and cache invalidation; how a given pipeline reports a validation failure, and where it writes it, is documented by that pipeline, not here.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
CHECK Assessment
2026-10-02 (commit 2d7e886b)
Chatter aims to provide sensible, understandable CHAT validation, not to imitate every behavior of CLAN CHECK. Grammar rejection and typed-model invariants can satisfy the same requirement with different diagnostics. Deliberate, justified divergences are resolved decisions, not unfinished work.
Current baseline
Inventoried assessment complete: 164 of 164 entries adjudicated; 0 unresolved gaps.
| Disposition | Entries |
|---|---|
| parity | 131 |
| divergence | 10 |
| no_obligation | 23 |
| gap | 0 |
Completion means every entry in this finite CHECK assessment inventory has been adjudicated. It does not certify every CHECK emission path, all possible CHAT input, 100% code coverage, or Chatter 1.0 readiness. The inventory is maintained against the committed CLAN source reference; it is not a claim about an unexamined upstream revision.
When to reopen
Reopen an obligation only for a concrete, independently justified CHAT requirement and evidence that Chatter misses invalid input or rejects legitimate input. A different diagnostic code, wording, count, order or rejection stage is not a gap. Upstream changes prompt scoped source and intent review, not automatic imitation.
Developer contract
Assess CHECK source and intent before adopting suspicious behavior. Keep recovery and reporting idiomatic to the tree-sitter CST and typed AST, with source-bound evidence and producer-owned context. Never add parallel parsers, reparsing or raw-text diagnostic classifiers to mimic CHECK. When available structural evidence cannot prove a narrower fault, report an honest broader error. Preserve invalid-input rejection and valid controls. Experimental re2c agreement is not a completion prerequisite.
Authority and verification
The single authored authority is crates/talkbank-parser-tests/tests/check_parity/manifest.json. This page is its generated view; counts and completion are derived, never entered separately. The shared talkbank-spec-vocabulary::check_assessment types are used by the report, fixture harness and spec-status command.
Regenerate with just check-mapping-gen (also included in just regen). Run cargo test -p talkbank-parser-tests --test integration check_mapping_audit to detect stale generated documentation. The normal integration suite includes that check.
chatter_matches_check tests the manifest’s fixture expectations; manifest_agrees_with_clan_reference checks source-reference consistency. Neither a generated page nor an adjudication count is a fresh runtime receipt. clan_check_grounding separately checks an identified CHECK executable through CHATTER_CLAN_RUN; it does not establish that CHECK’s policy is correct.
The separate docs/audits/check-parity-audit.md is only a code-mapping index, not a second assessment or completion authority. Missing mappings do not mean missing validation.
Adjudication ledger
The codes below pin current regression expectations, not a requirement to reproduce CHECK diagnostics. Each rationale belongs to the manifest.
CHECK 1 — parity
First substantive byte is an invalid top-level line start (x) before an otherwise ordinary CHAT file. CLAN emits CHECK 1; chatter rejects the bad line start with E316.
Fixture: CHECK_001_bad_initial_character.cha. Chatter regression codes: E316.
CHECK 2 — parity
A complete CHAT preamble is followed by a main tier whose speaker code is not terminated by a colon. CLAN emits CHECK 2; chatter rejects the malformed line with E316.
Fixture: CHECK_002_missing_main_tier_colon.cha. Chatter regression codes: E316.
CHECK 3 — parity
The character after the speaker-tier colon must be whitespace (check.cpp ~1756: !isSpace after check_FoundText); the fixture writes *CHI:hello . with no tab. CLAN emits CHECK 3; chatter rejects the line as unparsable content (E316).
Fixture: CHECK_003_no_space_after_colon.cha. Chatter regression codes: E316.
CHECK 4 — parity
A space instead of a TAB after the speaker code (*CHI: hi). CLAN 4; chatter requires the tab in-grammar and surfaces the recovery node as E316.
Fixture: CHECK_004_space_after_tier.cha. Chatter regression codes: E316.
CHECK 5 — parity
@Begin: (a stray colon on the no-colon @Begin header) parses as a valid @Begin followed by a stray : ERROR node. A whole-tree recovery-node backstop in parse_lines_with_old_tree surfaces any ERROR node that streaming lowering would drop as E316 UnparsableContent (recovery is not validity); the AST is still produced. CLAN rejects the file (codes 5 and 6), and so does chatter. CLAN CHECK 5 (co-emitted with 6).
Fixture: CHECK_005_begin_illegal_colon.cha. Chatter regression codes: E316.
CHECK 6 — parity
@Begins (a misspelled @Begin) parses as @Begin plus a trailing s ERROR node, which the whole-tree recovery-node backstop surfaces as E316. CLAN rejects it (codes 6 and 17). CLAN CHECK 6 (co-emitted with 17).
Fixture: CHECK_006_begin_malformed.cha. Chatter regression codes: E316.
CHECK 7 — parity
File has no @End header at all. CLAN 7 (@End missing at end of file); chatter E502 MissingEndHeader.
Fixture: CHECK_007_missing_end.cha. Chatter regression codes: E502.
CHECK 8 — parity
A body line that does not begin with @, %, * or TAB (free text). CLAN 8 (expected @ %% * TAB); chatter E326 surfaces the unparsable main-tier start.
Fixture: CHECK_008_bad_line_start.cha. Chatter regression codes: E326.
CHECK 9 — parity
Speaker/tier name before the colon exceeds SPEAKERLEN=1024 characters (check.cpp:2395, first pass). chatter rejects via undeclared-speaker (E522/E308/W108) and speaker-ID-length E307 (max 7 chars), not via a dedicated raw length cap; CHECK aborts first pass (second pass not attempted).
Fixture: CHECK_009_tier_name_too_long.cha. Chatter regression codes: E307, E308, E522.
CHECK 10 — divergence
UTTLINELEN buffer bound (18000): a logical tier whose text across continuation lines exceeds 18000 chars draws check_err(10) in the check_OverAll char scan (check.cpp 2403). A SINGLE physical line over 18000 instead hits a separate hard-exit guard (check.cpp 2329, ‘Speaker turn is longer then’, cutt_exit(1), no (NN) code), so the fixture splits ~20k chars over tab continuation lines. Adjudication: the limit is a fixed C buffer size, an implementation artifact, not CHAT semantics; chatter has no tier-length limit by design (long tiers are valid and supported) and intentionally accepts.
Fixture: CHECK_010_tier_text_too_long.cha. Chatter regression codes: none.
CHECK 11 — divergence
External depfile membership is not a universal CHAT validity requirement. Empty POS fields are a separate structural fault, rejected by morphology recovery (E702; see the E760 examples).
Fixture: CHECK_011_undeclared_symbol_pct_fac.cha. Chatter regression codes: none.
CHECK 12 — parity
An @ID tier whose 3rd field (speaker code) or 8th field (role) is empty (check.cpp:5060/5084). E511 ‘ID header speaker field cannot be empty’ is the direct equivalent. CLAN also emits 142 on the same fixture.
Fixture: CHECK_012_missing_speaker_in_id.cha. Chatter regression codes: E511, E523.
CHECK 13 — parity
Same speaker declared twice in @Participants. chatter E549 DuplicateSpeakerDeclaration.
Fixture: CHECK_013_duplicate_speaker.cha. Chatter regression codes: E549.
CHECK 13 — parity
Two @ID lines naming the same speaker (CHI). Distinct from the @Participants-duplicate variant (CHECK_013_duplicate_speaker.cha): CLAN 13 fires for duplicate @ID lines too; chatter reports E549 via check_duplicate_id_headers. CLAN CHECK 13.
Fixture: CHECK_013_duplicate_id.cha. Chatter regression codes: E549.
CHECK 14 — parity
Leading spaces before the *tier code. CLAN 14 (also 4); chatter E326.
Fixture: CHECK_014_spaces_before_tier.cha. Chatter regression codes: E326.
CHECK 15 — parity
An @ID tier whose role field does not match any role listed in the depfile @Participants template (check.cpp:5087). E532 ‘Invalid participant role’ is the direct equivalent. CLAN also emits 142 on the same fixture.
Fixture: CHECK_015_illegal_role.cha. Chatter regression codes: E532.
CHECK 16 — parity
A speaker CODE containing non-printable-ASCII in @Participants (cnt==0 field scan, chars outside 33..126; check.cpp 2084-2086). Emits via the check_trans_err MACRO (check.cpp 74-79), invisible to the literal check_err(N,) counter. Empirically grounded: real CLAN emits (16) on this fixture. Non-ASCII speaker IDs are invalid by the 2026-09-23 maintainer ruling. Chatter reports syntax code E307 through the shared speaker assessment, not the unrelated E522 participant-join code; header recovery can emit additional diagnostics.
Fixture: CHECK_016_extended_chars_speaker_code.cha. Chatter regression codes: E307.
CHECK 17 — parity
Undeclared dependent tier %zzz: (neither %x* nor a chatter-supported standard tier). chatter rejects it via E605 UnsupportedDependentTier (Error severity). The supported set deliberately excludes the legacy %grt/%tra/%trn/%umor tiers retired with the UD-%mor consolidation, so chatter is intentionally stricter than CLAN there (documented divergence). %x* is always valid. CLAN CHECK 17.
Fixture: CHECK_017_undeclared_dep_tier.cha. Chatter regression codes: E605.
CHECK 18 — parity
An @ID tier whose speaker code is not declared on the @Participants tier (check.cpp:5062). E523 ‘@ID header for X but speaker not in @Participants’ is the direct equivalent. Note CLAN prints the literal tier name ‘@ID’ in the message (quirk of check_mess case 18 printing utterance->speaker).
Fixture: CHECK_018_id_speaker_not_in_participants.cha. Chatter regression codes: E522, E523.
CHECK 19 — parity
A token (here the retrace code ‘[/]’) is immediately followed by a non-space, non-delimiter character with no space between (check.cpp:4590/4592); also fires for a time bullet glued to a word. Full CLAN message is two lines: ‘Illegal use of delimiter in a word.’ + ‘Or a SPACE should be added after it.(19)’. Grounded also with a bullet glued to a word (u019d.cha). chatter rejects it with E757 (CodeGluedToFollowingContent), by span-adjacency on the retrace span, which covers the marker brackets. Wild-data impact: zero kept files.
Fixture: CHECK_019_code_glued_to_word.cha. Chatter regression codes: E757.
CHECK 20 — divergence
Documented divergence (maintainer ruling, 2026-07-12): a depfile-membership check. A token undeclared in a corpus optional/global depfile is not universally invalid CHAT; chatter validates universal CHAT-format validity, never external-config membership (CHAT validity must never depend on any config file). chatter deliberately accepts.
Fixture: CHECK_020_undeclared_suffix_compound.cha. Chatter regression codes: none.
CHECK 21 — parity
A main tier with text but no utterance terminator (check.cpp:4921; tier UTD==1, no delimiter found). E305 is chatter’s missing-terminator error.
Fixture: CHECK_021_missing_terminator.cha. Chatter regression codes: E305.
CHECK 22 — parity
An unclosed replacement scope remains invalid. The grammar rejects the fixture with E316; no raw-text replacement scanner or dedicated E311 diagnostic is required. This witnesses the unclosed replacement shape, not every CHECK 22 emission site.
Fixture: CHECK_022_unclosed_replacement.cha. Chatter regression codes: E316.
CHECK 23 — parity
Unmatched closing bracket is rejected by grammar recovery (E316); no bracket interpretation is reconstructed from recovery text.
Fixture: CHECK_023_unmatched_close_bracket.cha. Chatter regression codes: E316.
CHECK 24 — parity
Unmatched < on a main tier (open angle, no close). CLAN 24 (also 160); chatter surfaces the recovery node as E316. CLAN fires reliably here; the grounding helper guards against a depfile-load miss.
Fixture: CHECK_024_unmatched_lt.cha. Chatter regression codes: E316.
CHECK 25 — parity
Unmatched > on a main tier. CLAN 25; chatter surfaces the tree-sitter recovery node as E316 UnparsableContent.
Fixture: CHECK_025_unmatched_gt.cha. Chatter regression codes: E316.
CHECK 26 — parity
Unmatched { on a main tier. CLAN 26 (also 11, 48); chatter E316.
Fixture: CHECK_026_unmatched_open_brace.cha. Chatter regression codes: E316.
CHECK 27 — parity
Unmatched } on a main tier. CLAN 27 (also 48); chatter surfaces the recovery node as E316 UnparsableContent.
Fixture: CHECK_027_unmatched_close_brace.cha. Chatter regression codes: E316.
CHECK 28 — no_obligation
Call site commented out since 11-02-98 (check.cpp 4992, inside the dated comment block); unix and GUI builds alike can never emit it. Unbalanced parentheses surface through other codes today.
Exclusion reason: CommentedOut.
CHECK 29 — no_obligation
Call site commented out since 11-02-98 (check.cpp 4993, same comment block as CHECK 28); never emitted by any current build.
Exclusion reason: CommentedOut.
CHECK 30 — no_obligation
Compiled but unreachable in file mode: firing requires check_MatchTierName to return a colon-less declared header (ts->col FALSE) for a line that ALSO carries trailing text, but the match (check.cpp 1706, partcmp pattern mode) compares the WHOLE trimmed line against the declared name, so any trailing text makes the lookup fail first (empirically: ‘@Blank extra’, ‘@New Episode extra’, and tab-separated variants all emit (17) tier-not-declared, plus (132) for the tab forms, never (30)). No stock-depfile colon-less header contains a pattern wildcard that could match a longer line.
Exclusion reason: UnreachableInFileMode.
CHECK 31 — parity
A declared tier with a colon but no text after it (check.cpp:1748); ‘@’ headers are exempt (‘to always allow empty header tiers’), so the fixture uses an empty dependent tier ‘%fac:’. chatter answers E756 ‘Empty dependent tier’, which states the rule CHECK 31 is about. A 2026-08-15 maintainer ruling made E756 cover every free-text tier. An empty ‘*CHI:’ main tier instead yields CLAN (21)+(70), not (31). Whitespace-only content is E756 on every free-text tier, matching check_FoundText (check.cpp:1639), which skips ASCII space, tab and newline before looking for text. Two deliberate divergences, chatter stricter in both: CHECK exempts user-defined %x tiers from 31 and chatter reports them; chatter’s whitespace is Unicode whitespace, so a payload of non-breaking spaces alone is E756 while CHECK (verified 2026-09-08) reports nothing.
Fixture: CHECK_031_missing_text_after_colon.cha. Chatter regression codes: E756.
CHECK 32 — divergence
Documented divergence (maintainer ruling, 2026-07-12): a depfile-membership check. A token undeclared in a corpus optional/global depfile is not universally invalid CHAT; chatter validates universal CHAT-format validity, never external-config membership (CHAT validity must never depend on any config file). chatter deliberately accepts.
Fixture: CHECK_032_undeclared_dollar_code.cha. Chatter regression codes: none.
CHECK 33 — no_obligation
Compiled but unreachable with the stock depfile: check_matchplate returns 33 only when a tier’s declared pattern list mixes an @d/@t date-time pattern with OTHER patterns after it and the value fails everything (check.cpp 1543); when the date-time pattern is last, the specific codes 34/35 return instead. Every stock depfile @d/@t list is date-time-only (@Date, @Birth of #, @Time Duration, %tim), so failures yield (34)/(35), never (33): empirically @Date/@Birth/@ID garbage all emit (34). Reachable only via a custom user depadd.
Exclusion reason: UnreachableInFileMode.
CHECK 34 — parity
@Date is not a valid date (‘notadate’). CLAN 34 (illegal date representation); chatter E518 InvalidDateFormat.
Fixture: CHECK_034_bad_date.cha. Chatter regression codes: E518.
CHECK 35 — parity
@Time Duration 99:99:99 is shape-valid (HH:MM:SS) but out of clock range. chatter validates the shape, and TimeDurationValue::has_out_of_range_component flags components outside hours 0-23 / minutes,seconds 0-59 and emits E540. CLAN range rule grounded empirically: 25:00:00, 99:30:00, 12:60:00, 12:00:60 all CLAN 35; 23:59:59 OK. CLAN CHECK 35.
Fixture: CHECK_035_time_duration_out_of_range.cha. Chatter regression codes: E540.
CHECK 36 — parity
Utterance-delimiter followed by more material on the main tier (check.cpp:4394/4396; DelFound set, then a non-bullet, non-[ token appears). E316 (unparsable/misplaced content) rejects the same shape.
Fixture: CHECK_036_delimiter_not_at_end.cha. Chatter regression codes: E316.
CHECK 37 — parity
INCIDENTAL parity, reached by chatter’s own reasoning rather than CHECK’s. The 2026-07-12 divergence ruling still governs the RULE: CHECK 37 is a depfile-membership check, a token undeclared in a corpus depfile is not universally invalid CHAT, and chatter never validates external-config membership. The parity comes from this FIXTURE. Its word a+un#go carries a word-internal prefix marker in an English utterance, and E763 rejects the marker in any language that does not use it (heb, ara). Same verdict as CHECK on this input, different and independent grounds. A depfile-undeclared prefix in Hebrew or Arabic would still be accepted, which is the divergence the 2026-07-12 ruling protects.
Fixture: CHECK_037_undeclared_prefix.cha. Chatter regression codes: E763.
CHECK 38 — parity
A word starting with a digit on a main tier (excluding ‘0’-forms) that matches no depfile number pattern (check.cpp:4811 via check_matchnumber). E220 is chatter’s digits-in-word error. CLAN also emits (47) on the same fixture.
Fixture: CHECK_038_digits_in_utterance.cha. Chatter regression codes: E220.
CHECK 39 — no_obligation
No emission path anywhere in clan/ or lib/ (whole-source sweep 2026-07-24 covering check_err, the check_trans_err macro, and helper return-code channels); only the case-39 printer in the check_mess table exists. Tier-ordering violations surface through other codes today.
Exclusion reason: NoEmissionPath.
CHECK 40 — no_obligation
No emission path anywhere in clan/ or lib/ (whole-source sweep 2026-07-24, same method as code 39). chatter models per-tier duplication rules on its own terms where they make sense.
Exclusion reason: NoEmissionPath.
CHECK 41 — no_obligation
No emission path anywhere in clan/ or lib/ (whole-source sweep 2026-07-24). Parenthesis misuse surfaces via codes 55/111 (grounded elsewhere in this manifest); modern CHAT parenthetical shortenings are a parsed construct in chatter.
Exclusion reason: NoEmissionPath.
CHECK 42 — parity
Malformed ampersand and parenthesis input is rejected by grammar recovery (E316), without reconstructing annotation syntax.
Fixture: CHECK_042_amp_and_parens.cha. Chatter regression codes: E316.
CHECK 43 — parity
No @Begin header. chatter rejects with E504 (Missing required @Begin header in file preamble). Grounded vs real CLAN CHECK 43; audit name-matcher had marked it a gap.
Fixture: CHECK_043_missing_begin.cha. Chatter regression codes: E504.
CHECK 44 — parity
Speaker content after the @End tier. CLAN 44 (file must end with @End); chatter flags E502 MissingEndHeader (and E501). Both reject the file.
Fixture: CHECK_044_content_after_end.cha. Chatter regression codes: E502.
CHECK 45 — parity
At end of file, one or more @Bg tiers were never closed by a matching @Eg (check.cpp:5758); CHECK lists the unmatched @Bg contents after the message. E526 is chatter’s unmatched-@Bg error.
Fixture: CHECK_045_bg_without_eg.cha. Chatter regression codes: E526.
CHECK 46 — parity
An @Eg tier with no open @Bg carrying the same text (check.cpp:5391/5400). E527 is chatter’s unmatched-@Eg error.
Fixture: CHECK_046_eg_without_bg.cha. Chatter regression codes: E527.
CHECK 47 — parity
A digit inside a word (he11o). CLAN 47 (numbers not allowed inside words); chatter E220. This fixture agrees, not every scoped-language case: E220 examples 7/8 retain a paired enclosing-language substitution. CHECK 21-Sep-2026 rejects both with 47, while Chatter accepts the zho-span control under its established governing-language policy and rejects the eng-span mutation. Direct mixed/ambiguous word markers in E220 examples 3-6 agree with CHECK.
Fixture: CHECK_047_number_in_word.cha. Chatter regression codes: E220.
CHECK 48 — parity
A standalone illegal character on a main tier; ‘|’ between words reaches CHECK 48. The original bare-pipe acceptance gap was closed 2026-07-16 by E243. CLAN also emits 11 on that fixture. The separately grounded semicolon shape is covered by E769 specs: current HTML CHAT disallows main-tier semicolons, and CHECK 21-Sep-2026 rejects top-level, nested, and retraced cases with 48. Other CHECK-48 shapes remain adjudicated per-shape; this row is not a blanket claim for every call site.
Fixture: CHECK_048_illegal_char_pipe.cha. Chatter regression codes: E243.
CHECK 49 — no_obligation
Call site commented out 2019-04-23 (‘do not check for upper case letters’, check.cpp 3447 inside that comment block); never emitted by current builds. chatter’s own capitalization rules live in the validator, adjudicated separately.
Exclusion reason: CommentedOut.
CHECK 50 — parity
A second utterance delimiter after one was already found on a main tier (check.cpp:4351/4354/4554). E316 rejects ‘hello . .’.
Fixture: CHECK_050_redundant_delimiter.cha. Chatter regression codes: E316.
CHECK 51 — parity
CLAN prints this message WITHOUT a (51) suffix (format string lacks (%d)); grounding matches by message text. A ‘<…>’ angle group not followed by a ‘[…]’ code, detected at end of tier or at the next word (check.cpp:4376/4568, anb flag from check_CheckBrakets). QUIRK: CLAN prints this message WITHOUT the numeric ‘(51)’ suffix; check_mess case 51 (check.cpp) has no (%d) in its format string. Grounding is the verbatim unique message text. chatter names the same construct as a missing required element: the grammar requires a code after the group, tree-sitter inserts a MISSING placeholder for it, and the whole-tree pass reports E342. The annotation decoder skips MISSING nodes, so the group stays bare and E342 is the whole verdict (a decoder that read the placeholder’s kind would build a Full retrace nobody wrote and add E370 to a construct the file does not contain).
Fixture: CHECK_051_angle_without_square.cha. Chatter regression codes: E342.
CHECK 52 — parity
A postfix retrace has no preceding material. The grammar rejects it and source-bound recovery reports E316. Dedicated CHECK-style classification adds no validity enforcement.
Fixture: CHECK_052_retrace_without_text.cha. Chatter regression codes: E316.
CHECK 53 — parity
The duplicate-begin fixture is rejected by grammar recovery (E316); duplicate-header semantics remain on typed header validation.
Fixture: CHECK_053_duplicate_begin.cha. Chatter regression codes: E316.
CHECK 54 — no_obligation
The duplicate-end fixture is rejected by grammar recovery (E316); a recovery prefix alone does not prove a second parsed end header.
Fixture: CHECK_054_duplicate_end.cha. Chatter regression codes: E316.
Exclusion reason: UnreachableInFileMode.
CHECK 55 — parity
More ‘(’ than ‘)’ inside a single word (check.cpp:4571, pb>lpb after word scan). chatter rejects ‘ab(c d .’ via unparsable-content/terminator errors. E305 REMOVED 2026-08-11: it was a spurious “missing terminator” on a line that ends with one, caused by the lowering discarding the tier_body an ERROR node displaced. See docs/audits/2026-08-11-utterance-initial-annotation-adjudication.md.
Fixture: CHECK_055_unmatched_open_paren_word.cha. Chatter regression codes: E316.
CHECK 56 — parity
More ‘)’ than ‘(’ inside a single word (check.cpp:4572, pb<lpb). chatter rejects ‘ab)c d .’ via unparsable-content/terminator errors. E305 REMOVED 2026-08-11: it was a spurious “missing terminator” on a line that ends with one, caused by the lowering discarding the tier_body an ERROR node displaced. See docs/audits/2026-08-11-utterance-initial-annotation-adjudication.md.
Fixture: CHECK_056_unmatched_close_paren_word.cha. Chatter regression codes: E316.
CHECK 57 — parity
The word-glued pause fixture is rejected with E751. The grammar retains a word and a separate pause; model validation check_pause_glued_to_word compares their source spans in the typed content walk (spec E751). This is structural evidence, not a scan or reparse of recovery text. Under the architecture-native assessment contract, rejection stage and diagnostic identity are not parity obligations. CLAN also emits 48 on the retained fixture; no new CLAN observation is claimed.
Fixture: CHECK_057_pause_glued_to_word.cha. Chatter regression codes: E751.
CHECK 58 — no_obligation
No emission path anywhere in clan/ or lib/ (whole-source sweep 2026-07-24). The 8-character tier-name limit is a retired-era constraint; chatter’s tier-name rules are its own.
Exclusion reason: NoEmissionPath.
CHECK 59 — parity
A time bullet (\x15) opened but never closed before end of scan; hidenc stays TRUE at word end (check.cpp:4049/4797). Fixture: ‘%pho:\thello \x15123’ with no closing \x15. QUIRK: on the unix build CLAN prints the error banner with an EMPTY message line: check_mess case 59’s printf is inside ‘#if _MAC_CODE/#elif _WIN32’, so no (59) trailer and no message text ever appear. Grounded banner_only: run restricted to ‘+e59’ (block prints) and cross-checked ‘+e1’ (clean), re-confirmed 2026-07-09.
Fixture: CHECK_059_unclosed_bullet.cha. Chatter regression codes: E316.
CHECK 60 — parity
File has no @ID header. CLAN 60; chatter E522 (missing @ID header).
Fixture: CHECK_060_missing_id.cha. Chatter regression codes: E522.
CHECK 61 — parity
No @Participants header. chatter rejects with E504 (Missing required @Participants header), plus E543/E522. Grounded vs real CLAN CHECK 61.
Fixture: CHECK_061_missing_participants.cha. Chatter regression codes: E504.
CHECK 62 — parity
An @ID tier with an empty language (first) field (check.cpp:5054; also 2536 for an empty ‘word@s:’ language marker and 5023). E510 is chatter’s empty-@ID-language error. CLAN prints (62) twice on the fixture.
Fixture: CHECK_062_missing_language_in_id.cha. Chatter regression codes: E510.
CHECK 63 — parity
Empty corpus field (2nd) in @ID. chatter E514 EmptyIDCorpus. The corpus field is a required CorpusName like its siblings, and the Validate trait flags empty. CLAN CHECK 63.
Fixture: CHECK_063_empty_corpus.cha. Chatter regression codes: E514.
CHECK 64 — parity
@ID sex field is neither ‘male’ nor ‘female’ (‘x’). CLAN 64; chatter E542 UnsupportedSex. Requires CLAN’s depfile loaded (auto-resolved from the lib dir).
Fixture: CHECK_064_bad_gender.cha. Chatter regression codes: E542.
CHECK 65 — parity
A trailing @ in hello@ remains rejected by grammar recovery with E316. This witnesses one of CHECK 65’s 31 emission sites, not every special-form or morphology-marker case. E202 remains independently available at the model-validation boundary.
Fixture: CHECK_065_trailing_at_sign.cha. Chatter regression codes: E316.
CHECK 66 — parity
The ampersand-inside-word fixture is rejected by grammar recovery (E316), without guessing a word annotation or repair.
Fixture: CHECK_066_amp_inside_word.cha. Chatter regression codes: E316.
CHECK 67 — parity
A word-final ‘+’ (or ‘-’) followed by end-of-word / certain punctuation must be followed by text (check.cpp 3319/3338; the CA-char site 3025 is inside #ifndef UNX). Grounded via ‘hey+ .’; CLAN message spans two output lines, recorded with embedded newline.
Fixture: CHECK_067_trailing_plus.cha. Chatter regression codes: E233.
CHECK 68 — divergence
Deliberate divergence, resolved July 8: CHECK’s non-default +g2 option requires a CHI Target_Child participant unless @Options: notarget is present. Default-mode CHECK accepts the retained fixture, and Chatter intentionally accepts it. This option-specific requirement is not a baseline CHAT obligation. The open-policy question is resolved; clan_flags preserves the grounding configuration.
Fixture: CHECK_068_no_target_child.cha. Chatter regression codes: none.
CHECK 69 — parity
File has no @UTF8 header. CLAN 69; chatter E503 MissingUTF8Header.
Fixture: CHECK_069_missing_utf8.cha. Chatter regression codes: E503.
CHECK 70 — parity
A speaker tier containing no text word (only delimiters/pauses/brackets and not the ‘0’ placeholder) at tier end (check.cpp 4984, isTextFound never set). Fixture ‘*CHI:\t.’; chatter rejects with three codes including E306 (no meaningful content). SCOPE (2026-07-24, empirically verified live): CLAN 70 does NOT fire on pause-only ((.) .) or event-only (&=laughs .) utterances; both validators ACCEPT those shapes, so the old ‘pause-only utterance’ policy-gap framing was wrong.
Fixture: CHECK_070_no_text.cha. Chatter regression codes: E253, E306, E342.
CHECK 70 — parity
An annotation-only utterance ([=! laughs] .) is rejected by structural recovery, E316 and E342. CHECK grounding emits 73 and 70. Pause-only and event-only utterances were separately accepted by both validators; this case does not establish their invalidity. No missing-terminator or missing-whitespace claim is inferred.
Fixture: CHECK_070_annotation_only_utterance.cha. Chatter regression codes: E316, E342.
CHECK 71 — parity
INCIDENTAL parity, reached by chatter’s own reasoning rather than CHECK’s. The 2026-07-12 divergence ruling still governs the RULE: CHECK 71 enforces an ORDERING between [/] and the legacy # pause marker, a construct chatter does not have (chatter uses (.)/(..)/(…)), so chatter does not implement that ordering rule and never will. The parity comes from this FIXTURE. Its main tier is he # [/] he ., where the standalone # sits inside the retrace scope, and E762 rejects a word that is nothing but the prefix marker in any language. Same verdict as CHECK on this input, unrelated grounds. The fixture fails because word validation recurses into retraces and validates the # inside the retrace.
Fixture: CHECK_071_retrace_after_pause.cha. Chatter regression codes: E762.
CHECK 72 — parity
A pause ‘(…)’ or non-postcode ‘[…]’ item appearing after the utterance delimiter (check.cpp 4740/4744). Fixture ‘hey . (.)’; chatter rejects (unparsable content after terminator).
Fixture: CHECK_072_pause_after_delim.cha. Chatter regression codes: E316.
CHECK 73 — parity
A bracketed code as the first item of a tier with nothing before it (also bullets not preceded by text; multiple sites: check.cpp 4045, 4781, 4861; CA sites 3030/3033 are #ifndef UNX). Fixture ‘[x 2] hey .’; chatter rejects it via E316 (grammar recovery). CLAN additionally emits (11). (E747 is BlankLineNotAllowed and is not the expectation here.) Additional leading-bullet grounding is owned by spec/errors/E770.md: simple, retraced-group, linker-prefixed and repeated leading bullets report E770; explicit-zero, event, pause and preceding-group-material controls do not. CHECK 21-Sep-2026 corroborates these shapes. Parse-recovered main tiers retain their parse diagnostic rather than an E770 absence claim. This row’s bracket fixture alone does not prove every CHECK 73 site: empty inter-bullet scope and legacy multi-option behavior are not adjudicated by this fixture. They remain outside this finite baseline, not demonstrated validation defects or queued diagnostic-emulation work. A concrete independently justified CHAT requirement is needed to open a new obligation.
Fixture: CHECK_073_bracket_first.cha. Chatter regression codes: E316.
CHECK 74 — no_obligation
Sole call site commented out on 2004-04-27 (check.cpp 1762, ‘// 2004-04-27 if (check_err(74,…))’): deliberately retired 22 years ago; no build can emit it. chatter’s tier-separator rules (single :\t) are enforced by the grammar.
Exclusion reason: CommentedOut.
CHECK 75 — parity
A postcode ‘[+ …]’ occurring before the utterance delimiter (check.cpp 4747). Fixture ‘hey [+ exc] .’; chatter rejects (postcode only legal after terminator).
Fixture: CHECK_075_postcode_before_delim.cha. Chatter regression codes: E305, E316.
CHECK 76 — no_obligation
An ‘@l’ letter-marker word whose stem is more than one letter. RETIRED ON BOTH SIDES. CLAN retired error 76 on 2026-08-07 (“CHECK: allow multiple characters before @l”); check_isOneLetter and its call site sit inside a dated comment block, so no unix build emits it. chatter retired its own E754 on 2026-08-11, independently and for a stated reason: the rule counted CHARACTERS, a digraph is one letter written with two (Welsh ‘ll’, Dutch ‘ij’), so it reported a spurious error on valid CHAT and could not do better because it never saw the word’s language. Both validators accept the construct, so expected_chatter_codes is empty and the fixture pins that agreement.
Fixture: CHECK_076_multi_letter_at_l.cha. Chatter regression codes: none.
Exclusion reason: CommentedOut.
CHECK 77 — parity
File has no @Languages header. CLAN 77 (also 61); chatter E504 MissingRequiredHeader. chatter uses one generic missing-required-header code where CLAN has per-header codes; parity is on validity.
Fixture: CHECK_077_missing_languages.cha. Chatter regression codes: E504.
CHECK 78 — parity
A special delimiter word ‘+,’ ‘+“’ ‘+^’ ‘+<’ ‘++’ (exactly two chars) appearing after the first word of the tier (check.cpp 4769). Fixture ‘hey +, you .’; chatter rejects (unparsable ’ +’). Expected codes refreshed 2026-07-30 (E316 -> E766): the misplaced-linker grammar change names this construct E766 (linker not utterance-initial), the direct counterpart of CHECK 78’s message.
Fixture: CHECK_078_pluscomma_mid.cha. Chatter regression codes: E766.
CHECK 79 — parity
A %mor word with two | separators (n|dog|cat). CLAN 79 (only one | per word); chatter surfaces the malformed %mor item as E702 (a recovery node below the tier’s direct children is classified with the tier’s name).
Fixture: CHECK_079_mor_two_pipes.cha. Chatter regression codes: E702.
CHECK 80 — parity
A %mor word with no | separator (bare ‘dog’). CLAN 80 (must be at least one |); chatter E702.
Fixture: CHECK_080_mor_no_pipe.cha. Chatter regression codes: E702.
CHECK 81 — no_obligation
Call site commented out 2025-07-04 (check.cpp 4791, dated comment block); bullet-placement validity is covered by the still-live bullet codes (73, 118 was also retired the same day).
Exclusion reason: CommentedOut.
CHECK 82 — parity
A media bullet whose END time is <= its BEG time (check.cpp 3879, new-format path with @Media header). Bullet \x15500_100\x15. Note @Media file name must match the .cha basename or CLAN adds (157) noise.
Fixture: CHECK_082_end_before_beg.cha. Chatter regression codes: E362.
CHECK 83 — parity
A tier’s first bullet BEG earlier than the previous tier’s first-bullet BEG (check_tierBegTime), reachable only from the third bulleted tier on because tEndTime==0 skips the block on the first (check.cpp 3885/3890). Needs three tiers (1000_2000, 3000_4000, then 2500_2600): with only two, the tracking variable is still 0 and no error fires; a two-tier same-speaker attempt yields (133) instead. CLAN message has no trailing period.
Fixture: CHECK_083_beg_backwards.cha. Chatter regression codes: E362.
CHECK 84 — divergence
Deliberate divergence, resolved July 8: CHECK +c0 rejects a current tier’s begin time before the previous tier’s end (check.cpp:3901). Default-mode CHECK accepts the retained fixture, and Chatter intentionally permits conversational overlap. This cross-tier option-specific restriction does not replace Chatter’s independent same-speaker timing rules. clan_flags preserves the grounding configuration.
Fixture: CHECK_084_overlap.cha. Chatter regression codes: none.
CHECK 85 — divergence
Deliberate divergence, resolved July 8: CHECK +c0 rejects a gap between successive tiers’ bullets (check.cpp:3905). Default-mode CHECK accepts the retained fixture, and Chatter intentionally permits gaps. Requiring continuous coverage of the recording is not a baseline CHAT obligation. clan_flags preserves the grounding configuration.
Fixture: CHECK_085_gap.cha. Chatter regression codes: none.
CHECK 86 — parity
Shape-scoped agreement: U+E000 inside a word produces CHECK 86 and E243. Chatter independently rejects Unicode private-use scalars and noncharacters in lexical words, including supplementary planes, with no CLAN-internal exemption. Ordinary assigned high-BMP characters and U+10000 are not rejected by this policy. The recorded 21-Sep-2026 CHECK U+10000 rejection remains an intentional divergence. Canonical E243_unicode_boundaries examples cover all noncharacters, private-use endpoints, former exemption endpoints and valid controls; this row does not claim universal Unicode parity or new CHECK observations.
Fixture: CHECK_086_private_use_char.cha. Chatter regression codes: E243.
CHECK 87 — parity
A %mor/%xmor/%trn (or %cnl) word whose compound structure is malformed: a ‘+’ not in canonical ‘pfx|+a|b’ shape drives numOfCompounds to -2 or 0 (check.cpp 3592/3646). Fixture ‘%mor: n|cat+n|fish .’ (compound missing the ‘|+’ marker after the head POS). chatter rejects: E702 on the item and E600 at the tier.
Fixture: CHECK_087_malformed_compound.cha. Chatter regression codes: E702.
CHECK 88 — no_obligation
Call site commented out since 2009-06-08 (check.cpp 3414, dated comment block); never emitted by current builds. chatter’s compound/form-marker rules are adjudicated per live CHECK codes 65/66 instead.
Exclusion reason: CommentedOut.
CHECK 89 — parity
A new-format bullet whose content is not digits_digits (e.g. wrong separator) (check.cpp 3863, res==2 from check_getMediaTagInfo). Bullet \x15100x200\x15 with @Media header. chatter rejects.
Fixture: CHECK_089_bad_bullet_char.cha. Chatter regression codes: E316, E544.
CHECK 90 — parity
A bullet time component written with a leading zero before another digit, e.g. 012 (check.cpp check_getMediaTagInfo res==3, call site 3865). Fixture: \x15012_200\x15 with @Media header. chatter emits E748 at parse time from the raw component text (both bullet paths in talkbank-parser media_bullet.rs; token scan in talkbank-parser-re2c file.rs). Spec: E748.md. A bare 0 stays legal. ADJUDICATED (maintainer decision, 2026-07-10): STYLE error. Leading-zero bullet times are unambiguous; kept as an error by project policy: style violations are errors.
Fixture: CHECK_090_leading_zero_time.cha. Chatter regression codes: E748.
CHECK 91 — parity
A blank line in the transcript (CLAN flags one between utterances AND between headers, verified). CLAN CHECK 91. The grammar’s newline is a single line break (/\r\n|[\r\n]/) and a blank_line CST node represents the blank line; the parser emits E747 BlankLineNotAllowed from that node (helpers.rs line dispatch), with no source or line scan. (A fused newline: /[\r\n]+/ would erase the blank line before any node could represent it.) chatter rejects via E747.
Fixture: CHECK_091_blank_line.cha. Chatter regression codes: E747.
CHECK 92 — parity
Comma not followed by space or end-of-line on a speaker tier (check.cpp 4309-4320, with CA-mark exemptions). GAP CLOSED 2026-07-09: model validation check_comma_glued_to_next (span adjacency over walk_content; dummy spans skipped) plus a re2c token-stream mirror. Spec: E749.md. Narrowed to word-next: constructs placing their own character after the comma are exempt, matching CLAN. ADJUDICATED (maintainer decision, 2026-07-10): STYLE error. Comma-glue is unambiguous under our grammar; kept as an error by project policy: style violations are errors.
Fixture: CHECK_092_comma_no_space.cha. Chatter regression codes: E749.
CHECK 93 — parity
‘[’ immediately preceded by ‘>’ or ‘]’ with no space (check.cpp 4469/4472; CA start-marks site 3179 is #ifndef UNX). Fixture <hey>[/] hey .; CLAN also emits (161). chatter rejects via E316. (E747 is BlankLineNotAllowed and is not the expectation here.)
Fixture: CHECK_093_bracket_no_space.cha. Chatter regression codes: E316.
CHECK 94 — parity
%mor tier whose utterance delimiter differs from the speaker tier’s recorded delimiter (check.cpp 4338 simple delimiters; 4522 ‘+…’-style). ‘*CHI: hey .’ with ‘%mor: co|hey ?’. chatter rejects with the dedicated E716.
Fixture: CHECK_094_mor_delim_mismatch.cha. Chatter regression codes: E716.
CHECK 95 — no_obligation
Call site commented out (check.cpp 3300, undated comment block guarding check_err(95) behind the retired capWFound flag, itself part of the 2019-04-23 upper-case retirement); never emitted by current builds.
Exclusion reason: CommentedOut.
CHECK 96 — no_obligation
Compiled but unreachable: the emit site (check.cpp 3462) is the @Languages branch of check_CheckWords, but @Languages values are validated by check_isLangMatch (check.cpp 2165), which deliberately SPLITS each token at ‘=’ and validates only the language-code part, accepting the legacy word-color suffix; the word-level walker never visits @Languages in file mode. Empirically ‘@Languages: eng=blue’ and ‘eng, fra=blue’ pass unix CHECK clean. chatter rejects the construct as malformed header content (E320), correctly stricter.
Fixture: CHECK_096_word_color_languages.cha. Chatter regression codes: E320.
Exclusion reason: UnreachableInFileMode.
CHECK 97 — parity
‘@’ encountered inside parentheses within a word on a speaker tier (check.cpp 3360, paransFound). Fixture ‘c(a@b)t .’; CLAN also emits (147). Chatter rejects the structural error with E316; recovery does not establish an independent form-marker node for E203.
Fixture: CHECK_097_at_in_parens.cha. Chatter regression codes: E316.
CHECK 98 — parity
An old-format bullet (no @Media header) ‘\x15%snd:“NAME”_B_E\x15’ whose media file name contains a space (check.cpp 3923 via res==4; also reachable in the new-format path 3867). chatter rejects the old-style %snd bullet outright (E360), so anything CHECK flags inside it is also rejected; parity by strictness.
Fixture: CHECK_098_space_in_name.cha. Chatter regression codes: E360.
CHECK 99 — parity
Media file name ending with an extension: in old-format bullets (check.cpp 3925 via res==5, used here), in new-format path (3869), or on the @Media header (5589). Grounded via old-format bullet ‘\x15%snd:“test.wav”_100_200\x15’; chatter rejects the %snd bullet (E360). The @Media-header variant of (99) was not separately grounded.
Fixture: CHECK_099_extension_in_name.cha. Chatter regression codes: E360.
CHECK 100 — parity
Trailing comma at end of @Participants. The grammar already surfaces the dangling comma as a tree-sitter ERROR node inside the participants header; the parser reports it as E550 TrailingCommaInParticipants (parse-stage diagnostic). CLAN CHECK 100.
Fixture: CHECK_100_trailing_comma_participants.cha. Chatter regression codes: E550.
CHECK 101 — no_obligation
GUI-only: the sole call site (check.cpp 3016) sits inside the #ifndef UNX CA-character region of check_CheckWords, so unifdef -DUNX blanks it and no unix binary contains it. DERIVABLE: the generator runs unifdef before scanning, so the reference reports zero live sites and this reason is checked rather than trusted.
Fixture: CHECK_101_ca_marker_no_stem.cha. Chatter regression codes: E316.
Exclusion reason: GuiOnly.
CHECK 102 — parity
Any character carrying the italic text attribute, encoded in the file as ATTMARKER 0x02 + 0x03 (italic start) / 0x02 + 0x04 (italic end) (check.cpp: multiple is_italic() sites, e.g. 4289/4433). Fixture embeds raw bytes \x02\x03hey\x02\x04; CLAN emits (102) twice (once per attributed char run). chatter names each of the four control bytes as E315 (the lexical rule permits only the underline pairs 0x02 0x01 / 0x02 0x02) and the broken word as E316, so legacy attribute encoding cannot pass.
Fixture: CHECK_102_italics.cha. Chatter regression codes: E315, E316.
CHECK 103 — parity
@Options header (which must immediately follow @Participants) listing both ‘CA’ and ‘IPA’ (check.cpp 5478/5537). First attempt placed @Options after @ID and got first-pass (125) instead; @Options must sit right after @Participants. CLAN also emits (11) because IPA is not in the depfile @Options template. chatter rejects with E534.
Fixture: CHECK_103_ca_and_ipa.cha. Chatter regression codes: E534.
CHECK 104 — no_obligation
Both call sites commented out (check.cpp 5483/5542, inside comment blocks); the font-selection advice is GUI-editor-era and never emitted by current builds. Fonts are not a CHAT-validity property; chatter has no analogue by design.
Exclusion reason: CommentedOut.
CHECK 105 — no_obligation
Both call sites commented out (check.cpp 5487/5546, same comment blocks as CHECK 104); never emitted. Same adjudication as 104: font choice is not validity.
Exclusion reason: CommentedOut.
CHECK 106 — parity
A scoped code split across a newline ([% comment then a continuation line continued]) leaves an ERROR node in the main-tier content, which the whole-tree recovery-node backstop surfaces as E316. CLAN CHECK 106.
Fixture: CHECK_106_code_spans_newline.cha. Chatter regression codes: E316.
CHECK 107 — parity
Three or more consecutive commas on a speaker tier (check.cpp 4322; two commas give (156) instead). Fixture ‘hey ,,, you .’; CLAN also emits (156). chatter rejects (E258 twice).
Fixture: CHECK_107_triple_comma.cha. Chatter regression codes: E258.
CHECK 108 — parity
A postcode [+ trn] placed after the final time bullet leaves an ERROR node in the utterance, which the whole-tree recovery-node backstop surfaces as E316. CLAN rejects it (codes 108 and 112). CLAN CHECK 108 (co-emitted with 112).
Fixture: CHECK_108_postcode_after_bullet.cha. Chatter regression codes: E316.
CHECK 109 — divergence
A postcode-shaped token [+ trn] on a dependent tier (%com). CLAN CHECK 109 (“Postcodes are not allowed on dependent tiers”, check_CheckWords, check.cpp:3471-3690) flags it on any non-%x dependent tier. chatter does not model postcodes on dependent tiers: ordinary dependent tiers like %com parse as TextTier with no Postcode AST node, so there is nothing typed to detect, and flagging it would require a banned raw-text scan. Empirically (2026-06-25) CLAN’s own analysis tools do not choke on dependent-tier postcodes: FREQ and MLU exclude the [+ ...] token from counts exactly as on the main tier, KWAL displays the line; only CHECK flags it. CHECK 109 is a CLAN-internal pedantic check, not a CHAT-validity rule, so chatter deliberately diverges (accepts) per the parity mandate’s ‘unless CHECK is buggy / not a real validity rule’ clause. Divergence (chatter intentionally looser), not a gap.
Fixture: CHECK_109_postcode_on_deptier.cha. Chatter regression codes: none.
CHECK 110 — divergence
Deliberate divergence, resolved July 8: with non-default +cN, CHECK requires a bullet on every speaker tier (check.cpp:4915). Default-mode CHECK accepts the retained fixture, and Chatter intentionally accepts an untimed tier. Universal timing is not a baseline CHAT obligation. The retained observation uses +c1, recorded in clan_flags.
Fixture: CHECK_110_no_bullet.cha. Chatter regression codes: none.
CHECK 111 — parity
An illegal pause format: (5) (digits with no dot) instead of (.)/(..)/(…)/(n.). CLAN 111 (pause must have ‘.’); chatter rejects with E209 (also E220).
Fixture: CHECK_111_illegal_pause.cha. Chatter regression codes: E209.
CHECK 112 — parity
Timing without an @Media header is rejected by E752 (spec E752). The original CLI acceptance defect was closed July 15 under the independently justified media-declaration rule. Since September 8 the timing-evidence union includes bullets inside utterances as well as terminal bullets. CHECK also counts inline bullets; Chatter additionally counts %wor timing. These differences do not create an obligation to copy CHECK’s text scan.
Fixture: CHECK_112_missing_media.cha. Chatter regression codes: E752.
CHECK 113 — parity
A comma-separated field after the file name on @Media that is not a depfile-listed keyword (audio/video/unlinked/…). E536 is the same rule (Unsupported @Media status: ‘junkword’). E531 is extra chatter strictness: media filename ‘test’ vs fixture file name.
Fixture: CHECK_113_media_bad_keyword.cha. Chatter regression codes: E536.
CHECK 114 — parity
@Media names a media file but lists neither ‘audio’ nor ‘video’; CLAN emits (114) on the @Media line. FIXTURE REPAIRED 2026-07-09: the survey-batch copy was truncated mid-@Media-line (no utterance, no @End), so CLAN aborted first-pass with (7) and could never reach 114; the drift guard caught it. chatter rejects the keywordless @Media at parse time (E316 unparsable header line; an incidental E747 blank-line cascade fires on the newline the recovery leaves behind, deliberately not asserted).
Fixture: CHECK_114_media_no_keyword.cha. Chatter regression codes: E316.
CHECK 115 — parity
With an @Media header present, a time bullet whose content starts with %s/%m, i.e. the old ‘%snd:“file”_beg_end’ bullet format. Old-format bullet is unparsable to chatter (E316) and the file then has no timing info (E544). E531 is fixture-name noise.
Fixture: CHECK_115_old_bullet_format.cha. Chatter regression codes: E316, E544.
CHECK 116 — parity
A legacy per-line font directive is not a valid scoped comment. Grammar recovery rejects the fixture with E316; raw-text prefix classification is unnecessary.
Fixture: CHECK_116_fnt_code.cha. Chatter regression codes: E316.
CHECK 117 — parity
A pairwise CA marker (faster U+2206, slower, creaky, whisper, left-arrow-circle U+21AB, etc.) occurring an odd number of times on a main tier. Dedicated chatter code: E230 ‘Unbalanced CA delimiter (Faster): missing closing delimiter’. CLAN tracks 15 pair-wise marker types; fixture uses faster U+2206. This fixture’s parity does not imply identical replacement-scope behavior: E230 examples 2/3 retain Chatter’s once-per-authored-marker pairing. CHECK 21-Sep-2026 scans replacement text twice in check_ParseWords (whole bracket token, then target-word re-entry), cancelling a single target marker and inverting the paired/unpaired verdicts. See E230’s source-path analysis.
Fixture: CHECK_117_unmatched_faster.cha. Chatter regression codes: E230.
CHECK 118 — no_obligation
Call site commented out 2025-07-04 (check.cpp 4623, dated comment block, retired together with 81); never emitted by current builds. Delimiter-before-bullet ordering remains covered by the live bullet/delimiter codes.
Exclusion reason: CommentedOut.
CHECK 119 — parity
Retrace marker [/] with nothing after it. chatter E370 StructuralOrderError.
Fixture: CHECK_119_dangling_retrace.cha. Chatter regression codes: E370.
CHECK 120 — parity
A 2-letter (or 2-letter plus dash subtag) language code where a 3-letter ISO-639 code is expected, on @Languages, the @ID language field, [- lng] precodes, or @s codes. Dedicated chatter code: E519 “Language code ‘en’ should be 3 characters (got 2)”. Fixture uses en consistently on @Languages and @ID.
Fixture: CHECK_120_two_letter_lang.cha. Chatter regression codes: E519.
CHECK 121 — parity
An UNASSIGNED three-letter code (qzz, outside the qaa-qtz private-use range) on @Languages/@ID. CLAN emits (121) from its legacy language table; chatter rejects via the full ISO 639-3 registry (E519, InvalidLanguageCode), the authoritative modern standard. PARITY per the 2026-07-14 ruling; fixture re-grounded 2026-07-15 from qqq to qzz because qqq is registry-VALID (private use), which the original fixture conflated; see the companion private-use divergence entry.
Fixture: CHECK_121_unknown_lang.cha. Chatter regression codes: E519.
CHECK 121 — divergence
A PRIVATE-USE code (qqq, inside ISO 639-3’s reserved qaa-qtz range) on @Languages/@ID. CLAN’s fixed table lacks the private-use range and emits (121); the ISO 639-3 registry defines qaa-qtz as valid codes reserved for local use, so chatter accepts them: an intentional registry-grounded divergence, not a gap. Wild grounding 2026-07-15: no kept file uses a private-use code (typed scan of every data-json languages field), so the divergence is currently theoretical.
Fixture: CHECK_121_private_use_lang.cha. Chatter regression codes: none.
CHECK 122 — parity
@ID declares language ‘eng’ that is not present in the @Languages header (which declares ‘fra’). CLAN 122; chatter E519. NOTE: code 122 has n_call_sites=0 in the generated catalogue (emitted via an indirect call the literal check_err grep misses) but IS emittable: verified vs real file-mode CLAN. The 8-code ‘dead’ list therefore needs empirical re-verification.
Fixture: CHECK_122_id_lang_not_declared.cha. Chatter regression codes: E519.
CHECK 123 — parity
In a non-CA file, a space immediately after the speaker-prefix TAB is rejected by E758. The acceptance defect was closed July 16; @Options: CA permits alignment whitespace. The retained source analysis distinguishes the live check.cpp:1669 site from the unreachable 1673 branch. This row preserves that scoped decision, not a raw-text implementation requirement.
Fixture: CHECK_123_tab_space.cha. Chatter regression codes: E758.
CHECK 124 — parity
@Media declares unlinked but the main tier has a time bullet, so the media is in fact linked. chatter emits E552 MediaUnlinkedWithTiming via check_media_unlinked_has_no_timing, the inverse of the existing E544 (linkage without timing). Grounded empirically: unlinked+bullet -> CLAN 124; unlinked+no-bullet and linked+bullet -> no 124. CLAN CHECK 124.
Fixture: CHECK_124_media_unlinked_with_bullet.cha. Chatter regression codes: E552.
CHECK 125 — parity
@Options placed after @ID instead of immediately after @Participants. chatter E551 OptionsHeaderOutOfOrder, mirroring the @ID rule (E548/CHECK 126). CLAN CHECK 125.
Fixture: CHECK_125_options_after_id.cha. Chatter regression codes: E551.
CHECK 126 — parity
A changeable header (@Comment) sits between @Participants and the @ID block.
Fixture: CHECK_126_id_header_out_of_order.cha. Chatter regression codes: E548.
CHECK 127 — parity
A changeable header sits between the @ID block and @Birth of.
Fixture: CHECK_127_constant_header_out_of_order.cha. Chatter regression codes: E547.
CHECK 128 — parity
Group-start marker U+2039 opened on a main tier and never closed (gm > 0 at end of tier). No dedicated code; unmatched opener makes the tier unparsable (E316).
Fixture: CHECK_128_unmatched_group_start.cha. Chatter regression codes: E316.
CHECK 129 — parity
Group-end marker U+203A encountered with no open U+2039 (gm goes negative). Rejected at parse level: E316 unparsable content on the stray closer plus E305 missing terminator. Expected codes refreshed 2026-07-30 (E305, E316 -> E316): the 2026-07-29/30 stack (parse-taint cascade gate, finer recovery, specific codes) changed which codes fire; the fixture is still rejected.
Fixture: CHECK_129_unmatched_group_end.cha. Chatter regression codes: E316.
CHECK 130 — parity
Sign-group-start marker U+3014 opened and never closed (sgm > 0 at end of tier). No dedicated code; rejected as unparsable (E316).
Fixture: CHECK_130_unmatched_sign_group_start.cha. Chatter regression codes: E316.
CHECK 131 — parity
Sign-group-end marker U+3015 with no opener (sgm goes negative). Rejected at parse level (E316 + E305). Expected codes refreshed 2026-07-30 (E305, E316 -> E316): the 2026-07-29/30 stack (parse-taint cascade gate, finer recovery, specific codes) changed which codes fire; the fixture is still rejected.
Fixture: CHECK_131_unmatched_sign_group_end.cha. Chatter regression codes: E316.
CHECK 132 — parity
A tab in the middle of a main tier. chatter rejects it (E370, unexpected tab). Grounded vs real CLAN CHECK 132.
Fixture: CHECK_132_tab_midline.cha. Chatter regression codes: E370.
CHECK 133 — parity
With linked media, a bullet whose BEG time is more than 500 ms earlier than the same speaker’s previous bullet END time. Dedicated chatter code: E704 ‘Speaker CHI overlaps with self: utterance 1 ends at 5000ms…’. E531 is fixture-name noise. CLAN threshold is >500ms (check.cpp:3892-3896); chatter’s exact threshold not probed by this survey.
Fixture: CHECK_133_overlap_same_speaker.cha. Chatter regression codes: E704.
CHECK 134 — parity
A %mor/%trn (or %cnl) item of the form unk|xxx, unk|yyy, or *|www (mor output for untranscribed material). chatter rejects via alignment (E706: main tier has 0 alignable items, %mor has 1) because xxx is not %mor-alignable; there is no dedicated unk|xxx rule. A hypothetical fixture where counts still align was not probed.
Fixture: CHECK_134_mor_unk_xxx.cha. Chatter regression codes: E706.
CHECK 135 — parity
A main-tier word that is exactly ‘xx’ or ‘yy’ (case-insensitive); %cnl tiers also flag bare ‘gen’ and ‘^’. Dedicated chatter code: E241 ‘“xx” is not legal; did you mean to use “xxx”?’. Only the main-tier xx call site (check.cpp:3138) was grounded; the %cnl sites (3613, 3632) were not separately probed.
Fixture: CHECK_135_bare_xx.cha. Chatter regression codes: E241.
CHECK 136 — parity
Open curly quote U+201C on a tier with no matching close quote (qt > 0 at end of tier). E242 is decided from CST structure and names the unmatched OPEN quote as well as the close-quote case (CHECK 137).
Fixture: CHECK_136_unmatched_open_quote.cha. Chatter regression codes: E242.
CHECK 137 — parity
Close curly quote U+201D with no preceding open quote (qt goes negative). Dedicated chatter code: E242 ‘Unmatched curly quote ” found on the tier’. Expected codes refreshed 2026-07-30 (E242, E305 -> E242): the 2026-07-29/30 stack (parse-taint cascade gate, finer recovery, specific codes) changed which codes fire; the fixture is still rejected.
Fixture: CHECK_137_unmatched_close_quote.cha. Chatter regression codes: E242.
CHECK 138 — parity
A right single curly quote U+2019 inside a word. CLAN 138; chatter E256 IllegalCurlyQuote (both parsers reject U+2018/U+2019).
Fixture: CHECK_138_curly_u2019.cha. Chatter regression codes: E256.
CHECK 139 — parity
A left single curly quote U+2018 at the start of a word. CLAN 139; chatter E256 IllegalCurlyQuote.
Fixture: CHECK_139_curly_u2018.cha. Chatter regression codes: E256.
CHECK 140 — parity
%mor tier does not align item-for-item with its speaker tier (createMorUttline sets mor_link.error_found). Dedicated chatter code: E705 ‘Main tier has 2 alignable items, but %mor tier has 1 items’.
Fixture: CHECK_140_mor_size_mismatch.cha. Chatter regression codes: E705.
CHECK 141 — parity
A [: text] replacement preceded by a group ‘>’, another ‘]’, nothing, a pause, or an &/+ form instead of exactly one plain word. chatter’s grammar does not accept a replacement after a group, so it rejects at parse level (E316 on <big dog> [: doggy]); no dedicated code. Expected codes refreshed 2026-07-30 (E316, E747 -> E316): the 2026-07-29/30 stack (parse-taint cascade gate, finer recovery, specific codes) changed which codes fire; the fixture is still rejected.
Fixture: CHECK_141_replacement_after_group.cha. Chatter regression codes: E316.
CHECK 142 — parity
The role field (field 8) of @ID differs from the role declared for that speaker on @Participants. Dedicated chatter code: E532 “Speaker ‘CHI’ has role ‘Mother’ on @ID but ‘Target_Child’ on @Participants”.
Fixture: CHECK_142_id_role_mismatch.cha. Chatter regression codes: E532.
CHECK 143 — parity
@ID line whose pipe-separated field count is not exactly 10. chatter rejects via E342 (missing required ‘pipe’), a parse-level equivalent.
Fixture: CHECK_143_id_field_count.cha. Chatter regression codes: E342.
CHECK 144 — parity
@ID SES field is an undeclared value (‘XYZ’). CLAN 144 (illegal SES / not declared in depfile); chatter E546 UnsupportedSesValue. The parity claim covers unknown tokens, not identical separator policy. CHECK’s check_ses skips commas/whitespace and validates each remaining token against either vocabulary (check_matchplate/check_SES_item); it has no pair-completeness check. E546_vocabulary records acceptance of ‘,MC’, ‘White,’ and comma-only input by CHECK 21-Sep-2026, versus Chatter’s existing stricter complete-pair policy.
Fixture: CHECK_144_bad_ses.cha. Chatter regression codes: E546.
CHECK 145 — no_obligation
Call site commented out 2019-04-17 with the comment ‘leave this to CHATTER’ (check.cpp 4602): CLAN explicitly delegated this rule’s future to this project. Never emitted by current builds. If the paired-marker interaction rule is wanted, it is a fresh chatter rule to design, not a parity obligation.
Exclusion reason: CommentedOut.
CHECK 146 — parity
A main-tier word that is exactly ‘&=’ with no code after the equals sign. Rejected at parse level: E316 ’Unparsable content on main tier: &= ’. Expected codes refreshed 2026-07-30 (E305, E316 -> E316): the 2026-07-29/30 stack (parse-taint cascade gate, finer recovery, specific codes) changed which codes fire; the fixture is still rejected.
Fixture: CHECK_146_bare_amp_eq.cha. Chatter regression codes: E316.
CHECK 147 — parity
An undeclared special-form marker foo@zzz. CLAN CHECK 147. The grammar intentionally consumes @zzz as one form_marker token (parse-don’t-validate), and the conversion requires the @z: colon (the only valid user-defined form is @z:label); @zXXX without the colon falls through to the E203 InvalidFormType branch in both the tree-sitter and re2c conversions. chatter rejects it via E203.
Fixture: CHECK_147_undeclared_form_marker.cha. Chatter regression codes: E203.
CHECK 148 — parity
Space before the comma in @Media. Chatter has its own named rule for it, E767, so the header parses and the diagnostic names the space (a header that failed to parse would fall back to E525, unknown header type, about a header chatter had recognised, plus E330, missing media_type, on a line visibly ending in , audio). Both validators reject, and both say why.
Fixture: CHECK_148_media_space_before_comma.cha. Chatter regression codes: E767.
CHECK 149 — parity
A skip character other than space/comma/‘<’/‘>’/newline (e.g. an utterance terminator ‘.’) sitting between a word and a following […] code. Fixture ‘dog . [: doggy] .’ also draws CLAN (50) extra-delimiter noise; chatter rejects the same configuration as unparsable (E316 on ‘[: …’).
Fixture: CHECK_149_char_before_code.cha. Chatter regression codes: E316.
CHECK 150 — parity
A ‘)’ with no matching ‘(’ in the non-skip material scanned back from a […] code. chatter rejects ‘dog) [/] dog’ as unparsable (E316 on ‘) [/] dog’). E305 REMOVED 2026-08-11: it was a spurious “missing terminator” on a line that ends with one, caused by the lowering discarding the tier_body an ERROR node displaced. See docs/audits/2026-08-11-utterance-initial-annotation-adjudication.md.
Fixture: CHECK_150_paren_before_code.cha. Chatter regression codes: E316.
CHECK 151 — no_obligation
Compiled into the unix build but never reached. Not gui_only: the check_err(151,…) site is at check.cpp 2980, inside helper check_isThereStem, which ends before the #ifndef UNX region begins at 2991. The site therefore survives unifdef -DUNX. What the region excludes is the helper’s ONLY caller, at 3013, so under -DUNX the function is compiled and never invoked. The emit path does not sit behind the #ifndef UNX region; only the caller does.
Fixture: CHECK_151_only_repetition_segments.cha. Chatter regression codes: E753.
Exclusion reason: UnreachableInFileMode.
CHECK 152 — parity
A [- CODE] utterance precode whose language is absent from @Languages. CLAN emits (152) via the indirect check_isLangMatch(wh=152) call at check.cpp 4194/5501, which the reference JSON’s literal check_err(N,) counter misses, so 152 is not in the dead list (the coverage tool auto-reclassifies once this entry exists). Adjudicated MEANINGFUL and closed as E755 per the @s-declaration ruling (docs design 2026-07-15, part 3): utterance-level presence is substantial and belongs in @Languages; wild grounding 0 of 7,167 precode-bearing files violate. Deliberate contrast: word-level @s:CODE carries NO declaration requirement (part 1 of the same ruling; the pre-2019 CHECK behavior and the retired E254 warning are explicitly NOT adopted: CLAN itself relaxed @s in 2019). SECOND FACET (2026-07-24): the other call site (check.cpp 5501) validates each token of an @New Language: header against CLAN’s static language table; @New Language is extinct in the kept corpus (0 files, rg sweep), and chatter rejects the whole header as unknown (E525 + E316), strictly subsuming the language-table facet.
Fixture: CHECK_152_undeclared_precode_lang.cha. Chatter regression codes: E755.
CHECK 153 — parity
@ID age with a component missing its leading zero (2;6. instead of 2;06.). CLAN 153; chatter E517 InvalidAgeFormat.
Fixture: CHECK_153_age_no_leading_zero.cha. Chatter regression codes: E517.
CHECK 154 — no_obligation
Call site commented out after four days in service (‘in 2018-10-03 out 2018-10-07’, check.cpp 5752); never emitted by current builds. Media/bullet consistency is adjudicated under the live codes (112 ruling of 2026-07-14 covers the meaningful direction: bullets require @Media).
Exclusion reason: CommentedOut.
CHECK 155 — parity
A non-pause word fully wrapped in parentheses, ‘(word)’, which should be written 0word (in a non-CA file). chatter rejects via E209 ‘Word has no spoken content’ (a fully-parenthesized word has no spoken material); not a dedicated use-0word rule but the same construct is refused.
Fixture: CHECK_155_paren_word.cha. Chatter regression codes: E209.
CHECK 156 — parity
A literal ‘,,’ (exactly two commas) on a main tier where the tag marker U+201E is required. Dedicated chatter code: E258 ‘Consecutive commas in utterance’.
Fixture: CHECK_156_double_comma.cha. Chatter regression codes: E258.
CHECK 157 — parity
@Media filename (‘differentname’) does not match the datafile basename. chatter has E531 MediaFilenameMismatch, which runs through the CLI because the file stem is threaded through validate_single_file_streaming. CLAN exempts remote URL @Media (verified: @Media: "https://..." -> no 157), so E531 exempts URLs too. CLAN CHECK 157.
Fixture: CHECK_157_media_name_mismatch.cha. Chatter regression codes: E531.
CHECK 158 — parity
A [: replacement] whose content is 0…, xxx/xx, yyy/yy, or www/ww instead of a real word. Dedicated chatter code: E391 ‘Replacement word cannot be untranscribed (xxx, yyy, www)’.
Fixture: CHECK_158_replacement_xxx.cha. Chatter regression codes: E391.
CHECK 159 — parity
A pause before its retrace marker is rejected by the grammar with E316. Canonical E375 pause-order controls retain valid marker-before-pause forms. No dedicated pause-order diagnostic or inferred missing-space advice is required.
Fixture: CHECK_159_pause_before_retrace.cha. Chatter regression codes: E316.
CHECK 160 — parity
Space directly after ‘<’ or before ‘>’ in an angle group (check.cpp 4300/4306, main + %wor tiers). The grammar models the tolerated whitespace as explicit optional CST nodes; the group parser reports E750 at each (both sides) rather than silently dropping them (which would also silently rewrite the text on normalize); re2c mirrors via LessThan/Whitespace/GreaterThan token windows. Spec: E750.md. Main-tier groups covered; %wor angle groups not separately grounded. ADJUDICATED (maintainer decision, 2026-07-10): STYLE error, no ambiguity whatsoever; legitimately caught as style discipline; kept as an error.
Fixture: CHECK_160_space_in_angle.cha. Chatter regression codes: E750.
CHECK 161 — parity
The grammar requires whitespace before a replacement. The glued fixture is rejected by E316; the spaced form remains legal. No replacement parser is reconstructed from recovery text.
Fixture: CHECK_161_no_space_before_replacement.cha. Chatter regression codes: E316.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Crate Reference
Status: Current Last modified: 2026-06-15 15:00 EDT
Summary of the main crates and packages in TalkBank/chatter.
Foundational crates
tree-sitter-talkbank
Rust binding crate for the generated TalkBank CHAT tree-sitter grammar. Exposes
LANGUAGE, NODE_TYPES, and the generated query constants used by editor and
parser integrations.
talkbank-model
The typed data model for CHAT files. Defines ChatFile, Utterance, DependentTier, MorTier, GraTier, and all other AST types. Includes validation logic, the WriteChat trait for CHAT serialization, serde support for JSON, and JsonSchema derivations. Also owns error types (ParseError, ErrorSink trait, Span, SourceLocation), diagnostic infrastructure, and ParseValidateOptions. Provides a closure-based content walker (walk_words / walk_words_mut) that centralizes recursive traversal of UtteranceContent and BracketedItem with domain-aware group gating.
talkbank-derive
Procedural macros for the model crate (SemanticEq, SemanticDiff, SpanShift, ValidationTagged, and the error_code_enum macro).
talkbank-cache
SQLite-backed validation and roundtrip cache used by higher-level validation and corpus workflows.
talkbank-parser
The canonical parser. Wraps the tree-sitter C parser and converts the concrete
syntax tree (CST) into ChatFile model types. Provides error recovery via
tree-sitter’s GLR algorithm and is the parser used by the CLI, LSP, transform
pipelines, and editor tooling.
talkbank-parser-re2c
Independent alternate parser used as an equivalence oracle against the tree-sitter parser. Primarily a testing and spec-hardening tool rather than a first-wave end-user surface.
talkbank-transform
High-level pipelines: parse+validate, CHAT-to-JSON, JSON-to-CHAT, normalization. Integrates the validation cache, JSON schema validation, and parallel directory validation.
Application and integration surfaces
chatter
The chatter CLI binary: validate, normalize, to-json, and corpus management.
talkbank-lsp
Language Server Protocol server with tree-sitter incremental parsing, real-time diagnostics, and semantic highlighting.
send2clan
Rust bindings for sending files to the CLAN application (macOS Apple Events,
Windows WM_APP). The crate exposes the safe send2clan API directly while
keeping the raw FFI in private modules.
chatter-desktop
Desktop validation app (Tauri v2, React). Mandates TUI parity with the CLI.
Test and spec-support crates
talkbank-parser-tests
Parser tests. Runs the parser over the reference corpus and validates the results. Also owns spec-generated tests, roundtrip tests, equivalence tests, and property tests.
spec/tools
Generator binaries for tree-sitter corpus tests, generated Rust tests, shared spec artifacts, and error documentation.
spec/runtime-tools
Runtime-aware spec tooling for validation, bootstrap, and corpus-mining tasks that should not live in the root Rust workspace.
This page last changed: 2026-06-23 (commit 06381dda). The whole book last changed: 2026-10-07 (commit 5e895791).
CLI Startup and the Program Stack
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
Why main() in crates/chatter/src/main.rs does not run the program
directly, and what every contributor adding CLI surface should know about
stack budgets.
The hazard this page addresses
The clap-derived command-tree construction (Cli::augment_args via
CommandFactory::command()) runs before argument parsing and is shared
by every subcommand. Its stack need grows with the number of declared
flags. On Windows in debug builds, where the main thread has 1 MiB, a
command tree large enough to cross that line makes every chatter
invocation fail with STATUS_STACK_OVERFLOW (exit code 0xC00000FD)
before argument parsing begins. Each added flag or subcommand moves the
tree closer to that line.
Why stack usage is not portable
Two multipliers vary independently, and the crash happens where they collide:
-
Platform main-thread allowance. There is no single default:
Context Main/default stack Windows main thread 1 MiB (set in the PE header at link time) macOS main thread 8 MiB Linux main thread typically 8 MiB ( ulimit -s)Rust spawned threads 2 MiB unless stack_sizeis givenShipping cross-platform means your real budget is the smallest of these: Windows’ 1 MiB.
-
Build profile. At opt-level 0, rustc gives every temporary in a function body its own stack slot and does not coalesce them, so a function’s frame is roughly the SUM of all its temporaries, not the maximum simultaneously alive. clap’s derive expands to one enormous builder function per args struct (one multi-call chain per flag, each
Arg/Commandtemporary a few hundred bytes by value), which is exactly the shape this penalizes. Release builds coalesce slots and inline, shrinking the same frames by one to two orders of magnitude.
Consequence: identical code can be fine in release on macOS (8 MiB budget, small frames) and fatal in debug on Windows (1 MiB budget, fat frames). Debug test binaries cross the line first, which is why CI subprocess tests are where a too-large command tree shows up, well before release binaries do.
The design: an explicitly sized program thread
main() spawns the entire program onto a thread with an explicit,
documented stack size (PROGRAM_STACK_BYTES, 16 MiB) and only joins and
re-raises panics, so exit semantics are unchanged. This removes the
dependency on platform main-stack defaults altogether instead of
holding the budget under an invisible, platform-dependent line that the
CLAN parity roadmap (roughly sixty commands’ worth of flags still to
come) would cross again. rustc itself uses the
same pattern for the same reasons.
flowchart TD
main["main()\n(crates/chatter/src/main.rs)"]
spawn["thread::Builder::stack_size(PROGRAM_STACK_BYTES)\n.spawn(program_main)"]
prog["program_main()\nclap tree build + parse + cli::run"]
join{"join() result?"}
ok["process exits normally"]
panic["resume_unwind(payload)\n(same exit behavior as a panic in main)"]
fail["spawn failed (OS resource):\neprintln + exit(1)"]
main --> spawn
spawn -->|"Ok(handle)"| prog
prog --> join
join -->|"Ok(())"| ok
join -->|"Err(payload)"| panic
spawn -.->|"Err(e)"| fail
The reservation is virtual address space; physical pages are committed only as they are touched, so the 16 MiB costs nothing measurable. The extra thread spawn at startup is microseconds.
Regression gates
crates/chatter/tests/stack_limit_tests.rsruns the real binary under a Windows-sized 1 MiB stack (sh -c 'ulimit -s 1024') on Unix, so macOS and Linux CI enforce the Windows constraint on every run. Without this, the constraint would be tested only by the windows-latest job.- The windows-latest cross-platform job is the native test of the real
1 MiB main stack (which does not constrain the program thread, but
guards the
main()shim itself).
Guidance for contributors
- Do not move program logic back onto the bare OS main thread; anything
before the
spawnruns under the platform’s smallest default. - Adding flags and subcommands is normal and expected; the budget is
the explicit
PROGRAM_STACK_BYTESconstant. If deep recursion or generated code ever approaches it, raise the constant deliberately in a reviewed change rather than discovering the limit in CI. - The same two multipliers apply to any worker threads you spawn:
Rust’s 2 MiB spawned-thread default is also finite. Do not spawn your
own: fan work out through
talkbank_transform::worker_pool::fan_out, whose workers run onCHAT_THREAD_STACK_BYTES(16 MiB). The program thread’sPROGRAM_STACK_BYTESis that same constant, not a copy of it.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Repository Architecture and Boundaries
Status: Current Last modified: 2026-09-28 20:59 EDT
Top-level layout
spec/ canonical syntax and error spec source
spec/tools/ deterministic generators + validators (separate Cargo workspace)
grammar/ tree-sitter grammar source + generated parser artifacts
crates/ all Rust crates (root Cargo workspace)
talkbank-model/ data model, validation, alignment, errors, parser API trait
talkbank-derive/ proc macros (SemanticEq, SpanShift, ValidationTagged, error_code_enum)
talkbank-parser/ canonical parser (tree-sitter)
talkbank-parser-re2c/ experimental alternate parser (opt-in)
talkbank-parser-tests/ parser equivalence and roundtrip tests
talkbank-transform/ pipelines, CHAT↔JSON, caching, parallel validation
chatter/ the `chatter` CLI binary
talkbank-lsp/ LSP server
send2clan/ Rust bindings to the legacy CLAN app bridge
talkbank-cache/ validation + roundtrip cache
apps/ desktop app (Tauri v2 + React): chatter-desktop
corpus/ reference corpus (must pass 100%)
schema/ JSON Schema for ChatFile AST
tests/ workspace-level integration tests and fixtures
book/ mdBook documentation source
docs/ strategy docs, proposals, and investigations
Architectural principles
- Clear boundaries between specification, generation, runtime logic, and documentation.
- Generated artifacts and hand-authored code are kept separate with
hard guardrails,
parser.c,node-types.json, generated tests and error-doc artifacts are never edited by hand. - Each crate has a single clear responsibility.
- Entry-point docs guide new contributors to authoritative references quickly.
Canonical ownership rules
spec/owns the language intent and accepted examples, what CHAT means.grammar/owns tokenization and CST shape only, not semantic validation policy.talkbank-modelowns semantic validity, serialization invariants, error types, and parser API contracts.talkbank-transformowns pipelines and JSON schema validation.talkbank-cacheowns the shared SQLite-backed validation and roundtrip cache.
Dependency direction rules
specdoes not depend on runtime crates.grammaris consumed by parser crates, not vice versa.talkbank-modelis dependency-minimal and stable; all other talkbank-* crates depend on it.- CLI / LSP / desktop apps depend on stable internal APIs, never directly on unstable internals of other crates.
- Generator tools may read specs and grammar metadata but do not become runtime dependencies.
Acceptance criteria
- Every top-level directory has a clear purpose statement.
- No crate depends on internal modules outside declared boundaries.
- No generated artifact is edited manually.
- New contributors can identify authoritative docs in less than five minutes.
This page last changed: 2026-09-28 (commit 2cb42a45). The whole book last changed: 2026-10-07 (commit 5e895791).
Grammar System and Token Governance
Status: Current Last modified: 2026-09-28 13:01 EDT
Current Reality
grammar/grammar.js encodes substantial implicit language knowledge directly in regex exclusions,
reserved symbol lists, and leniency decisions. Example areas:
- word segment forbidden start/rest classes,
- CA delimiter/element symbol groups,
- event segment exclusions,
- hand-maintained coupling between comments and token rules.
This is currently powerful but fragile.
Primary Failure Modes
- New symbolic token added in one place but not in exclusion sets.
- Parser behavior changes silently due to regex class edits.
- Generated node types drift from assumptions in spec tooling.
- Lenient parsing choices become undocumented policy.
Current Design
The generated symbol registry is the single source of token constraints.
The pipeline has shipped, just symbols-gen rebuilds it.
Registry Artifacts
spec/symbols/symbol_registry.json(human-authored intent):- symbol string
- category (delimiter, continuation, overlap, punctuation, etc.)
- contexts where reserved/allowed
- parse role and precedence notes
- Generated outputs:
grammar/src/generated_symbol_sets.jscrates/talkbank-model/src/generated/symbol_sets.rsspec/tools/src/generated/symbol_sets.rs- docs: Symbol Registry
Grammar Refactor Requirements
- Replace large manual regex strings with generated character classes.
- Keep final grammar readable by preserving semantic names in generated constants.
- Distinguish clearly between:
- syntax permissiveness,
- semantic validation restrictions.
- Add comments only for design rationale, not for duplicating manual references.
Node Type Drift Controls
- Enforce regeneration and consistency checks:
- grammar source change must regenerate parser and node types,
- node type constants consumed by
spec/toolsand parser code must compile, - CI fails if generated files differ from committed state.
Leniency Policy
Explicitly classify every lenient parse behavior:
- Parse-lenient + validate-strict.
- Parse-lenient + validate-warning.
- Parse-strict (hard fail).
Document this matrix in the Leniency Policy.
Recognized but unsupported headers
Grammar recognition does not guarantee a supported model representation.
@Thumbnail is recognized structurally but remains unsupported: source-bound
lowering reports E525 over the declaration and refuses to admit that header.
Recovery preserves following speech, but the recovered document is not a
strictly accepted document or a lossless serialization of the input. The E525
spec pairs this refusal with an otherwise identical supported @Comment
control. This is a Chatter support boundary, not a claim that thumbnail headers
are forbidden by the wider CHAT language.
Date and time token selection
Date/time headers declare strict lexical alternatives and whole-line fallbacks. Their selection must respect both complete values and malformed suffixes. Higher lexical precedence is not a harmless tie-break: it can select a strict prefix before a longer malformed value. Rule-order changes must also preserve empty-header recovery, not merely improve selection on valid examples. Follow Tree-sitter’s conflicting-token rules and review the complete diagnostic snapshot after any such change.
These CST types prove lexical shape, not calendar or clock validity. In
particular, strict_time includes the shared digit/separator alphabet; the
checked model and header-specific validator still decide whether a value is
supported and in range. Selection tests need complete reference values, suffix
error specimens and the established empty-header policies. A declared strict
alternative alone does not prove that the compiled lexer ever selects it.
Empty fields have header-specific policy: an empty @Date is E516, whereas
empty birth/start/duration values retain the existing omission policy. Their
spec controls require the header and following speech to survive parsing and
JSON replay with byte-exact CHAT output. Nonempty malformed suffixes remain
E518/E540/E541. Do not infer midnight from the legacy start-time model’s zero
components when its preserved value is empty; that representation is not
evidence of a known time.
Grammar Test Strategy
- Keep corpus tests generated from
spec/constructs. - Add targeted hand-authored edge tests for symbol boundary interactions.
- Add mutation-style tests for forbidden-character regressions.
- Add parser equivalence tests for tokenizer-sensitive cases.
Acceptance Criteria
- No manual reserved-symbol duplication in
grammar.js. - Symbol registry is generated to all required consumers.
- Grammar modifications cannot land with stale generated artifacts.
- Every special token category has explicit policy documentation.
This page last changed: 2026-09-28 (commit 2cb42a45). The whole book last changed: 2026-10-07 (commit 5e895791).
Parser, Model, and API Contracts
Status: Current Last updated: 2026-09-28 20:59 EDT
Single-handle parser API
talkbank-parser provides TreeSitterParser as the canonical API
handle for all parsing, full-file and fragment methods live directly
on the struct. Callers create one instance and pass
&TreeSitterParser everywhere. The alternate talkbank-parser-re2c
is opt-in and experimental (independent comparison and batch parsing)
and produces the same ChatFile model.
Contract for Batchalign
The Batchalign runtime (the batchalign crate) consumes these
guarantees from the talkbank-* core crates:
- parsing produces a typed
ChatFileor an explicit parse-status signal - parse-health taint is visible to alignment consumers
- alignment helpers operate on semantic model types, not raw text hacks
- recovery never fabricates valid-looking placeholder semantics for malformed input
The parser/model boundary stays honest enough for downstream
workflows, align, compare, benchmark, morphotagging, to make
their own validity decisions.
Canonical Contract Model
Public Contract Layers
- Parse API Contract:
- stable function signatures,
- deterministic parse result envelope,
- clear partial-success semantics.
- Semantic Model Contract:
- stable core model fields,
- explicit unstable/internal fields policy.
- Diagnostic Contract:
- stable error code IDs and severity semantics,
- best-effort message text compatibility.
- Serialization Contract:
- deterministic output constraints,
- normalized formatting policy.
Required Types
ParseOutcome<T>value: T | omitted-by-statusdiagnostics: Vec<Diagnostic>status: Success | Partial | Failed
Diagnosticcode,severity,category,message,location,context,suggestion
Parser Role
talkbank-parser: the sole parser, used by CLI/LSP/API/batchalign3.TreeSitterParseris the only API handle, callers create one and pass&TreeSitterParsereverywhere.- Tree-sitter GLR provides error recovery; the Rust traversal code converts CST to typed model.
- Full-file methods:
parser.parse_chat_file(),parser.parse_chat_file_streaming(). - Fragment methods:
parser.parse_word_fragment(),parser.parse_main_tier_fragment(), etc.
Invariants
- Parsing with offset must shift all spans consistently.
- Parse-level and validation-level diagnostics must remain distinguishable.
- Serialization should preserve semantic equivalence and documented formatting rules.
- Roundtrip behavior must be testable per parser implementation.
- Parser functions that accept
ErrorSinkshould not returnOption<T>for fallible parse state.
API Versioning Policy (Pre-1.0, Strict)
- Three intended contract levels:
- Stable-for-integrators
- Stable-internal
- Experimental
- Mark every public function/type by contract level.
This classification is not yet codified in a separate manifest file; the levels above are the working policy. Integrators should treat any unmarked surface as Experimental until contract levels are formally published.
Acceptance Criteria
- Single canonical parse outcome envelope exposed for integrators.
- Parser implementations conform to shared contract tests.
- Contract-level annotations exist for all public API surfaces.
- Documentation for parse/validate/serialize lifecycle is centralized and current.
Recovery Contract: No Fabricated Semantic Values
The parser contract must forbid sentinel semantic values during error recovery.
Disallowed recovery behavior:
- returning arbitrary enum variants as fallback for unknown/missing nodes,
- returning empty strings as stand-ins for required fields,
- constructing fake words/chunks like
"missing","error", or other placeholders.
Required recovery behavior:
- Emit structured diagnostic with precise span and expected node kind.
- Return an explicit parse-status signal (
Partial/Failed) throughParseOutcome. - Omit invalid semantic node OR store it in explicit recovery metadata, never as a valid semantic value.
Current enforcement:
- CI guardrail script tracks and blocks introduction of new
ErrorSink + Optionsignatures. - See
scripts/check-errorsink-option-signatures.shandscripts/errorsink_option_allowlist.txt.
Rationale:
- fabricated semantic values create secondary, misleading diagnostics against synthetic data,
- downstream tools cannot distinguish real user content from parser-generated placeholders,
- equivalence and regression tests become noisy and non-actionable.
For batchalign3, this is especially important because alignment workflows
must be able to tell the difference between:
- a malformed input that should taint or block alignment
- a recoverable input where raw text can be preserved
- a clean input that should proceed through the align/compare pipeline
String Storage Policy
The model uses three string storage strategies:
Arc<str>interning (interned_newtype!): For high-frequency repeated values (POS tags, stems, speaker codes). Global interner avoids redundant allocations.SmolStr(string_newtype!): For short strings (median 10-15 chars) that benefit from inline storage. O(1) clone, no heap allocation for strings ≤23 bytes.String: Only for utility types outside the core model (e.g.,semantic_diff/).
This page last changed: 2026-09-28 (commit 2cb42a45). The whole book last changed: 2026-10-07 (commit 5e895791).
Parser Backends
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
TalkBank has two CHAT parser implementations. Both implement the ChatParser
trait and produce the same ChatFile model type, not necessarily identical
values or recovery results.
The --parser flag selects the backend at the CLI boundary; everything
downstream consumes the shared model API. Backend differences remain observable
in supported input, diagnostics and recovery:
flowchart TD
cli["chatter validate --parser <backend>\n(ParserBackend enum,\nchatter cli_types.rs)"]
sel{"which backend?\n(ParserKind,\ntalkbank-transform\nvalidation_runner/config.rs)"}
ts["TreeSitterParser\n(talkbank-parser:\nGLR, incremental)"]
re2c["Re2cParser\n(talkbank-parser-re2c:\nre2c DFA + chumsky)"]
trait["ChatParser trait\n(talkbank-model\nparser_api/chat_parser.rs)"]
model["ChatFile\n(shared model type;\nbackend-specific recovery)"]
cli --> sel
sel -->|"tree-sitter (default)"| ts
sel -->|"re2c"| re2c
ts -->|"ParserDispatch::TreeSitter\n(worker.rs) implements"| trait
re2c -->|"ParserDispatch::Re2c\n(worker.rs) implements"| trait
trait --> model
ParserDispatch::new(kind) (in validation_runner/worker.rs) is the single
place that constructs the chosen backend from a ParserKind; both variants
wrap a ChatParser implementor, so the validation runner never branches on
backend again.
The shared ChatParser trait
Backend compatibility mandate
This section is the authoritative policy for re2c/tree-sitter compatibility. Backend measurements and regression baselines record evidence; they do not override this contract.
Both implementations must pursue the same independently justified CHAT syntax, meaning and validity rules idiomatically within their own architectures. Neither parser is an oracle for the other. For supported valid input, the goal is semantic agreement in the shared model, not identical internal representation.
For developers, these are mandatory constraints:
- Use re2c lexer states, rich tokens, parser-owned recovery and typed, source-bound evidence. Use typestate, ownership and validated constructors to preserve admission and recovery distinctions.
- Do not emulate tree-sitter’s CST, MISSING nodes, recovery traversal or recovered model shape merely to satisfy an equality test. Do not add a parallel parser, reparse reconstructed text, or scan raw input after parsing to manufacture another backend’s diagnostics.
- Report the fault the available evidence supports, at a truthful source location. Preserve useful content where sound, but never fabricate valid structure, discard faults silently, or treat a recovered result as valid. A broader honest diagnostic is preferable to invented specificity.
- Adjudicate differences as a genuine semantic defect, an acceptable architectural difference, or an explicitly unsupported experimental feature. Retain valid controls and invalid-input rejection tests. Do not weaken a CHAT rule or change a baseline solely to make a comparison pass.
- Exact diagnostic codes, wording, counts, ordering, rejection stage and recovery output are not cross-backend requirements. Fix an inaccurate diagnostic for its user impact, not because the other backend differs.
For users switching --parser:
| Aspect | What to expect |
|---|---|
| Supported valid CHAT | The intended meaning and shared-model semantics should agree. A disagreement needs investigation. |
| Invalid input | Diagnostic detail, number, order and highlighted recovery region may differ; identical error reports are not promised. |
| Recovered content | Partial models and retained fragments may differ. Recovery is not validity and is not a portable data-repair contract. |
| Experimental limitations | re2c is incomplete and may reject supported CHAT or miss invalid input. A clean re2c result is not a substitute for default tree-sitter validation. |
Use the default tree-sitter backend for production validity decisions and the LSP. Consumers must not depend on cross-backend diagnostic identity or interchangeable recovered output. Report concrete validity or semantic disagreements with a minimal input and parser version; a diagnostic difference alone does not establish a bug.
Before 1.0, re2c improvements are welcome when justified and affordable, but remain lowest priority. Full backend equivalence is not a release prerequisite; re2c may remain explicitly experimental and incomplete. This does not waive truthful diagnostics or the prohibition on architecture-distorting workarounds.
API shape
Both backends implement talkbank_model::ChatParser directly (the
tree-sitter impl landed 2026-07-24 in
talkbank-parser/src/api/chat_parser_impl.rs; the re2c impl has carried it
from the start). The trait is the parser-agnostic API for every
granularity: whole files, headers, utterances, main tiers, %mor/%gra
and the other dependent tiers, down to single words and relations. Each
method takes (input, offset, errors) and returns a ParseOutcome;
diagnostics stream through the caller’s ErrorSink.
Downstream consumers should bind on the trait, not on a concrete backend:
fn analyze<P: ChatParser>(parser: &P, text: &str) { /* ... */ }
selects the backend with one generic bound, including cross-target setups
(tree-sitter natively, pure-Rust re2c on wasm, where compiling
tree-sitter’s C runtime is undesirable). No facade or cfg-gated dispatch
module is needed on the consumer side. The wasm half of that contract is
pinned in CI: the wasm job in ci.yml checks talkbank-model and
talkbank-parser-re2c for wasm32-unknown-unknown on every push.
Two notes on the trait’s shape:
- The trait has generic methods (
errors: &impl ErrorSink), so it is not dyn-compatible; runtime backend selection uses a small enum such asParserDispatchrather thanBox<dyn ChatParser>. - On
TreeSitterParser, every trait method delegates to the matching inherentparse_*_fragmentmethod, so trait-path and inherent-path behavior are identical by construction. The conformance gate istalkbank-parser/tests/chat_parser_trait.rs.
TreeSitterParser (default)
For isolated words, parse_word_with_context (or the inherent
parse_word_fragment_with_context) applies the enclosing document’s effective
@Options: CA to the admitted typed word. A standalone parenthesized word then
has the same CA-omission interpretation as whole-document parsing; mixed lexical
shortenings remain shortenings. This transition preserves raw spelling and
source coordinates and does not retry rejected input. The context-free word API
does not infer file options: callers that know them should supply the context.
- Crate:
talkbank-parser - Technology: tree-sitter GLR parser
- Grammar:
grammar/grammar.js→ generated C parser - Strengths: Incremental reparsing (LSP), robust error recovery (GLR), CST-level diagnostics
- Weaknesses: Slower on batch workloads,
!Send + !Sync(one parser per thread)
Used by the LSP, the default CLI, and all production validation.
Re2cParser
- Crate:
talkbank-parser-re2c - Technology: re2c DFA lexer + chumsky parser combinators
- Grammar: Translated from
grammar.jsrules → re2c conditions + chumsky combinators - Strengths: 4-8x faster,
Send + Sync, zero constructor cost, independent experimental implementation - Weaknesses: No incremental reparsing, incomplete diagnostic parity, and it is not ready to judge CHAT validity (see below)
Used for parser parity testing and performance benchmarking.
Source ownership and participant recovery
Parsed values borrow the caller’s source; token storage and temporary
recovery buffers are released after parsing. The backend leaks no memory
(no Box::leak).
File parsing receives a LexedSource that privately owns tokens and their
lexer locations alongside the borrowed source. Its only constructor lexes
that source, preventing callers from pairing unrelated token and location
arrays. Participant lists consume those located tokens through one parser
shared with the fragment entry point:
flowchart LR
source["Source text"] --> lexed["LexedSource: tokens and locations"]
lexed --> parser["Participant list state machine"]
parser --> entries["HeaderParsed::Participants: recovered entries"]
parser --> errors["ErrorSink: located diagnostics"]
entries --> model["Header::Participants"]
The list distinguishes its initial state, a nonempty entry, and a consumed comma awaiting another entry. A trailing comma therefore reports E550 while preserving the preceding participants. Conversion receives parsed entries instead of reparsing raw header tokens, and header fragments forward the same diagnostics with the caller’s offset. The internal AST snapshot records this distinction; it does not define a serialized CHAT format change.
The re2c newline token represents one LF, CRLF or lone CR, matching the canonical grammar. It does not fuse consecutive breaks, so blank-line structure is kept. Source-aware file dispatch reports an unconsumed blank newline at its lexer span. Generated error fixtures preserve their exact line-ending bytes in Git; published Markdown normalizes display line breaks and labels that presentation.
Annotation categories survive conversion
Token classification produces ParsedAnnotation::Scoped(ScopedAnnotationParsed)
for annotations that decorate content. Retraces, replacements, language codes
and postcodes remain distinct outer variants. The model converter accepts only
ScopedAnnotationParsed and returns a ContentAnnotation directly: structural
markers cannot enter that conversion and be silently discarded through None.
Replacement lookup likewise returns its payload rather than an index requiring
a second match or an unreachable branch.
The file-level E757 spacing check uses the same classified categories for
closing annotations and retraces. It reads adjacent tokens from LexedSource
and reports the following word’s complete lexer span when the code is glued to
that word. The specification includes both glued examples and a spaced control;
the cross-backend gate checks those generated cases. The same located-token pass
rejects replacements glued to rich or reconstructed words with E375/E316, matching
word_with_optional_annotations and CHECK 161. The bracket-location boundary
test compares both backends against the violation and its spaced legal control.
Canonical closing-bracket recovery excludes absorbed trailing whitespace from
its highlight and builds its context from the original source. The internal AST snapshot
changes to show the category, while reference-corpus model equivalence guards
serialized CHAT behavior.
Postcode admission and diagnostic offsets
A PostcodeToken owns its lexer’s full span and a private payload state:
nonempty content after trimming trailing whitespace, or recoverable missing
content. The lexer preserves leading payload whitespace. TierBody::postcodes
contains only these tokens, so lowering cannot accidentally treat another token
kind as a postcode. Missing content emits E363 and contributes no model postcode;
valid content retains its source span. Main-tier and utterance lowering require
an error sink explicitly.
The re2c trait implementation streams diagnostics through one offset adapter. File diagnostics, utterance fragments, main-tier lowering and header/participant fragments therefore use the same rebasing operation as their models. This removes temporary diagnostic collection in header fragments and the discarded utterance diagnostics. Remaining fragment entry points that do not yet produce diagnostics are still a separate parity gap. The postcode boundary test loads the authored E363 examples and checks actual token spans, nonzero offsets, recovered tier content and canonical-parser normalization.
Morphology admission and recovery
The morphology lexer distinguishes the stricter first lemma character from its
continuation characters and requires content after each feature separator.
The parser splits an admitted token without inventing an empty lemma fallback.
On failed %mor parsing, RejectedMorTier retains raw tokens and reports E600
at construction, alongside the primary syntax diagnostic. It cannot convert to
a model tier or masquerade as an unsupported dependent tier. Utterance lowering
retains morphology taint so alignment does not treat the dropped tier as clean.
The authored E316 examples and a legal lemma/feature control exercise this path.
This does not change the grammar’s allowance for angle brackets inside a lemma;
it rejects the forbidden leading angle bracket shown by the source examples.
Dependent-tier prefix admission
The lexer distinguishes a complete TierPrefix, including its required colon
and tab, from an IncompleteTierPrefix recovered from a label. File dispatch
recognizes both forms. Dependent-tier recovery consumes this classification
rather than inferring malformed syntax from an empty body and a suffix check.
It reports E602 over the original complete line and retains the recovered
content for inspection. A complete prefix with no body remains a separate
content-validation question (E756). The E602 boundary test loads both malformed
specification examples and the valid colon-tab control and checks source spans.
Separator provenance belongs to lexical admission
PrefixToken owns the matched payload and separator provenance. Prefix lexer
rules consume spaces after the required tab, so those spaces never become
header or tier content. The same carrier covers ordinary headers, embedded
speaker headers, dependent prefixes, and main-tier separators. AST header lines,
main tiers and dependent entries retain the admitted TierSeparator through
model lowering. Both backends use the shared file validator and its CA
policy; there is no main-tier-only whitespace scan or separate CA probe. Separator spans are omitted from serialized AST/model metadata, and
CHAT serialization writes the canonical tab in both CA and non-CA files.
The source-spec boundary test covers all E758 examples, padded CA headers and tiers, canonical non-CA controls, exact byte spans and nonzero source offsets. It compares canonical serialized CHAT with tree-sitter. The lexer prefix payload and header/dependent-entry AST shapes change; the file inspection snapshot records that API change.
Lengthening counts preserve their source
The grammar admits a nonempty run of colons with no 255-character limit.
WordLengthening::count and the re2c AST carry NonZeroUsize, measured at the
parser boundary. Neither backend narrows source length to u8, so a run of
any length neither overflows nor silently wraps. A zero count cannot be constructed or decoded
from JSON, and both the default constructor and omitted JSON count mean one
colon. Serialization does not repair zero counts with max(1).
The Rust count field and with_count argument are NonZeroUsize. JSON retains the integer count field, omitted for one colon,
but accepts longer runs and rejects zero. The schema describes that boundary.
The public-parser regression checks both source roundtrip and semantic equality
at the former u8 boundary (255 colons); equality alone would allow both
backends to lose the same information.
Remaining parity limits
The dated re2c measurements
include known silent invalid cases; do not treat zero-silence results as
guarantees. A clean --parser re2c run is not a general validity
guarantee. Backend disagreements
remain in diagnostic specificity, extra or missing diagnostics, and source
locations. The per-case authority is
tests/integration/error_parity/baseline.rs; run its gate for derived counts.
Both backends feed the shared model validator, but source information discarded before lowering cannot be checked there. Rejected morphology preserves taint; other recovery paths and remaining dummy diagnostic locations still need review.
CLI Usage
# Default: tree-sitter
chatter validate corpus/
# Opt into the experimental re2c backend
chatter validate --parser re2c corpus/
# Roundtrip with re2c
chatter validate --parser re2c --roundtrip corpus/
The --parser flag accepts tree-sitter (default) or re2c. Cache entries
are parser-specific, switching parsers does not invalidate the other’s cache.
Parity Status
The reference-corpus equivalence and roundtrip gates compare actual parsed
models and serialized output. The error-spec gate
backends_diverge_only_where_recorded separately compares diagnostic code
sets against a named, bidirectional baseline: a newly divergent case fails,
and a resolved case must be removed from that baseline. E550 is not in the
baseline because file and fragment participant recovery agree, and neither is
E747: both lexers preserve single logical line breaks, and both parsers locate a
blank line under LF, CRLF and lone-CR endings while retaining its surrounding
utterances.
A passing baseline means that disagreements are accounted for, not that both backends meet every spec. The harness distinguishes backend agreement from each backend’s conformance to the declared spec. Run its report with:
cargo test -p talkbank-parser-re2c --test integration backends_diverge_only_where_recorded --locked -- --nocapture
No wild-corpus percentage or performance figure is recorded here; measure the current build rather than relying on a stored number.
Performance
Run benchmarks: cargo bench -p talkbank-parser-re2c --bench parse_comparison
When to Use Which
| Use Case | Recommended Parser | Why |
|---|---|---|
| LSP / editor integration | tree-sitter | Incremental reparsing |
| Batch validation (>100 files) | tree-sitter | re2c is faster but is not a validity authority |
| CI validation | tree-sitter | The two backends are not interchangeable validity authorities |
| Error diagnostics (user-facing) | tree-sitter | More specific E3xx codes |
| Parser comparison testing | Both | Disagreements require adjudication against the specs; neither backend is an oracle |
| Profiling / benchmarking | re2c | DFA lexer gives a performance floor |
Shared Model Infrastructure
Both parsers convert to the same talkbank_model::ChatFile type and share
post-hoc promotion logic:
TierContent::extract_terminal_bullet(): trailing InternalBullet → utterance bulletparse_bullet_node_timestamps(): structured bullet CST → (start_ms, end_ms)
CA intonation arrows are not promoted to terminators at the
parser/model boundary; both parsers leave them as Separator items.
See CA Terminator Resolution.
Detailed Parity Report
See crates/talkbank-parser-re2c/docs/parity-report.md
for the full gap analysis, divergence categories, and remaining work items.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Parser Leniency Policy
Status: Current Last updated: 2026-10-07 (commit 5e895791)
This document is the single source of truth for how the tree-sitter grammar,
Rust validation layer, and CLI tooling divide responsibility for enforcing the
CHAT specification. It consolidates decisions scattered across grammar.js
comments, analysis documents, and code.
Scope: Documentation only. This document does not implement new validation rules; it records what exists, what is intentionally absent, and proposes a roadmap for closing gaps.
Philosophy: Parse, Don’t Validate
The tree-sitter grammar intentionally accepts a superset of valid CHAT. The rationale:
-
Maximise parse coverage: Real-world
.chafiles contain legacy patterns, whitespace variations, and edge cases. A grammar that rejects them produces no AST and therefore no diagnostics. Accepting them gives the validation layer something to work with. -
Separate syntax from semantics: The grammar captures structure (headers, utterances, tiers, annotations). The Rust validation layer enforces semantic rules (required headers, participant declarations, alignment counts).
-
Enable configurable strictness: Different consumers need different policies. A roundtrip pipeline can be strict; an editor providing live diagnostics should be lenient. Validation profiles (see Validation Profile Infrastructure) make this possible.
Three-Tier Classification
Every intentional leniency decision falls into one of three tiers:
| Tier | Label | Meaning |
|---|---|---|
| A | Parse-lenient + validate-strict | Grammar accepts it; validation rejects it as an error |
| B | Parse-lenient + validate-warning | Grammar accepts it; validation emits a warning |
| C | Parse-lenient only | Grammar accepts it; no validation needed: the construct is genuinely optional or the broad acceptance is by design |
This classification was proposed in an earlier grammar governance analysis and is formalised here.
Leniency Matrix
Master table of every documented leniency decision in the grammar. The Status column indicates whether downstream validation compensates for the grammar’s permissiveness.
| # | Grammar Construct | Spec Requirement | Grammar Behavior | Tier | Validation | Error Code | Status |
|---|---|---|---|---|---|---|---|
| 1 | @UTF8 header | Required, must be first line | Optional (not enforced) | A | Validated | E503 | OK |
| 2 | @Begin header | Required | Optional (grammar.js ~L104) | A | Validated | E504 | OK |
| 3 | @End header | Required | Optional (grammar.js ~L106) | A | Validated | E502 | OK |
| 4 | Pre-first-utterance header order | No enforced order (matches CLAN CHECK) | choice(), any order (grammar.js ~L122-135) | C | N/A (by design) | , | OK |
| 5 | Headers after utterances | Allowed (e.g. @Bg, @Eg, @G, @Comment) | Interleaved freely | C | N/A (by design) | , | OK |
| 6 | Content type context restrictions | Unified across contexts | Unified base_content_item (grammar.js ~L731-738) | C | N/A (by design); specific semantic rules (E371, E372) exist separately | , | OK |
| 7 | Terminator presence | Required (except CA mode) | Optional (grammar.js ~L691-692) | A | Validated | E305 | OK |
| 8 | Bare shortening as word | CA mode only | Accepted anywhere | A | Validated | E2xx | OK |
| 9 | Trailing whitespace in annotations | Not specified | Optional trailing space (grammar.js ~L957, 966, 975, 1004, 1013) | C | N/A | , | OK |
| 10 | MOR segment Unicode | Very permissive (broad language support) | Exclusion-based regex (grammar.js ~L1909-1915) | C | N/A (by design) | , | OK |
| 11 | MOR fusional suffixes with hyphens | ALNUM + IPA only | Allows hyphens (grammar.js ~L1942-1945) | C | N/A (by design) | , | OK |
| 12 | MOR nested translations | No nested structures | Allows () and [] nesting (grammar.js ~L1954-1966) | C | N/A (by design) | , | OK |
| 13 | Linkers / language codes | Truly optional | Optional | C | N/A | , | OK |
| 14 | Word annotations | Truly optional | Optional | C | N/A | , | OK |
| 15 | Media bullet | Truly optional | Optional | C | N/A | , | OK |
| 16 | Group whitespace (leading/trailing) | No whitespace inside < > | Optional (grammar.js ~L1097, 1099) | C | N/A | , | OK |
| 17 | Long feature label characters | Limited character set | /[A-Za-z0-9@%_-]+/ (grammar.js ~L1327) | C | N/A | , | OK |
| 18 | Catch-all headers ($.anything) | Structured content for some headers | /[^\r\n]+/ for ~19 header types | C | N/A (content is opaque) | , | OK |
| 19 | Header gap whitespace | Single space/tab | repeat1(choice(space, tab)) (grammar.js ~L467, 477, 489) | C | N/A | , | OK |
| 20 | @Types header whitespace | No spaces around commas | Optional whitespace around commas (grammar.js ~L584-592) | C | N/A | , | OK |
Permissiveness Regression Decisions
Several validation rules are deliberately permissive because stricter forms produce false positives against the reference corpus. Each decision is summarised here with its ruling and rationale (the permissiveness regression log, archived, holds the full record).
Decision 1: [*] bare annotation, E214 retired
- Behaviour: Bare
[*](an emptyContentAnnotation::Error) is accepted without error; no code is emitted for it. - Rationale: Reference files (
errormarkers.cha,compound.cha) use bare[*]as valid CHAT. - Why the number stays retired. A code that names two different rules
(the bare-
[*]rule and “the scoped-annotation LIST is empty”) leaves its documentation and its implementation free to disagree, and nothing detects the drift when neither rule can fire.AnnotatedContentAnnotationsis non-empty by construction, so the empty-list rule is unrepresentable rather than merely unimplemented, and the bare-[*]rule stays retired on this decision’s own reasoning. - Revisit: If coded error annotations become required, that is a NEW code
against the
ContentAnnotation::Errorpayload, behind an explicit strict profile. Do not reviveE214: it has meant two things already.
Decision 2: @t without @s:<lang>, E248 disabled
- Behaviour:
@tis accepted without requiring@s:<lang>;E248(the code for a@tmarker lacking an explicit language marker) is not emitted. - Implementation:
talkbank-model/src/validation/word/structure.rscarries no such check. - Rationale: Reference file
formmarkers.chacontainsa@tand is expected to be valid. - Revisit: Scope to explicit strict validation mode if desired.
Decision 3: Undeclared inline language codes, E254 retired
- Behaviour: An explicit word-level
@s:LANGmarker carries no requirement to be declared in@Languages(reference filelang-marker.chastays valid).E254(UndeclaredExplicitWordLanguage) is a retired code and is not emitted. - Rationale:
@Languagesdeclares the transcript’s substantial languages; a one-word insertion is not substantial presence. CLAN CHECK imposes no@sdeclaration requirement either. - Neighbouring rules:
E255(WholeUtteranceLanguageSwitchShouldUsePrecode) covers whole-utterance@sruns that should use[- lang]precodes, and a[- lang]utterance precode whose language is absent from@LanguagesisE755. - Revisit: The number is retired and not reused.
Decision 4: Mixed-language digit legality, permissive-any rule
- Behaviour: For mixed/ambiguous markers, digits are accepted if legal in at least one applicable language (not required to be legal in all).
- Implementation: an
any()over the applicable languages intalkbank-model/src/validation/word/language/digits.rs. - Rationale: Prevents false positives in mixed-language reference examples.
- Revisit: Confirm spec intent for mixed/ambiguous validation semantics.
Decision 5: @Bg nesting, same-label only
- Behaviour:
E529only fires when nesting the same label (or same unlabeled scope key). Different labels may nest hierarchically. - Implementation: a
same_scope_opentest (not “any scope open”) intalkbank-model/src/validation/header/structure.rs. - Rationale: Avoids false positives on hierarchical markup patterns (e.g., HSLLD corpus).
- Revisit: Decide whether nesting policy should be global or per-label.
Decision 6: Temporal bullets in CA mode, checked for every file
- Behaviour:
E701/E704run for every file, CA included; there is no CA-mode skip. - Rationale: the temporal rules carry a 500 ms tolerance and per-speaker
semantics, and with them the CA reference files and all 994 kept
CA-declared files validate clean (full-population measurement). A CA-mode
skip has no CLAN CHECK counterpart and would be internally incoherent:
E362bullet monotonicity runs on CA files, soE701/E704must too. - Revisit: closed. No CA-specific temporal policy is needed: no policy difference between CA and non-CA files exists.
Decision 7: Pipeline severity threshold, errors only
- Behaviour: Pipeline returns failure only if at least one diagnostic
has
Severity::Error. - Implementation:
talkbank-transform/src/pipeline/parse.rs. - Rationale: Warnings should not block parse/transform/export pipelines.
- Revisit: Keep as default; add explicit
--strictflag/profile if needed.
Decision 8: Spacing warnings W210/W211, retired
- Behaviour: No style-level spacing warning runs around terminators and
overlap markers; the core main-tier validation path in
talkbank-model/src/model/content/main_tier.rshas no such pass. - Rationale: Such warnings produce unexpected diagnostics on files treated as valid in the reference workflow.
- Revisit: CLOSED (maintainer ruling). Real CLAN CHECK accepts the W210 construct (glued terminator), overlap markers hug their content by design so W211’s shape is valid CA notation, and no production code emits either. The numbers are retired and not reused; no lint profile will reintroduce them. The living spacing rules are E243, E749, E750, E751, E757, and E758.
Decision 9: Repeated @Date headers, accepted
- Behaviour: A file may carry any number of
@Dateheaders, anywhere headers are allowed, with the same or different values, adjacent or separated. Each is validated on its own (E516empty,E518malformed); no rule relates one@Dateto another. - Parity: CLAN CHECK accepts the same three shapes: two identical
@Datelines together, two different ones together, and a later@Dateafter utterances begin. Measured against CLANV 21-Sep-2026 11:00and chatter 0.27.0 with three minimal files, all accepted by both. - Why so permissive: Repeated dates are legitimate CHAT in real
corpora. Episode-structured transcripts put an
@Datebefore each recording session, so the same date repeats whenever several episodes share a day. Diary corpora open each day with its own@Date, including days that produced no utterances. Sessions recorded over two days carry both dates. A narrower rule (“two adjacent@Datelines are an error”) was proposed in 2026-10 for identical adjacent duplicates introduced by an export tool. The maintainer ruling is that chatter does not add@Daterules ahead of CLAN CHECK. - Revisit: Only when CLAN CHECK adds a repeated-
@Daterule. Chatter then follows it for parity, under a new code.
Decision 10: %wor word intervals, reversed rejected, zero-duration legal
- Behaviour: A
%worword bullet that ends before it starts (300_100) isE362(check_word_interval). A zero-duration word bullet (100_100) is legal, and word bullets may overlap or start out of order. - Relation to CHECK: CLAN CHECK
21-Sep-2026checks no%worbullets: it accepts both cases above, and reports its error 82 (“BEG mark of bullet must be smaller than END mark”) only for main-tier bullets. Measured with theE362spec examples renamed to match@Media. Chatter goes beyond CHECK for the reversed case only. - Why: the maintainer, 2026-10-06: “do what actually makes sense. CHECK is not
God and neither are we.” A reversed interval contradicts itself and has no
reading as timing; no corpus file was found with one, so rejecting it costs
nothing. A zero-duration interval places a word at an instant; at least
185,540 such bullets in 1,102 corpus files pass today, and rejecting them
would invalidate those files for no reader’s benefit. A consumer that needs
a positive interval for each word refuses it at its own boundary
(
assess_wor_timing_sequenceyieldsRejectedwith the slot named). - Planned: zero-duration word bullets in existing files are aligner
artifacts (words that were not located, recorded as instants). Once the
affected files are regenerated, a zero-duration word bullet becomes
E362as well, as a zero-duration main-tier bullet already is: a bullet that covers no time locates nothing.
Validation Gap Roadmap
Concrete items where the grammar is lenient but no validation compensates. Each proposes a new error code and priority.
Priority 1: @UTF8 Presence (E503), DONE
@UTF8 Presence (E503)- Grammar:
@UTF8is optional. - Spec: Required, must be the first line.
- Implemented:
E503(MissingUTF8Header) added tocheck_headers()intalkbank-model/src/validation/header/structure.rs. - Severity: Error.
- Note: All 340 reference corpus files contain
@UTF8, zero roundtrip impact.
Priority 2: Pre-First-Utterance Header Order (proposed E534), Not a Gap
- Grammar:
choice()accepts headers in any order between@Beginand the first utterance. - Assessment: CLAN CHECK does not enforce any ordering for post-
@Beginheaders; it validates presence and format only. Our grammar’s flexible ordering matches CHECK’s behavior. - Status: Reclassified from Tier B (GAP) to Tier C (by design).
Priority 3: Content Type Context Validation, Not a Gap
- Grammar: Unified
base_content_itemaccepts any content type in any context. - Assessment: The unified rule is correct by design. Nested groups are legal
CHAT (e.g.,
<the <dag> [: dog]> [= something]). The two specific semantic restrictions that do exist (no pauses in pho groups, E371; no nested quotations, E372) are already validated. - Status: Reclassified from Tier A (PARTIAL) to Tier C (by design).
Validation Profile Infrastructure
What Exists
Two kinds of setting, two types, two crates
What the validator COMPUTES and what a reader SEES are different questions, and conflating them is not a style matter: it decides what a cached verdict means. They are separate types, and deliberately not in the same crate.
RuleSelection (talkbank-model/src/errors/config.rs)
Which rules run. Every field here changes the diagnostics that exist, which is why this type, and only this type, derives the validation cache key.
let rules = RuleSelection::new().with_strict_linkers(); // turns on the [Opt-in] codes
new(): every always-on check, no opt-in checkwith_strict_linkers(): run the cross-utterance linker checks (chainable)strict_linkers_enabled() -> bool: querycache_key_fragment() -> String: the canonical text folded intotalkbank_cache::RulesVersion::current_with_rule_selection. DestructuresSelfwith no..rest pattern, so a new field is a compile error until someone folds it in.
PresentationPolicy (talkbank-transform/src/presentation.rs)
What a reader is shown, and at what severity, applied to diagnostics the
validator has ALREADY produced. --suppress lands here.
let policy = PresentationPolicy::new()
.downgrade(ErrorCode::IllegalUntranscribed, Severity::Warning)
.disable(ErrorCode::InvalidOverlapIndex)
.upgrade(ErrorCode::UnknownAnnotation, Severity::Error);
API: new(), downgrade(code, severity), disable(code),
upgrade(code, severity), set_severity(code, Option<Severity>),
effective_severity(code, original) -> Option<Severity>,
is_disabled(code) -> bool, shows_everything() -> bool,
apply(diagnostic) -> Option<ParseError>, apply_all(Vec<ParseError>).
Pre-built profiles:
lenient(): showsIllegalUntranscribedandInvalidOverlapIndexas warnings. For gradual migration of legacy corpora.strict(): shows unmapped warnings as errors. Explicit per-code overrides still take precedence, so a caller can opt a specific code back toSeverity::Warning.
Why the crate split. talkbank-transform depends on talkbank-cache, so
the cache crate cannot name PresentationPolicy. Folding a display preference
into the cache key is therefore a dependency cycle rather than a judgement call.
Were a display preference part of the cache key, --suppress would partition
the cache: two runs differing only in what they printed would share no
entries, and a second pass over a large corpus would re-validate all of it
from cold.
What this makes true of a cache row. The stored fact is “this file produced no diagnostics at all under this rule selection”. No presentation policy can change that, which is what lets one cache serve suppressed and unsuppressed runs alike.
ConfigurableErrorSink (talkbank-transform/src/presentation.rs)
Wrapper that applies a PresentationPolicy to diagnostics on their way to an
inner ErrorSink, for surfaces that stream to a reader as they arrive.
let inner = ErrorCollector::new();
let sink = ConfigurableErrorSink::new(&inner, policy);
It must never wrap a sink whose output feeds a cache write or a run tally: those consume the complete diagnostic set.
Runner-Level Flags (talkbank-transform, chatter)
| Flag | Effect |
|---|---|
--skip-alignment | Skip tier alignment validation |
--roundtrip | Test serialization idempotency after validation |
--force | Clear cache for path and revalidate |
--max-errors N | Stop once N errors (never warnings) are found (N is at least 1) |
What Is Missing
| Gap | Description | Effort |
|---|---|---|
No --profile CLI flag | Users cannot select strict / lenient / lint from the command line | Medium |
| No profile serialization | Cannot load profiles from TOML/JSON config files | Medium |
| No corpus-specific profiles | E.g., HSLLD-specific rules | Future |
Proposed Profiles
From the permissiveness regression log:
| Profile | Purpose | Behaviour |
|---|---|---|
reference-compatible | Current permissive baseline | Default, matches current validation behaviour |
strict-chat | Full spec enforcement | Re-enable selected tightenings (E248, etc.; E254 and E214 are retired codes and are not candidates) |
The roundtrip gate should be pinned to an agreed profile to prevent future ambiguity about what “pass” means.
Silent Recovery Points (NLP Pipelines)
An earlier Python-Rust boundary audit identified several
places where batchalign-core silently massages data without diagnostics. These
are related to leniency because they represent permissive acceptance without
transparency.
| Pipeline | Recovery Mechanism | Diagnostics? |
|---|---|---|
| Stanza morphosyntax | retokenize.rs DP alignment; Word::new_unchecked fallback | No |
| Whisper/Wave2Vec FA | forced_alignment.rs DP “best fit” | No |
| Google Translate | Imported verbatim into %xtra | No filtering |
| Stanza segmentation | Silent abort on assignment mismatch | No |
Key infrastructure gap: ParseHealth exists in talkbank-model (per-utterance
tier cleanliness flags with taint(), is_clean(), can_align_main_to_mor()
methods). It is used by the tree-sitter and direct parsers during parsing.
However, batchalign-core does not read, write, or propagate ParseHealth
during any mutation (morphosyntax injection, FA injection, retokenisation). The
infrastructure exists in the model layer but is not connected to the pipeline
layer.
Cross-References
| Source | What It Contains |
|---|---|
| Grammar governance analysis (archived) | Proposed this document; leniency matrix concept; three-tier classification |
| Permissiveness regression log (archived) | 8 permissiveness regression decisions with rationale |
| Python-Rust boundary audit (archived) | Silent recovery points; ParseHealth gap; NLP pipeline audit |
grammar/grammar.js | Inline comments on each leniency decision (line references in matrix above) |
talkbank-model/src/errors/config.rs | RuleSelection API (and the cache key derived from it) |
talkbank-transform/src/presentation.rs | PresentationPolicy and the ConfigurableErrorSink adapter |
talkbank-model/src/validation/header/structure.rs | Header validation: E501, E502, E503, E504-E533 |
talkbank-model/src/validation/temporal.rs | Temporal constraint checks (E701, E704), run for every file |
talkbank-model/src/model/content/main_tier.rs | Main-tier validation path; carries no W210/W211 spacing pass |
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Error Diagnostics UX Standard
Status: Current Last modified: 2026-05-30 07:08 EDT
Workspace-wide standard for diagnostic shape, severity, recovery
behavior, span correctness, and integrator output formats. Applies
to the CHAT-core error system. Upstream
batchalign-runtime errors follow the same shape and are documented
separately in the batchalign3 project.
Objective
Make diagnostics precise, explainable, and actionable for both developers and non-technical editors, while keeping machine readability for downstream tools.
Open concerns
- Message quality across the error catalog is not yet governed by one central style standard. Different error codes were authored at different times and converge unevenly on the message-quality guidance below.
Canonical Diagnostic Schema
Diagnostic {
code: String,
severity: Error | Warning | Info,
category: Parse | Validation | Alignment | Header | Tier | Internal,
location: SourceLocation,
context: ErrorContext,
message: String,
suggestion: Option<String>,
related: Vec<RelatedLocation>
}
Message Quality Standard
Each diagnostic must answer:
- What failed.
- Where it failed.
- Why it likely failed.
- What to do next.
Avoid internal jargon unless accompanied by user-facing explanation.
Severity Policy
Error: blocks parse/validation outcome.Warning: content is usable but has quality/compliance concerns.Info: optional guidance and migration hints.
Severity must not be overloaded for tooling convenience.
Recovery Policy: Diagnostic-First, Not Sentinel-First
When parser recovery is required:
- do not invent semantic fallback values to keep type construction convenient,
- do not use empty strings or arbitrary enum defaults as recovered content.
Instead:
- Report a diagnostic with expected/actual node context.
- Preserve span information for tooling and UI.
- Propagate partial/failure status explicitly.
Any synthetic placeholders that are unavoidable for internal plumbing must be:
- non-semantic (not exposed as real model content),
- marked internal-only,
- excluded from user-facing diagnostics and serialization.
Sentinel vs error-variant rule
If an unexpected condition changes semantic trust in parsed content:
- Represent that explicitly as an error-bearing state (enum variant, parse-taint flag, or explicit outcome type).
- Never represent it as
Noneor a default payload that can be mistaken for valid content.
This applies both to parser outputs and to runtime metadata consumed during validation.
Diagnostic construction
Use shared constructors/helpers for common diagnostics to reduce drift:
- span-only diagnostics (
code + severity + span + message), - source-backed diagnostics (
code + severity + span + source + offending + message).
Benefits: consistent location/context population, fewer ad-hoc
ParseError::new(...) call shapes, simpler migration to richer
miette rendering.
Error Code Governance
- Central registry under
talkbank-model(errors module). - One authoritative description and example per code.
- Deprecated codes remain mapped with explicit migration notes.
- CI check forbids duplicate code definitions or orphaned docs.
Span and Location Correctness
- All diagnostics use consistent line/column and byte-offset definitions.
- Golden tests cover:
- single-byte and multi-byte UTF-8 content,
- embedded content offsets,
- continuation lines and tabs.
Integrator Output Formats
- Human-readable CLI diagnostics.
- Machine-readable JSON diagnostics.
- LSP diagnostic mapping.
All formats share the same underlying diagnostic schema.
Acceptance Criteria
- Every emitted diagnostic includes code, severity, location, and suggestion policy.
- Error code documentation and runtime definitions are synchronized automatically.
- Span correctness is covered by dedicated tests.
- CLI and JSON outputs are contract-tested for schema compliance.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Wide Struct Audit
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
A repository-wide audit rule for struct shape. Applies to the crates in
TalkBank/chatter (model, parser, transform, CLI, CLAN, LSP, cache, and
related tooling). This page is scoped to that repository.
A struct with many fields is not automatically wrong. The smell is:
- many unrelated concerns packed into one value
- several related booleans that act like implicit policy enums
- repeated field-name prefixes that point to missing sub-structs
- parallel vectors or stringly runtime fields
- runtime code reaching into many unrelated fields of the same value
The repo therefore treats 10 or more named fields as an audit threshold, not as an automatic ban.
Categories
Wide structs fall into four categories.
1. Boundary shim, may stay wide
CLI, JSON, or clap boundary types. Acceptable if they are converted into typed policies or sub-structs before entering core runtime code.
Examples: ValidateDirectoryOptions, clap-facing CLI arg structs, JSON boundary
records.
2. Transport or schema record, may stay wide
DB rows, HTTP response shapes, JSON schema mirrors. Acceptable as long as they don’t become the internal runtime shape.
Examples: WordJsonSchema, DbMetadata, CoverageReport.
3. Real aggregate, may stay wide
Domain values whose fields all answer one coherent question and whose callers consume the whole rather than spelunking through unrelated subsets.
Examples: metric/report records like SpeakerEval, SpeakerKideval,
SpeakerComplexity, and SpeakerFluency (report records, not runtime
coordination).
4. Refactor target, must be split
Mix of policy and state, multiple responsibilities, or callers needing to know the whole subsystem to use a subset of fields.
Design Rules
- Treat 10 or more named fields as an audit trigger.
- Treat 3 or more related boolean fields as a smell even below that threshold.
- Boundary and transport records may stay wide when they mirror a real external shape.
- Runtime coordination structs prefer named sub-structs over flat bags.
- Replace parallel vectors with per-item records where possible.
- If a wide struct stays wide, record the reason in the surrounding design docs, audit notes, or code review rather than letting it remain unexplained.
Refactor Examples
ValidateDirectoryOptions (chatter)
Format, cache, roundtrip, parser, audit and TUI settings are grouped by concern rather than held as a flat bag:
ValidationRulesValidationExecutionValidationPresentation(one enum:Streamed(Lines | Json | Audit)orTui, resolved once from--format,--quiet,--auditand--tui-mode; one value stands in for a format, a quiet flag, an interface flag and an audit path, so no precedence between them is ever evaluated)- the resolved
--suppresscodes
(There is no traversal-mode field: every invocation is an explicit file list, so no traversal choice exists to carry.)
Shape this audit wants for policy-rich CLI boundaries: one small top-level struct with explicit sub-objects and enums rather than a dozen flat fields.
ParseHealth (talkbank-model)
Stores taint as a compact tier bitset keyed by ParseHealthTier (no
per-tier boolean fields), the shape this audit expects for fixed domain
sets.
flowchart LR
tier["ParseHealthTier"] --> set["Tier health set"]
set --> checks["Alignment safety checks"]
Open Hotspots
TUI state bags
Real state owners that still want grouping by concern (selection vs. progress vs. render flags vs. status):
crates/chatter/src/ui/validation_tui/state.rsTuiState
Backend (talkbank-lsp)
crates/talkbank-lsp/src/backend/state.rs is a service-root aggregate.
Defensible, but still wants grouping such as document caches, parse caches,
validation state, language services.
Metric structs
SpeakerEval/SpeakerKideval are acceptable as report records. If output
renderers keep needing subsets (lexical metrics, morphosyntax metrics, error
counts, derived scores), those records should eventually nest along those
lines.
Audit Guardrail
There is currently no repo-local automated wide-struct lint in
TalkBank/chatter. Treat this page as a manual review checklist and refactor
trigger: when a type grows past the threshold, decide explicitly whether it is
an acceptable boundary/schema aggregate or a real split target.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Spec Tooling
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
What the generator crates ARE. For the spec system’s contract, which is what you need to write or change a spec, read Spec System; for the procedure, Spec Workflow. This page covers only the tooling, so the three do not overlap.
Two crates, one workspace
spec/ is its own cargo workspace, so every command needs
--manifest-path spec/Cargo.toml.
| Crate | Owns | Depends on the parser? |
|---|---|---|
spec/tools (generators) | reading specs and emitting artifacts | No |
spec/runtime-tools | anything needing the live parser or model | Yes |
That split is the point. spec/tools reads markdown and JSON and produces
tests, fixtures, docs and generated Rust; it never parses CHAT. Work that has to
actually run the parser (verifying a spec example emits its codes, mining the
corpus) lives in spec/runtime-tools.
The artifact registry is split along the same line, and for the same reason:
generators::artifacts::ARTIFACTS holds everything derivable from markdown
alone (plus, since R4, the observation snapshot as a data-file input), and the
runtime half holds the artifacts that need the live parser or ErrorCode
enum: the observation snapshot itself (which regenerates FIRST, being an input
to the tree-sitter corpus), the DiagnosticKind registry, and the book’s
artifact table. spec_gen runs both halves in dependency order, so a
contributor sees one command and one list; the generated artifact table included in the spec-system chapter is the
live inventory.
Layout of spec/tools
src/
bin/ one binary per generator
spec/ markdown spec loaders (constructs, errors)
output/ formatters (tree-sitter corpus, Rust tests, docs)
form_markers/ the form-marker registry: typed model, renderers, drift gate
templates/ Tera templates wrapping fragments into whole CHAT files
generated/ generated symbol sets (never edited by hand)
Determinism, and what enforces it
Generation must be idempotent: a re-run with no source change produces no diff. Three things make that true rather than hoped for.
- Generators write only when content differs, so a no-op run does not churn mtimes.
- Rust output is formatted by the generator, which runs
rustfmtitself. Otherwisejust fmtand the generator each rewrite the same bytes forever, both correct. Both registries do this. - Drift gates compare committed artifacts against what the generators produce, calling the real generators rather than a second description of their output. See Spec System for the full list.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Symbol Registry Architecture
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
Purpose
spec/symbols/symbol_registry.json is the canonical source of the token and
symbol classes CHAT tokenization policy depends on. It is one of two closed
vocabularies owned under spec/; the other is
the form-marker registry.
Scope
The registry holds two kinds of entry, and they are deliberately different shapes.
symbols are ENTITIES. Each has an identity (a codepoint), a name, a
meaning, a parse role, a notation family and a runnable example, and each maps
1:1 to a Rust enum variant. These are the 25 word-attached and paired-stretch
symbols.
character_classes are SETS. Bags of characters with no individual
identity, used to build the grammar’s word and event regexes: word-segment
forbidden (start, rest, common) and event-segment forbidden (base, common).
Storing a set as a list of records, or an entity as a bare character, would be the same category error in opposite directions.
The paired_stretch_symbols and word_attached_symbols arrays the grammar and
the model consume are derived from parse_role and appear nowhere in the
file, and they are NAMED for the role they are derived from, not for a notation family:
a name asserting provenance on a value holding a parse role would be the
collapse the next section warns about.
parse_role is not provenance
parse_role says what the GRAMMAR does with a symbol. notation_family says
where it comes from. They are independent, and collapsing them would file two
disfluency marks (≠ blocking, ↫ segment repetition) as
Conversation Analysis notation; CLAN names them NOTCA_CROSSED_EQUAL and
NOTCA_LEFT_ARROW_CIRCLE. Code that needs to know “is this CA” calls
notation_family(); never infer it from the name of a ca_* array.
Rules
- Symbols change in
spec/symbols/symbol_registry.jsonand nowhere else. - Regenerate after any change:
just symbols-gen, which validates the registry and then runs both generators. - Generated files are never edited by hand.
What is enforced, and where
Most structural checking happens in spec/symbols/registry.js, which every
generator reads the registry through, so a malformed registry cannot reach a
generator even if nobody runs the validator: required fields present, ids
snake_case and unique, codepoints well-formed and unique, parse_role and
notation_family from their closed sets, and every example containing its own
symbol. validate_symbol_registry.js adds the character-class checks (single
Unicode scalar values, no duplicates) and prints the report.
There is no disjointness check. The two derived arrays come from a single
parse_role field, so a symbol in both is unrepresentable and there is
nothing to assert.
Every example is additionally PARSED AND VALIDATED by
crates/talkbank-parser/tests/integration/symbol_registry_examples.rs, so a
documented usage that stops being valid CHAT fails the build. The uniform
example template is valid for 24 of the 25 symbols and invalid for ↫, which
needs a stem outside its brackets.
Lexicographic ordering is NOT required, and no category is sorted. Nothing enforces it, and the validator says in its own comment that semantic grouping is more useful than forced ordering. Reordering would make a large diff that buys nothing.
Generated outputs
| Output | Consumer |
|---|---|
grammar/src/generated_symbol_sets.js | imported by grammar/grammar.js |
crates/talkbank-model/src/generated/symbol_sets.rs | model and validation |
spec/tools/src/generated/symbol_sets.rs | spec tooling |
crates/talkbank-model/src/generated/ca_symbols.rs | CAElementType, CADelimiterType, NotationFamily |
book/src/chat-format/generated/ca-symbols.md | included by the book’s symbols page |
The Rust outputs are formatted by the generator itself, which runs rustfmt
before writing. That is not tidiness: without it just fmt re-wraps the const
arrays, re-running the generator un-wraps them, and the two rewrite the same
bytes forever with both sides correct. Generating and formatting have to be one
state. The form-marker generator does the same, for the same reason.
The drift gate
generated_symbol_sets_are_current, in
spec/tools/src/form_markers/mod.rs, runs each generator in --check mode
(render, compare, write nothing, exit non-zero on drift) and fails if any
committed output disagrees with the registry. It runs in CI under
cargo test --manifest-path spec/Cargo.toml --workspace.
It runs the REAL generators rather than re-describing their output, so there is
no second description to drift. They are JavaScript, so the gate shells out to
node.
The generator list is DISCOVERED, not written down. A gate that lists what
it covers stops covering things silently, so the gate globs
spec/symbols/generate_*.js and refuses to report at all if the glob finds
fewer than two.
A hand-edit to a generated symbol set is therefore detected by the gate, which runs in CI.
Change workflow
- Edit the registry JSON.
just symbols-gen.- Regenerate the parser if the grammar’s tokenization changed: see Grammar Workflow.
- Run the gates:
cargo test --manifest-path spec/Cargo.toml --workspaceandjust test. - Commit the registry and every regenerated output together.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Bullet Validation
Status: Current Last updated: 2026-09-24 00:21 EDT
Media bullets are timestamps embedded in CHAT utterances that link transcript
text to audio/video. They appear as •start_end• at the end of a main tier
line (e.g., *CHI: hello . •1000_2000•). Validating that these timestamps are
internally consistent is one of the more subtle parts of CHAT validation,
because the “obvious” rules turn out to be wrong for multi-party conversation.
This chapter documents what CLAN CHECK does, where its implementation falls
short of its own intent, and how chatter validate interprets and improves on
that intent.
The three temporal checks
There are three distinct temporal constraints that can be checked on bullet timestamps. They differ in scope, severity, and whether they should run by default.
E701: Same-speaker start-time monotonicity (CLAN Error 83)
Rule: For each speaker, their utterances’ start times must be non-decreasing. If speaker CHI has utterance A starting at 10,000ms and utterance B (later in document order) starting at 8,000ms, that is an error, CHI’s timeline has gone backward.
Scope: Per-speaker. Cross-speaker non-monotonicity is allowed (see Why cross-speaker non-monotonicity is not an error).
Severity: Error.
E704: Same-speaker self-overlap (CLAN Error 133)
Rule: For each speaker, the current utterance’s start time must not be more than 500ms before the same speaker’s previous utterance’s end time. In other words, a speaker cannot overlap with themselves by more than 500ms.
Scope: Per-speaker. The 500ms tolerance accounts for annotation rounding and minor timing imprecision at boundaries.
Severity: Error.
E729: Cross-speaker overlap (CLAN Error 84)
Rule: The current utterance’s start time must not be before the previous utterance’s (any speaker) end time. This checks for any temporal overlap between adjacent utterances, regardless of speaker.
Scope: Global (cross-speaker). Only fires with CLAN’s +c0 flag.
Severity: Warning. Not part of default validation.
This check is part of CLAN’s “strict timeline contiguity” mode, which requires that every utterance’s start time equals the previous utterance’s end time, no gaps (Error 85) and no overlaps (Error 84). It is designed for a very specific use case: verifying that audio has been exhaustively and non-redundantly segmented. In normal conversational transcripts, cross-speaker overlap is ubiquitous, so this check would be absurd as a default.
What CLAN CHECK does
CLAN CHECK implements bullet validation in the function
check_checkBulletsConsist() in check.cpp. Understanding its implementation
is essential because it has several accidental behaviors that affect the error
counts users see.
The snapshot-and-compare pattern
The function uses a global pair (check_SNDBeg, check_SNDEnd) to hold the
“current” bullet timing, and saves the previous values into local variables
(tBegTime, tEndTime) at the start of each call. The comparison flow is:
1. Save previous: tBegTime = check_SNDBeg, tEndTime = check_SNDEnd
2. Parse new bullet into check_SNDBeg, check_SNDEnd
3. Check error 83: check_SNDBeg < tBegTime? (cross-speaker comparison)
4. Check error 133: speaker's last END - check_SNDBeg > 500? (same-speaker)
5. If +c0 mode: check error 84 (overlap) and error 85 (gap)
6. Update speaker's last END time via check_setLastTime()
The early-return shadowing bug
The critical implementation detail is that error 83 fires via return(83) at
step 3. This causes the function to exit immediately, skipping steps 4
through 6. Two consequences follow:
-
Error 83 shadows error 133. An utterance that triggers error 83 (global non-monotonicity) can never also trigger error 133 (same-speaker overlap) in the same call, even if both conditions are true. This is not intentional, it is an artifact of C-style early-return control flow.
-
Speaker state goes stale. Step 6 (
check_setLastTime) updates the speaker’s per-speaker tracking in theSPLISTlinked list. When error 83 fires, this update is skipped. All subsequent error-133 checks for that speaker compare against a staleendTimevalue, causing cascading state corruption that suppresses legitimate error 133 reports.
Error 83 is global, not per-speaker
CLAN fires error 83 by comparing the current utterance’s start time against the previous utterance’s start time, regardless of speaker. In a multi-party conversation:
*PIL: something . •100000_102000•
*UEL: response . •99500_101000• ← Error 83: 99500 < 100000
This fires error 83 because UEL’s start time (99,500ms) is before PIL’s start
time (100,000ms). But this is just two people talking at the same time, normal
conversational overlap. The [>] and [<] markers in CHAT explicitly annotate
this as intentional simultaneous speech.
In files with many speakers (the Koine/bre corpus has 7-9 speakers per file, including children talking over each other), this fires on a huge fraction of utterances. CLAN’s accidental shadowing partially masks the problem by suppressing downstream error-133 reports when error 83 fires.
Why cross-speaker non-monotonicity is not an error
Consider a classroom recording with a teacher (PIL) and seven children. The teacher asks a question, and three children answer simultaneously:
*PIL: qué es esto ? •50000_52000•
*UEL: un coche . •51200_52500• ← started during PIL's question
*MAR: coches . •51000_51800• ← started even earlier
*REN: es un coche grande . •51500_53000• ← started between UEL and MAR
In document order, the start times are: 50000, 51200, 51000, 51500. This is non-monotonic (51000 < 51200), but there is nothing wrong with this data. The children are simply talking at the same time. No amount of reordering the utterances in the file would make all start times monotonically increasing while preserving the speaker-turn structure.
Cross-speaker non-monotonicity is an inherent property of multi-party conversation, not a data error. Flagging it as an error produces thousands of false positives on any corpus with overlapping speech.
When IS non-monotonic start time an error?
Same-speaker non-monotonicity IS an error. If CHI speaks at 10,000ms, then later in the file CHI speaks again at 8,000ms, CHI’s timeline has gone backward. This almost certainly indicates a transcription or alignment mistake.
The test is simple: within the same speaker’s utterance sequence, start times
must be non-decreasing. This is what chatter validate checks for E701.
How chatter validate implements bullet validation
E701: Per-speaker monotonicity (not global)
chatter validate tracks each speaker’s last start time in a HashMap. E701
only fires when the same speaker’s start time goes backward. Cross-speaker
non-monotonicity is silently accepted.
This is an intentional semantic divergence from CLAN CHECK, which fires error 83
globally. We believe CLAN’s global check reflects the implementation
(comparing against a single global tBegTime) rather than the intent (detecting
disordered timestamps). The per-speaker version matches the intent without
drowning users in false positives from normal conversational overlap.
E704: Per-speaker overlap with 500ms tolerance
chatter validate tracks each speaker’s last end time in a HashMap. E704
fires when the overlap exceeds 500ms (same threshold as CLAN Error 133).
Unlike CLAN, E704 runs independently of E701. An utterance can trigger both errors if it is both non-monotonic (E701) and self-overlapping (E704). CLAN’s early-return pattern prevents error 133 from firing when error 83 fires, which is a bug, not a feature.
Speaker state is always updated regardless of whether errors fire. This avoids the cascading state corruption that CLAN’s implementation suffers from.
E729: Not in default validation
E729 (CLAN Error 84, cross-speaker overlap) is reserved and unimplemented.
Chatter does not expose a strict-bullet mode equivalent to CHECK’s +c0.
The parity manifest records a deliberate divergence for option-gated CHECK
84/85/110: these are not requirements to reject ordinary conversational CHAT.
E730/E732 are likewise reserved; E731 does not add a zero-tolerance duplicate
of the implemented E704 timing rule. The four specs retain real bullet-bearing
default-mode controls, including E704’s accepted exact-500-ms boundary.
Untranscribed utterances are skipped
Utterances containing only untranscribed markers (www, xxx, yyy) are
skipped for E704 checks. These utterances often carry broad segment bullets
(covering a long span of background speech) that would create false self-overlap
reports. This matches CLAN CHECK’s behavior, where untranscribed tiers do not
contribute to timing comparisons.
CA mode disables all temporal checks
When the file header includes @Options: CA, all temporal validation is
skipped. Conversation Analysis mode intentionally relaxes timing constraints
because CA transcription conventions use overlapping and non-sequential timing
as part of the analytic notation.
Comparison: CLAN CHECK vs chatter validate
The following table summarizes the behavioral differences:
┌────────────────────────────┬──────────────┬─────────────────┐
│ Behavior │ CLAN CHECK │ chatter validate│
├────────────────────────────┼──────────────┼─────────────────┤
│ Error 83 / E701 scope │ Global │ Per-speaker │
│ Error 133 / E704 scope │ Per-speaker │ Per-speaker │
│ Error 84 / E729 default │ Off (+c0) │ Off │
│ 83 shadows 133 │ Yes (bug) │ No │
│ 83 corrupts speaker state │ Yes (bug) │ No │
│ E701 + E704 independent │ No │ Yes │
│ Speaker state always fresh │ No │ Yes │
│ Untranscribed skipped │ Implicit │ Explicit │
│ CA mode bypass │ Yes │ Yes │
│ 500ms tolerance (E704) │ Yes │ Yes │
└────────────────────────────┴──────────────┴─────────────────┘
Expected count differences
On multi-party files with overlapping speech:
-
E701 count will be lower than CLAN’s error 83 count. CLAN fires error 83 on cross-speaker non-monotonicity; we don’t. The difference represents legitimate conversational overlap that we intentionally do not flag.
-
E704 count will be higher than CLAN’s error 133 count. CLAN’s early-return shadowing prevents error 133 from firing when error 83 fires, and the stale speaker state causes further suppression. Our correctly maintained per-speaker tracking reports all genuine self-overlaps.
On single-speaker files or files with minimal overlap, the counts should be very close or identical.
Implementation details
The implementation lives in
crates/talkbank-model/src/validation/temporal.rs.
Data flow
flowchart TD
A["collect_bullets(file)\n(temporal.rs:101)"] -->|"Vec<BulletInfo>"| B
B["validate_global_timeline()\n(temporal.rs:169)"] -->|"Per-speaker HashMap"| C["E701 errors"]
A -->|"Vec<BulletInfo>"| D
D["validate_speaker_timelines()\n(temporal.rs:212)"] -->|"Per-speaker HashMap"| E["E704 errors"]
BulletInfo
Each utterance with a bullet produces a BulletInfo containing:
utterance_idx: 0-based index in the filespeaker: the speaker code (e.g.,"CHI","PIL")bullet: theBulletstruct withstart_msandend_ms
Only main speaker tiers are collected. Dependent tiers (%mor, %gra, etc.)
are excluded. Possession of a collected bullet supplies the timing evidence;
there is no lexical-eligibility flag. Untranscribed speech (xxx, yyy, www)
still occupies time and constrains the same speaker’s following turn.
Per-speaker tracking
Both E701 and E704 use HashMap<&str, ...> keyed by speaker code:
- E701: stores
(utterance_idx, start_ms), the speaker’s most recent start time - E704: stores
(utterance_idx, end_ms), the speaker’s most recent end time
State is always updated after processing each bullet, regardless of whether an error was reported. This ensures clean tracking for subsequent comparisons.
CLAN source reference
For readers who want to trace the CLAN implementation:
- Function:
check_checkBulletsConsist()inOSX-CLAN/src/clan/check.cpp, lines 3849-3967 - Error 83: lines 3883-3890 (early
return(83)) - Error 133: lines 3892-3895 (only reached if error 83 did not fire)
- Speaker state update: line 3909 (
check_setLastTime), only reached if no error fired - Per-speaker tracking:
SPLISTlinked list, lookup viacheck_getLatTime()/check_setLastTime() - +c0 mode:
checkBulletsflag, set via+c0command-line option (line 5920), guards errors 84/85 at lines 3897 and 3953 - Call site:
check_ParseWords()line 4801, guarded byutterance->speaker[0] == '*'(main tiers only)
This page last changed: 2026-09-24 (commit 0089c3eb). The whole book last changed: 2026-10-07 (commit 5e895791).
CA Terminator Resolution
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
How CA markers are split between separators and linkers in the parser/model.
Current rule
The parser/model does not promote CA markers into utterance terminators.
The supported split is:
- Standard utterance terminators remain the CHAT terminators such as
.?!+...+/.and related final punctuation tokens. - CA intonation arrows (
⇗ ↗ → ↘ ⇘) staySeparatorcontent items. - CA TCU markers (
≈ ≋) staySeparatorcontent items. - CA TCU linker forms (
+≈ +≋) stayLinkeritems.
This means a trailing →, ≈, or ≋ remains in main-tier content rather
than being retyped as Terminator.
Parser/model consequences
- Tree-sitter grammar keeps arrows and
≈/≋on theseparatorpath. - The tree parser converts those nodes directly into
Separatorvariants. - The re2c parser classifies
≈/≋as separators and+≈/+≋as linkers. - No post-hoc pass promotes CA markers to terminators.
Terminator::try_from_chat_str()intentionally rejects CA arrows,≈,≋,+≈, and+≋.
Data Model
The active surface split is:
| Kind | CHAT tokens |
|---|---|
Terminator | . ? ! +... +/. +//. +/? +!? +"/. +". +//? +..? +. |
Separator | ⇗ ↗ → ↘ ⇘ ≈ ≋ plus the other CA/content separators |
Linker | +≈ +≋ plus the other utterance linkers |
Legacy CA-only Terminator variants still exist in the type for backward
compatibility with older serialized data, but new parser/classifier code does
not construct them from CHAT text.
Regression coverage
The regression surface for this split is:
ca_symbols_are_not_chat_terminatorsintalkbank-modeltrailing_ca_arrow_stays_separatorintalkbank-parsertrailing_ca_no_break_stays_separatorintalkbank-parsertrailing_ca_technical_break_stays_separatorintalkbank-parser
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Validation Cache
Status: Current Last modified: 2026-10-02 (commit 5c414cf6)
The persistent CHAT validation cache, used by chatter validate and the
desktop validation runner. The LSP maintains its own in-memory document cache. Distinct from the audio-task cache used by upstream
batchalign3 for FA / UTR ASR / media conversion (documented
separately in that project): this cache stores parse + validate
results keyed by file path + options.
crates/talkbank-cache/.
Architecture
flowchart TD
read["Worker reads the file once"]
hash["ContentHash::of(bytes read)"]
key["Cache key\n(path/parser namespace + RulesVersion\n+ AlignmentValidation + ContentHash)"]
db["SQLite WAL\n~/.cache/talkbank-chat/\ntalkbank-cache.db"]
hit["CacheLookup::Hit(outcome)\n→ serve a Valid verdict"]
miss["CacheLookup::Miss\n→ validate the bytes read, store for that hash"]
err["Err(CacheError)\n→ validate without the cache,\ncount it in cache_errors"]
read --> hash --> key --> db
db -->|"row for this version, coverage and content"| hit
db -->|"no row, rules changed, or other content"| miss
db -->|"database failed"| err
miss --> db
The worker hashes the bytes it is about to validate and passes that hash to
every lookup and write, so a verdict is always stored for the content that
was validated, even if the file changes during the run; the cache never
re-reads the path. A lookup returns Result<CacheLookup<V>, CacheError>,
V being the verdict’s own type: CacheOutcome { Valid, Invalid } for
validation and RoundtripOutcome { Passed, Failed } for roundtrip, so a
failed roundtrip cannot be read as an invalid file. A locked or corrupt
database is an Err, which the runner counts in the run’s
cache_errors and reports (stderr, the JSON summary, the desktop summary),
and the file is then validated without the cache. It is never a miss.
The methods of VerdictReader (get, get_roundtrip) and
ValidationCache (set, set_roundtrip) are all required. A cache that
keeps no roundtrip verdicts answers Miss and stores nothing in its own body,
rather than inheriting a silent default.
Configuration
| Config | Value | Why |
|---|---|---|
| Backend | SQLite via sqlx | Concurrent reads (WAL), atomic writes, zero-config |
| Pool size | 16 connections | Matches validation worker count |
mmap | 256 MB | Fast random access for 95k+ entries |
| Invalidation | Rules-version field + content hash + 30-day TTL | Rule-set or schema changes auto-invalidate; content edits invalidate per-file; stale entries pruned |
| Reachability prune | On validation open: keep the opening version plus one predecessor | Rows under any other version can never be bound again; without this the file grew by a corpus per release |
| Bridge | Embedded single-threaded tokio runtime, entered only via blocking::block_on | Sync workers block on async SQLite. Never Runtime::block_on directly: a caller that is itself driving a runtime (a Tauri async fn command) would nest one runtime in another and panic. Such a call is run on a thread with no ambient runtime instead |
| Init serialization | Advisory file lock (talkbank-cache.init.lock) | Exactly one opener performs first-time create + migrate; see below |
Schema
file_cache table (see
crates/talkbank-cache/migrations/20260101000000_initial.sql):
| Column | Role |
|---|---|
path_hash | The CacheKey: a blake3 hash of the path and its namespace (validation:tree-sitter or validation:re2c for validation, the parser’s label for roundtrip), as 64 hex digits |
file_path | The row’s ResolvedPath as text, indexed for cache clear --prefix and the missing-file purge; keys never read it |
content_hash | Hash of the file content; mismatch invalidates the entry |
version | Cache-compatibility version (RulesVersion): the cache crate version folded together with a fingerprint of the active validation rule set. A mismatch invalidates the entry |
cached_at | Insertion timestamp |
check_alignment | The AlignmentValidation coverage, as 0 (Structure) or 1 (IncludeTierAlignment); the only place it is a number |
is_valid | Cached validation outcome (0/1) |
roundtrip_tested | Whether roundtrip equivalence was checked |
roundtrip_passed | Roundtrip result when tested |
parser_kind | Roundtrip backend discriminator; NULL for validation, whose parser is in path_hash |
Validation uses a partial unique index on (path_hash, version, check_alignment)
where parser_kind IS NULL; roundtrip uses a second partial index including
parser_kind where it is non-NULL. file_path remains a maintenance index.
The path: one ResolvedPath, however it was spelled
Every cache method takes a ResolvedPath (talkbank_model::resolved_path,
re-exported by the cache): the file’s parent directory resolved by the
operating system (std::fs::canonicalize: links, .. and /tmp versus
/private/tmp resolved) joined with the file’s own name as stored. The last
component is never resolved: a link to a transcript is a transcript under
its own name, which validation compares with @Media. Its only constructors
read the filesystem, so no caller can mint one from a spelling of its own.
Resolution uses the operating system’s path rules, not a portable POSIX
interpretation. Windows normalizes .. in ordinary paths during native
absolute-path conversion; POSIX retains it and refuses traversal through a
missing parent. Windows canonicalization can also produce an extended-length
path prefix. Cache identities and test doubles retain ResolvedPath, rather
than comparing its native spelling against an unresolved argument. These
types enforce admission and identity flow; platform-specific filesystem
behavior still requires native-platform tests.
A StoredTranscript makes its ResolvedPath when it is admitted (by the
walk that found it, from one resolution per directory, or by the argument
resolver, from one per parent), so the validation worker keys by it,
validate --force clears by it, watch keys by it, and the argument
expansion counts a transcript once by it:
flowchart LR
arg["argument or walk\n(a.cha, ./a.cha, /tmp/c/a.cha, sub/../a.cha)"] --> stored["StoredTranscript\n(stored name, ResolvedPath)"]
stored --> key["CacheKey (lookups, stores)"]
stored --> clear["clear_paths: every namespace's key (--force)"]
stored --> dedup["one transcript per ResolvedPath"]
A directory that cannot be resolved makes its transcripts unreadable inputs
when they are found, the one policy for every consumer. cache clear --prefix resolves its prefix into a ResolvedPrefix (ResolvedPrefix::of,
its own type, so a ResolvedPath always names a file): an existing
directory resolved whole, anything else as a file; a missing directory is
resolved through its deepest existing ancestor. --force clears by key, in
every namespace a row can be written under, so two non-UTF-8 names that
share a lossy display spelling cannot clear each other’s rows.
The key
path_hash is a CacheKey, and CacheKey::of(resolved_path, namespace) in
cache_utils.rs is the only way to make one; validation and roundtrip both
call it, with CacheIdentity::validation_namespace() or
roundtrip_namespace(). The hash is specified, so a key computed by one
build is the key every later build computes:
flowchart LR
path["ResolvedPath"] --> comps["components()"]
comps --> enc["per component: kind tag\n1 prefix, 2 root, 3 ., 4 .., 5 name\nthen for a prefix or name:\nlength (u64 LE) + bytes"]
ns["KeyNamespace label"] --> sep["0 byte, then\nlength (u64 LE) + bytes"]
enc --> h["blake3"]
sep --> h
h --> key["CacheKey\n(64 hex digits)"]
A name’s bytes are the raw OsStr bytes on Unix and the UTF-16 code units
as little-endian bytes on Windows. The length prefixes keep ab + c and
a + bc apart. The path is already resolved, so every spelling of a file
is one path; hashing its components, not the display string, keeps native path equality
beyond that (Windows / and \ share a key) and keeps distinct non-Unicode
names distinct. Symlinks are not resolved and filename Unicode is not
normalized here; the stored-name resolution before it settles which name a
transcript has. Unit tests pin one key’s exact digits on Unix and one on
Windows, so a change to either encoding fails.
The key is a specified hash (blake3), not std’s DefaultHasher, whose
algorithm Rust does not promise to keep: every repository tracks stable Rust,
and a toolchain update must not be able to turn the persistent cache into
misses with nothing to say so.
The key scheme is part of the version. rules_version::KEY_SCHEME
(+keys.blake3-components-1) is folded into every RulesVersion, so a change
to CacheKey’s encoding is a version change: rows written under another
scheme carry another version, are never served, and leave the database
through the same reachability prune as any superseded rule set (kept for one
generation, then deleted on open; the 30-day expiry removes them in any case).
Rows are retired deliberately rather than lingering as unexplained misses.
What a run may do with the cache
A cache is one RunCache value: Absent, ReadOnly(Arc<dyn VerdictReader>) or ReadWrite(Arc<dyn ValidationCache>). The cache trait is
split: VerdictReader (identity, get, get_roundtrip) and
ValidationCache: VerdictReader (set, set_roundtrip), so a read-only run
holds a value with no write method, and whether a run may write is one value
rather than a mode beside an optional cache that could disagree with it.
Every runner entry point (validate_directory_streaming,
validate_files_streaming, validate_arguments_streaming) takes a
ValidationRun: the run’s ValidationConfig bound to its RunCache. The
only routes to one are ValidationRun::uncached(config) and
ValidationRun::new(config, cache), which refuses (CacheIdentityMismatch)
a cache whose identity() is not config.cache_identity(). A cache opened
for one rule set (strict linkers, another parser) therefore cannot serve its
verdicts to a run under another, and the bound configuration cannot be
changed afterwards. The CLI opens the cache from the run’s configuration and
binds it (initialize_validation_cache returns the bound run); the desktop
binds the pool it memoized for the request’s identity.
The CLI decides the mode once, as a CachePolicy:
flowchart LR
pres["presentation"] --> pol["CachePolicy::for_run"]
force["--force"] --> pol
pol -->|"Audit, no --force"| ro["ReadOnly\nReadOnlyCache::open: an existing,\ncurrent cache only; creates, migrates,\nprunes, clears and writes nothing"]
pol -->|"Audit + --force"| err["usage error (exit 2)"]
pol -->|"anything else"| rw["ReadWrite { refresh }\nCachePool::new: prune reported;\n--force clears the run's files"]
ro --> run["RunCache"]
rw --> run
ReadOnlyCache (CachePool<ReadOnlyScope>) is the read-only typestate: it
implements VerdictReader, has count and stats, and has no clear,
clear_paths or purge_nonexistent (those exist only for scopes that may
delete). Opening it writes nothing: it creates no directory, lock file or
database, runs no migration, and connects read-only. A cache that does not
exist (CacheError::NoCacheDatabase), whose schema is older than this
build’s (CacheError::SchemaNotCurrent) or that a newer build wrote
(CacheError::SchemaNewer) is refused, and the audit runs without a cache and
says so; a writing run creates or upgrades an older one. SQLite may create
its write-ahead-log index files beside an existing database, which any
reader of it needs.
How each file used the cache is a CacheUse (Hit, Miss,
NotConsulted) on its FileCompleteEvent, a fact separate from its
FileStatus. Hit means the verdict that decided the status came from the
cache: a cached Valid verdict, or a cached roundtrip verdict (also after a
fresh validation). A run with no cache consulted nothing and counts no misses,
and cache_hit_rate() is None when nothing was consulted.
Identity and handle states
Every cache scope offers close(self): a consuming transition that waits for
the pool’s connections and SQLite workers to shut down. Ordinary drop may
leave background cleanup in flight. Close all handles you own before removing
their database; the consumed handle cannot be queried again. This does not
close another pool or process’s handles or establish exclusive ownership of a
filesystem path. Fresh inspection, not the closed handle, admits the resulting
directory state; filesystem operations remain fallible.
Schema inspection transfers an open capability only for a current schema. Older-schema observations and schema-read refusals await shutdown before returning, so an observation with no handle does not leave its inspection workers behind. Failed migration or validation-scope admission likewise closes the unretained pool before returning its error.
CachePool::new(identity) returns Result<CachePool, CacheError>. Callers handle
that result before wrapping a successful pool in Arc. The CLI keeps the concrete
opening error until presentation. A failed cache open leaves validation active
and produces a structured warning in JSON mode, without writing prose to stderr.
The desktop app sends it to the frontend as a cacheUnavailable event before
the run’s own events, and the run’s summary says it.
CacheIdentity owns a RulesVersion and the shared ParserKind vocabulary.
ValidationConfig::cache_identity() derives both from the request’s semantic
configuration, excluding suppression and display policy. Every validation-cache
constructor requires this identity. Validation and roundtrip operations use the
bound parser, so an independent string argument cannot select another backend.
The desktop memoizes by this complete identity, including the parser toggle.
Parser choice is in the row namespace, not the retained generation. Rotating default/strict rules across both parsers therefore keeps four configurations inside the two-generation window. The CLI regression reproduces a real cross-backend cache hit; desktop and SQLite regressions verify separate hits, contradictory stored verdicts, and repeated rotations without eviction.
Administration starts from one look at the directory, CacheOnDisk::inspect,
which creates, migrates and writes nothing. What it finds is the one value a
preview and the operation it previews both start from, so the two cannot
disagree about the schema:
flowchart LR
dir["cache directory"] -->|"CacheOnDisk::inspect<br/>(writes nothing)"| found{"CacheOnDisk"}
found -->|"no database file"| absent["Absent(NoDatabase)"]
found -->|"ledger older than this build's,<br/>or no ledger"| older["OlderSchema"]
found -->|"ledger is this build's"| current["Current(InspectionCache)<br/>read-only: count, stats"]
found -->|"ledger past every migration<br/>this build knows"| newer["CacheError::SchemaNewer"]
older -->|"migrate()"| maint["MaintenanceCache<br/>count, stats, clear, purge"]
current -->|"into_maintenance()"| maint
InspectionCache (CachePool<InspectionScope>) is the read-only view of the
whole cache: count and stats, no verdicts (it has no identity) and no
clear. It is connected exactly as ReadOnlyCache is, through the one
read-only look: an existing database of this build’s schema, connected
read-only, with nothing created, migrated or pruned (SQLite may still create
its -shm and -wal files beside it). ReadOnlyCache maps the two states it
cannot read to CacheError::NoCacheDatabase and
CacheError::SchemaNotCurrent.
MaintenanceCache is a distinct CachePool state. It can read statistics and
perform explicit clear/purge operations but has no validation/roundtrip
methods. It has no public constructor: it is reached only from what the look
found, by OlderSchema::migrate (which migrates under the initialization
lock) or InspectionCache::into_maintenance (which reopens a current
database writable), both without automatic expiration or generation pruning.
A directory with no database has no route to it, so maintenance never creates
a cache to act on. Reading statistics records no generation of its own, so it
cannot displace one of the real retained generations.
chatter cache stats and chatter cache clear match on the found value:
| Found | stats | clear --dry-run | clear |
|---|---|---|---|
Absent | “No cache database at PATH”, exit 0 | would clear 0, exit 0 | cleared 0, creates nothing, exit 0 |
OlderSchema | says so, migrates nothing, exit 0 | says it would migrate first; no count, since migrating can remove rows | migrates, then clears |
Current | the entry count and file facts | counts read-only | reopens writable, clears |
SchemaNewer | exit 1 | exit 1 | exit 1 |
Rows under the unqualified validation suffix, which names no parser, are
never served through the parser-qualified namespace and remain eligible for
normal age/generation cleanup. No migration rewrites such a row to claim a
parser identity it does not record.
Concurrent initialization
Multiple chatter processes (or test processes) can open the same cache
directory simultaneously. Steady-state reads and writes are serialized by
SQLite itself (WAL journal mode plus a busy_timeout on every connection),
but the one-time first-open of a FRESH database is not: sqlx’s SQLite
migrator has no cross-connection lock (its Migrate::lock is a no-op for
SQLite), so two openers racing an empty database would both apply migration
version 1 and the loser would fail with UNIQUE constraint failed: _sqlx_migrations.version; concurrent first-connection WAL setup can collide
the same way.
The cache therefore serializes initialization explicitly:
sequenceDiagram
participant A as "Opener A\n(CachePool::with_directory)"
participant L as "Lockfile\n(talkbank-cache.init.lock)"
participant D as "SQLite db\n(talkbank-cache.db)"
participant B as "Opener B\n(CachePool::with_directory)"
A->>L: try_lock (exclusive) succeeds
B->>L: try_lock fails, bounded poll wait
A->>D: create + WAL setup + migrate
A->>L: unlock (drop InitLock)
B->>L: try_lock succeeds
B->>D: connect, migrator sees applied versions, no-ops
B->>L: unlock
- The lock (
InitLockincrates/talkbank-cache/src/init_lock.rs) is an exclusive advisory file lock (stdFile::try_lock:flock(2)on Unix,LockFileExon Windows) ontalkbank-cache.init.lockbeside the database. It is held only across pool connect + migrate, never across cache operation, so steady-state concurrency is unchanged. - Acquisition is a bounded try-lock poll, not a blocking OS wait: if the
deadline (10 s) expires, opening fails with the typed
CacheError::InitLockTimeoutinstead of hanging, and callers such as the CLI degrade to running uncached. Cache initialization can never block a caller indefinitely. - The OS releases the lock when the holder’s handle closes, including on crash, so a dead initializer cannot strand the lock.
- A bounded retry inside the pool-open path is retained as a backstop for
openers that do not honor the lock protocol (for example an older
chatterbuild sharing the same cache directory): once any winner has migrated the database, a re-attempt connects to a ready database and the migrator no-ops.
Regression coverage: tests/concurrent_open.rs (many threads, one
process) and tests/concurrent_process_open.rs (many processes racing one
fresh directory, with a hard deadline so a wedge fails instead of hanging
the suite).
What the cached value means, and what does NOT key it
A row records ONE fact: this file produced no diagnostics at all under this rule selection. That is a property of the bytes and the rules, so it is the same answer for every run, whatever any given run chooses to display.
Only RuleSelection therefore reaches the key
(RulesVersion::current_with_rule_selection). A PresentationPolicy
(--suppress, severity remapping) never does: it is applied to diagnostics that
have already been computed and have already decided what gets cached.
This is a fact of the crate graph, not a convention: talkbank-transform (home
of PresentationPolicy) depends on talkbank-cache, so the cache crate cannot
name the type, and folding one into the key is a dependency cycle rather than a
judgement call. A suppression set in the key would make chatter validate
followed by chatter validate --suppress xphon re-validate a whole corpus
from cold.
Only a clean file skips work, and that asymmetry is deliberate
A cache hit on a VALID file skips the parse entirely: the row says the file produced no diagnostics, and “no diagnostics” is the whole of what a caller needs, so there is nothing left to reconstruct.
A file recorded as INVALID is re-parsed and re-validated on every run
(worker.rs, the CacheOutcome::Valid arm is the only one that short-circuits).
The row stores one bit, not the diagnostics, so the bit alone cannot produce the
codes, spans, source snippets, or suggestions the user actually asked for. The
cache can say THAT a file failed; only a real run can say HOW.
This is intended, and it should not be “fixed” by caching diagnostics. The reasons, in order of weight:
- A diagnostic is not a fact about the file alone. It carries spans into the file’s bytes and rendered source context, so a cached diagnostic is only valid against the exact bytes that produced it. That is already what the content hash guarantees, but it makes the cached value large and structured rather than one bit, and every change to a message, a span, or a suggestion silently invalidates a store that has no way to know it.
- The bit is the part that is stable across releases; the rendering is not. Diagnostics are deliberately improved release to release. A cache keyed on the rule selection correctly serves the verdict across such a change, but would serve STALE TEXT for the same key, which is worse than slow: a user would see last release’s wording and last release’s suggestion.
- The asymmetry costs nothing on a healthy corpus and self-corrects. The kept corpus is ~106,000 files with ~141 invalid, so re-validation touches 0.1% of the work; a full warm run is about 6 seconds. As files get fixed they move into the fast path on their own.
The cost is real only where MOST files are invalid, which is the case during a cleanup campaign or when a rule has just been tightened. If that ever needs to be fast, the answer is not to cache diagnostics but to make the invalid path cheaper, or to give the campaign its own narrower target than the whole corpus.
When measuring cache behaviour, do not build a synthetic corpus by copying
files under new names. Renaming breaks the @Media filename check (E531), so
the copies validate as INVALID, and a benchmark built that way measures the
re-validation path while appearing to measure the hit path. Measured on a real
subtree the difference is stark: 9,263 real files take 29.0 s cold and 0.5 s
warm at a 100% hit rate, while the same files flattened under generated names
report a 28% hit rate and a warm run barely faster than cold. Use a real corpus
subtree; scripts/debug/chatter_validate_scaling.sh in the operator workspace
documents this and the sorted-file-list trap beside it.
Reachability pruning
Deleting by AGE and deleting by REACHABILITY are different questions, and the cache answers both on open.
The 30-day TTL removes rows that are stale, not rows that are merely unreachable; without a second rule every release would strand a complete copy of the corpus under its retired version, which no reader could ever bind.
Opening deletes every row whose version is outside a two-generation window:
- the version the pool binds, and
- the most recently written OTHER version.
The predecessor is kept deliberately. Pruning strictly to the current version makes a downgrade cold, which is a real cost during a bisect or a rollback, and it would make two chatter builds sharing a machine delete each other’s rows on every open. One generation of grace bounds the file at about two copies of the corpus while keeping both of those cases cheap.
When rows are deleted the database is rewritten (VACUUM) so the space returns
to the filesystem: SQLite otherwise frees pages for reuse without shrinking the
file, and an operator checking with du would reasonably conclude nothing
happened. A rewrite blocked by another process is not an error (the rows are
gone either way); the pages stay reusable and the next quiet open rewrites.
The outcome is reported (CachePool::version_prune) rather than logged from
inside the library, and chatter validate prints it: reclaiming most of a
user’s cache file in silence is indistinguishable from a bug.
Database location
| Platform | Path |
|---|---|
| macOS | ~/Library/Caches/talkbank-chat/talkbank-cache.db |
| Linux | ~/.cache/talkbank-chat/talkbank-cache.db |
| Windows | %LocalAppData%\talkbank-chat\talkbank-cache.db |
Statistics
CachePool::stats (and so InspectionCache::stats and
MaintenanceCache::stats) returns a CacheStats:
the entry count and a StorageStats that the cache crate reads itself, so a
caller renders it and never re-derives the database file from the directory.
flowchart LR
open["CachePool opened"] --> storage{"CacheStorage<br/>(private, fixed at open)"}
storage -->|InMemory| inmem["StorageStats::InMemory"]
storage -->|"Directory(dir)"| read["DatabaseFile::read(cache_db_path(dir))"]
read -->|NotFound| missing["Directory { database: Missing }"]
read -->|metadata| present["Directory { database: Present { size_bytes, modified } }"]
read -->|"other I/O error"| err["CacheError::Io"]
- In memory is a variant, not an absent directory, and a missing file is a
variant, not a zero size. Only a database that exists is opened, so
Missingmeans the file was removed after the cache opened. modifiedis ajiff::Timestamp, never optional, admitted when the file is read: a platform that keeps no modification time is aCacheError::Io(none of the release platforms is one), and a time outside the years -9999 to 9999 isCacheError::ModifiedOutOfRange. A renderer therefore formats it with no failure case of its own;chatter cache statsrenders the report directly, with no converted copy of it.CacheStatsand theDirectoryandPresentvariants are#[non_exhaustive], so no other crate can assemble a report from raw parts; tests outside the crate get one from a real cache.- Maintenance takes a
CacheScope:All, orCacheScope::Under(prefix), aResolvedPrefix(built byResolvedPrefix::of, which resolves the directory whole), for a path and everything under it by whole components.count(&scope)reads the database alone, sochatter cache clear --dry-runcannot fail on the database file’s metadata;clear(&scope)returns the rows its own delete removed. TheWHEREclause for a scope is written in one place.chatter cache cleargets its scope and mode from clap:ClearScopeandClearModeimplement clap’sArgs(cli/args/cache_clear_args.rs), so--prefix(any path the system accepts, UTF-8 or not) is resolved where it is parsed and the command matches values rather than re-deciding flags. - The reachability prune measures the file around its
VACUUMwith the sameDatabaseFile::read, from the pool’sCacheStorage; a size it cannot read isSpaceReclaimed::VacuumedSizeUnknown, never a reported 0.
Invalidation
-
Validation-rule changes: the
versioncolumn holds aRulesVersion, which folds thetalkbank-cachecrate version together with a fingerprint of the active validation rule set (an FNV-1a hash over everyErrorCodethe validator can emit, viatalkbank_model::validation_rules_fingerprint). Adding, removing, or renaming a rule (for example introducing error code E370, “retrace marker must be followed by material”) changes the fingerprint, hence theRulesVersion, hence the lookup key, so verdicts cached under the old rule set become a cache MISS and are re-validated instead of served stale. This is the mechanism that keepschatter validate(the authority on CHAT validity) from returning a stale “Valid” after the rules tighten.Rows under a superseded version are then UNREACHABLE: no query any binary can issue will match them again. Opening the cache deletes them (see “Reachability pruning” below), keeping one predecessor generation.
-
Content changes: each entry stores the file’s
content_hash; a mismatch is a per-file miss. -
Time-based: entries older than 30 days are pruned.
-
Reachability: rows under versions outside the two-generation window are deleted on open (see above). This is about unbounded growth, not correctness: those rows were already invisible.
-
Manual: pass
--forceto bypass cache lookups for a particular validation run.
Per repository policy, do not delete the cache directory without explicit
request. Use --force when you want fresh validation for specific paths
without destroying the whole cache.
See also
- Upstream
batchalign3documents its own audio-task cache for FA / UTR ASR / media conversion.
Parser implementation changes
Production cache generations include the grammar fingerprint and complete
source fingerprints from both parser crates, including recovery, conversion,
and the authored and vendored re2c lexer. The build-only talkbank-build
helper hashes sorted relative paths and exact bytes inside each owning package;
it never searches for a sibling checkout. The same helper fingerprints the
model source tree. Unreadable entries and symbolic links fail the build.
parser_behavior_fingerprint() composes both backends into one generation.
Parser selection still separates rows inside that generation, preserving the
two-generation retention budget while switching backends. The legacy
GRAMMAR_FINGERPRINT re-export describes grammar changes only and is not the
production cache identity. A parser-only source edit causes a cold miss
without requiring a package version bump.
These are conservative source fingerprints, not binary attestations. Comment and test-only edits invalidate too. Compiler, dependency resolution, feature flags, and runtime environment are not independently fingerprinted.
This page last changed: 2026-10-02 (commit 5c414cf6). The whole book last changed: 2026-10-07 (commit 5e895791).
Alignment
Status: Current Last modified: 2026-10-02 06:50 EDT
Alignment in the toolchain operates at two structural layers, plus a separate overlap-marker pass. Tier alignment is structural (counting and pairing AST nodes); word extraction is positional (domain-ordered token indices).
| Layer | Where | Purpose |
|---|---|---|
| Tier alignment | talkbank-model::alignment | 1:1 mapping between main tier and structural dependent tiers (%mor, %pho, %sin, %gra) |
| Word timing binding | talkbank-model::alignment | Count-matched positional convention between main-tier lexical slots and %wor timing observations |
| Word extraction | talkbank-transform::extract | Pull NLP-ready words from the AST in domain order |
Tier Alignment
Validates that dependent tiers have the correct number and arrangement
of items relative to the main tier. Lives in
crates/talkbank-model/src/alignment/.
TierDomain and PositionalDomain
#![allow(unused)]
fn main() {
enum TierDomain { Mor, Pho, Sin, Wor } // walks, descent, word membership
enum PositionalDomain { Mor, Pho, Sin } // counts and extraction
}
TierDomain is the vocabulary of the walkers and the membership rule
(counts_for_tier), and it has Wor. PositionalDomain is what a count or
an extraction takes (count_tier_positions, collect_tier_items,
TierCountable, AlignableTier::DOMAIN, extract_words), and it has no
Wor on purpose: the %wor count and pairing are
WorMainTierProjection’s (MainTier::wor_projection, then bind_timing
for the count and corroborate_wor_timing for the words), so the count and
extraction functions carry no second implementation of that count. The
overlap-marker position walk in alignment/helpers/overlap.rs is on
the shared walker at the %wor domain, the projection’s own leaf set.
PositionalDomain converts into TierDomain infallibly; the reverse is a
TryFrom that refuses Wor.
Internally, the shared positional traversal carries a static domain through
both top-level and bracketed content. Its emitted payload is domain-indexed:
%pho positions can contain a word, a phonological group, or a pause, never a
sign group or action. %mor has no atomic-group payload; %sin can carry sign
groups and top-level actions. These are producer guarantees, not filters in a
phonology-specific second walk. Runtime-domain public APIs dispatch to this
same traversal; their signatures and alignment policies are unchanged.
Leaf admission emits the domain-typed position directly to the sink. The shared
walk does not receive an optional payload and re-test admission at each leaf;
only the policy owner decides whether a separator, pause or action contributes.
The measuring-group policy selects its atomic payload before constructing a position. Mutable word traversal asks that same policy for the decision without an AST payload, while immutable word traversal projects its answer to entered content. Neither owns a second domain table. Existing spec/reference contracts cover the container/domain matrix, grouped diagnostic presentation and shared phonology indices; type-level exclusions do not replace those policy tests.
The same utterance produces different counts per membership domain:
| Rule | Mor | Pho | Sin | Wor |
|---|---|---|---|---|
| Skip retrace groups | Yes | No | No | No |
| Count pauses | No | Yes | No | No |
| PhoGroup | Recurse | Atomic (1) | Skip (0) | Recurse |
| SinGroup | Recurse | Skip (0) | Atomic (1) | Recurse |
Include fragments (&+) | No | Yes | Yes | No |
Include nonwords (&~) | No | Yes | Yes | No |
Include fillers (&-) | No | Yes | Yes | Yes |
| Include untranscribed | No | Yes | Yes | No |
| Include tag-marker separators | Yes | No | No | No |
ReplacedWord aligns to | Replacement | Original | Original | Original |
For the underlying word filter (counts_for_tier,
should_skip_group), the content walker, and the ChatFile model itself,
see CHAT Data Model. The walker plus the
domain table together govern every tier-alignment count.
Retrace handling, alignment-critical
Retraces are the most alignment-critical content type. A Retrace node
wraps content the speaker said then corrected.
- Mor: skip entirely (count
0). The retrace was a false start; only the correction carries morphological analysis. - Pho, Sin: recurse, words were physically produced and have phonological / gestural data.
- Wor: recurse, retrace ancestry does not change
%wormembership.
Critical invariant: the parser must emit UtteranceContent::Retrace
for all retrace patterns, including single-word retraces with
replacements (word [: repl] [* err] [//]). If a retrace is
accidentally emitted as a bare ReplacedWord, it counts for %mor
alignment, causing false E705 errors. Enforced by
tests/retrace_replaced_word_regression.rs. Full data model + parsing
pipeline + CHAT examples in
Retraces and Repetitions.
AlignmentPair
#![allow(unused)]
fn main() {
struct AlignmentPair {
source_index: Option<usize>,
target_index: Option<usize>,
}
}
Universal index-pair primitive. Some/Some = matched. One None =
insertion / deletion placeholder for mismatch diagnostics.
is_complete(), both indices Some. is_placeholder(), unmatched.
Per-domain results
| Type | Function | Source → Target |
|---|---|---|
MorAlignment | align_main_to_mor() | Main → %mor items |
PhoAlignment | align_main_to_pho() | Main → %pho tokens |
SinAlignment | align_main_to_sin() | Main → %sin tokens |
GraAlignment | align_mor_to_gra() | %mor chunks → %gra relations |
%gra aligns to %mor chunks, not items. Clitics create additional
chunks (pro|it~v|be&PRES = 2 chunks: pre-clitic + main).
Trait abstractions
| Trait | Purpose | Implementors |
|---|---|---|
IndexPair | source()/target() on any pair type | AlignmentPair, GraAlignmentPair |
TierAlignmentResult | pairs()/errors()/push_*() accumulator | Structural alignment result types |
AlignableTier | What a structural tier provides for generic alignment | PhoTier, SinTier |
TierCountable | count_tier_positions() / collect_tier_items() methods, over a PositionalDomain | [UtteranceContent] |
The generic positional_align() function uses AlignableTier to
eliminate duplication: align_main_to_pho() and align_main_to_sin() are
thin wrappers around it. %mor does not use it because it has additional
terminator validation logic. %gra does not use it because its source is
MorTier, not MainTier.
%wor is not validated
%wor is a timing-annotation sidecar, not a structural dependent tier.
validate_alignments() does not reject a %wor word-count mismatch.
Old corpus files may have xxx, fragments, or nonwords in %wor
(pre-2026-04 behavior) without producing false errors.
Consumers that need timings call bind_wor_timing(). Its typestate result is
one of Missing, Drifted, or CountMatched. A
CountMatchedWorTimings value exposes only the common count after equal counts have
been observed under the named FilteredLexicalV1 membership policy. Position
permits the next comparison; it does not yet expose timing. Callers must pass
that state to corroborate_wor_timing(), which compares the parsed %wor
display tokens with the canonical display sequence derived from the main tier.
Only CorroboratedWorTimings exposes positional slots. Each such slot takes
lexical identity from the main tier and timing from the corresponding %wor
word bullet. %wor text may refuse unsafe reuse but cannot supply lexical
identity. A present but untimed slot is WorSlotTiming::Unaligned; it is not
conflated with a missing tier or count drift.
MainTier::wor_projection() is the single owner of current Wor-domain
selection. Both %wor generation and timing binding travel through that typed
projection, so membership disagreement between two implementations cannot be
represented. See
%wor Timing Semantics for the complete contract and research
boundary.
Phon tier-to-tier alignment
A second class of alignment that operates between dependent tiers:
| Source | Target | Code |
|---|---|---|
%modsyl | %mod | E725 |
%phosyl | %pho | E726 |
%phoaln | %mod | E727 |
%phoaln | %pho | E728 |
Derived-view alignments: %modsyl is a syllabified reannotation of
%mod, %phosyl of %pho; those counts match directly. %phoaln
advances the two source tiers independently: a one-sided pause consumes a
slot only on the tier bearing it. Utterance-owned PhoalnWordBinding values
carry the alignment word and its bound source items. Both effective counts
and reconstruction checks consume those bindings; neither consumer indexes
the source tiers by raw alignment-word position. Missing source items remain
distinct from slots intentionally absent for an opposite-side pause.
compute_alignments() runs after main-tier alignment and may report E727
and E728 simultaneously. Ordinary non-pause insertions/deletions still consume
both word slots.
Known data issue: Phon XML source data has orthography↔IPA word
count discrepancies in ~4% of files (518 / 12,340). Expected in child
phonology data. A subset of existing corpus CHAT files handle this
inconsistently across tiers: %mod/%pho are truncated to match
orthography, one word to one word, but %xmodsyl/%xphosyl/%xphoaln
carry the full IPA word set, undropped. Result: E725-E728 mismatches.
As of Phon 4.0.0-beta.9 (2026-06-25), Phon reads and writes CHAT
natively; we have not seen output from that native export and do not
know whether it reproduces the inconsistency.
Parse-health gating
Alignment diagnostics honor ParseHealth metadata. If a dependent
tier’s domain is parse-tainted, mismatch errors for that domain pair
are suppressed. Main-tier taint blocks all main→dependent alignments.
Dependent-tier taint blocks only that tier. Phon tier-to-tier checks
have their own gates (can_align_modsyl_to_mod,
can_align_phosyl_to_pho, can_align_phoaln).
ParseHealthState::Unknown, including after JSON import, cannot authorize any
alignment. Adding one tier’s taint or all dependent-tier taint preserves
Unknown: evidence of damage cannot establish that the remaining tiers were
parsed cleanly. Tainting an already parser-backed state only withdraws trust;
it never enables an alignment that was previously unavailable.
Explicit checked construction
is a separate admission path for assembled typed documents. Its Constructed
state permits alignment checking but does not claim parser provenance or source
span authority. Adding recovery taint withdraws construction admission entirely;
it must never turn construction into partially clean parser evidence.
The canonical provenance workflow checks single-tier, dependent-only and whole-utterance trust withdrawal against parsed spec/reference models. Broad recovery must discard cached pairs and WOR timings, retain the semantic content and tier presence, and produce stable diagnostics on repeated recomputation. Dependent-only recovery leaves main-tier provenance clean; whole-utterance recovery warns about both damaged alignment sides. These are public API recovery transitions, not claims that the clean seed files contain those parse errors.
Word Extraction
extract_words() (in crates/talkbank-transform/src/extract.rs) uses
the content walker to pull words from the AST in domain-specific order,
over a PositionalDomain (%wor words are the projection’s).
Returns Vec<ExtractedUtterance>; each entry carries its speaker, utterance
index, and ordered words. Each ExtractedWord carries cleaned text,
raw_text, utterance_word_index, and form_type. Its governing language
mark is private: use language_kind() or resolve_language() so resolution
retains the position captured from the source during traversal. Tag-marker
separators (, „ ‡) are included as words in Mor because they have
%mor items (cm|cm, end|end, beg|beg).
Canonical replacement/retrace references verify the extracted sequences in all three domains. Mor uses replacement words and omits retraced material; Pho/Sin use the eligible spoken originals, including retraced words. Ordinary explanatory annotations do not create extra extracted words. Extraction is read-only and must preserve the original CHAT serialization.
Overlap Marker Iteration
CA overlap markers (⌈⌉⌊⌋) appear at three content levels,
UtteranceContent (top-level), BracketedItem (inside groups), and
WordContent (intra-word, butt⌈er⌉). One API in
talkbank-model/src/alignment/helpers/overlap.rs, on the shared
walk_content at the %wor domain, so its word positions are the %wor
projection’s slot indices.
extract_overlap_info, region-based
The marker stream owns pairing state: opening boundaries, available closing boundaries, and consumed closing boundaries are distinct variants. Pairing consumes only a later closing boundary of the same kind and index; there is no parallel bitmap whose state can drift from the markers. Unmatched openings and closings remain explicit in the result. Opening regions retain source order, followed by orphaned closing regions in source order. This is positional analysis, not a validity claim or an observed acoustic onset time.
Pairs markers by (kind, index) into OverlapRegion structs. Each
region represents a matched ⌈…⌉ or ⌊…⌋ pair. Index-aware:
⌈2...⌉2 forms a separate region from ⌈...⌉. Mismatched indices
leave markers unpaired. Onset-only ⌈ (without ⌉) is a legitimate CA
convention, region has end_at_word = None,
is_well_paired() = false, but top_onset_fraction() still works.
Cross-utterance, analyze_file_overlaps
For whole-file analysis, in overlap_groups.rs. 1:N matching: one
top region from speaker A can match multiple bottom regions from
speakers B, C, etc. Used by E347 and chatter debug overlap-audit.
Overlap validation
| Code | Level | Check |
|---|---|---|
| E347 | Cross-utterance | Orphaned tops/bottoms with 1:N matching (warning) |
| E348 | Utterance | Unpaired markers within a single utterance (warning) |
| E373 | Utterance | Invalid overlap index values (must be 2-9) |
| E704 | Cross-utterance | Same speaker encoding both top and bottom (error) |
chatter debug overlap-audit <path> reports per-file statistics
(groups, bottoms, orphans, temporal consistency) in TSV format. Use
--database <path.jsonl> for a persistent JSON-lines database.
Coordinated %mor / %gra replacement
MorTier::splice_range_coordinated (and the single-item
splice_coordinated) replace a contiguous range of %mor items and the
matching %gra relations in one atomic edit. Morphotag’s L2 pass uses it
once per @s span: the span’s words, reparsed in their own language,
replace the primary parse’s items, and the replacement can change chunk
counts (it's becomes it~'s).
The replacement is a SplicedBlock, built by SplicedBlock::new(mors, relations) from the items and one relation per chunk in block-relative form.
Building it is the only route to a block, and it parses the relations once:
head 0 becomes the span root and every other head a BlockChunk (the
block’s own 1-based numbering, a different space from the host’s
SemanticWordIndex1). A block is a tree with exactly one root; a count
mismatch, a relation out of chunk order, a head outside the block, no root or
two, or a cycle is a SplicedBlockError, so the splice never sees one.
SplicedBlock::root_chunk() returns that admitted root as a BlockChunk.
Callers can use it for an explicit host redirect without inspecting or
validating the block’s relations again. The root and relations are immutable
after admission.
The host is admitted first: each of its relations must carry its own chunk as
its index (relation k of the tier, from 1, has index k), because every
head is read as a chunk number; a host numbered otherwise is refused
(HostIndexOutOfOrder). Inside the splice, the numberings it moves between
are separate private types (host_chunks.rs beside the splice): a host chunk
before the splice (PreChunk), after it (PostChunk), a block chunk
(BlockChunk), and a chunk or item of the replaced range. Geometry::locate
sorts a pre-splice chunk into kept or replaced, and Geometry::translate,
which takes only a kept chunk, is the one route from the pre-splice numbering
to the post-splice one; Geometry::place is the one route from a block chunk.
A post-splice number is never made from a bare integer, so a pre-splice index
cannot be written into the result. The whole new %gra is built from shared
borrows before either tier is written, which is what makes a refusal atomic.
Four kinds of %gra head are rewritten, each by its own rule:
- A head inside the block becomes the block chunk placed after the host chunks before the range.
- The span root attaches where the caller’s
SpanRootsays:UtteranceRoot(head0, relationROOT), orHostChunk { chunk, relation }, a host chunk named by its index BEFORE the splice, which the splice translates like any host head (shifted when it lies after the range), with the relation the span root takes under it. That relation is anAttachmentRelation, which refuses a root label, soROOTunder a host head cannot be written.SpanRoot::from_gra_head(head, relation)reads the anchor off a host relation’sGraHeadRef. - A host head past the replaced range shifts by
new_chunks - old_chunks. - A host head INTO the replaced range depended on a word, so it must land on
the chunk of the block that stands for that word. Only the caller knows how
old items correspond to new ones, so it states that as a
HostRedirectsvalue, and the splice validates the statement against the admitted host range and the block before it changes anything.
flowchart TD
plan["HostRedirects"] --> by{"ByItem or PerItem?"}
by -->|"ByItem (equal item counts)"| counterpart["every item: ItemTarget::Counterpart"]
by -->|"PerItem(targets)"| each{"targets[k]"}
each -->|"Chunk(c)"| stated["every old chunk of item k -> block chunk c"]
each -->|Counterpart| counterpart
counterpart --> shape{"old item k and block item k\nhave the same chunk count?"}
shape -->|yes| chunkwise["old chunk j -> new chunk j"]
shape -->|no| head{"exactly one chunk of block item k\nheaded outside it?"}
head -->|yes| headchunk["every old chunk -> that head chunk"]
head -->|no| ambiguous["no target: refused if a host\nrelation depends on item k"]
HostRedirects::ByItem refuses unequal item counts
(RedirectItemCountsDiffer). PerItem with the wrong number of targets, a
Chunk target outside the block, or a Counterpart for an item the block
does not have is refused (RedirectCountMismatch, RedirectOutOfBlock,
NoCounterpart). A dependent of an item with no unique head chunk is refused
(NoUniqueHeadChunk), but only when such a dependent exists. The validated
form, one target per replaced old chunk, is private to the splice and built
from the admitted host, so it cannot be validated against one range and
applied to another.
What the splice guarantees
The splice adds no cycle and no second root: if the host %gra was a tree,
the result is a tree. The block is a one-rooted tree by construction; the
span root attaches to the utterance’s root only when no host relation outside
the replaced range is already a root (UtteranceRootTaken), and to a host
chunk only when that chunk lies outside the range (SpanRootInReplacedRange),
within the host (SpanRootOutOfHost), and its own chain of heads does not
reach the range (SpanRootDependsOnSpan: with x@s y@s z . and z -> x,
the span x y cannot hang under z). Every refusal leaves both tiers
unchanged. The method’s rustdoc carries a worked L2 example: in dont@s:eng mal geh ., the span dont becomes do~n't, mal (a dependent of dont)
moves to the span root do, and the span root keeps depending on geh,
chunk 3 before the splice and 4 after it.
Design Principles
- No string hacking. All alignment operates on typed AST
structures (
Word,MorTier,AlignmentPair), never on serialized CHAT text. - Domain-aware from the start.
TierDomaingates traversal at the walker level. Downstream code never re-implements retrace / group skipping logic. - Deterministic over approximate. Tier alignment and word extraction use deterministic, positional algorithms over the typed AST.
- Dense indexed structures.
AlignmentPairusesOption<usize>rather than cloned data; index pairs are stored positionally, not in hash maps. - Exhaustive matching. Every
matchonUtteranceContent(24 variants) orBracketedItem(22 variants) lists all variants explicitly. New variants are a compile error, not a silent bug. - Walker as shared primitive.
walk_words()removed ~330 lines of duplicated traversal boilerplate across 7 call sites.
Downstream Consumers
| Consumer | Crate | Usage |
|---|---|---|
| Validation | talkbank-model | Cross-tier checks (E714/E715, E725-E728), overlap (E347/E348/E373/E704) |
| LSP hover | talkbank-lsp | Show aligned tier items for word under cursor |
| Word extraction | talkbank-transform | NLP-ready words from utterances |
| Overlap audit | chatter | chatter debug overlap-audit |
%wor generation | talkbank-model | Build %wor tier from main tier |
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
%wor Timing Semantics
Status: Current Last modified: 2026-10-07 (commit 5e895791)
Purpose
%wor is a timing sidecar over a named subset of main-tier word slots. It is
not an independent lexical transcript and it is not a structural dependent-tier
alignment like %mor, %pho, or %sin.
The main tier owns lexical identity. %wor contributes an optional inline
media bullet for each selected position. The visible word printed on %wor is
display material and optional corroborating evidence. It can prevent stale
timing reuse, but it never supplies lexical identity.
Generation and serialization are different operations. Newly generated %wor
words use the named cleaned-text display convention. Serializing an already
parsed tier preserves its typed word structure, including shortening, compounds
and markers, then writes its timing bullet. It must not silently regenerate a
parsed display token from cleaned text. The selective-name and pseudonymizer
source/expected references exercise this distinction through fragment,
round-trip, normalization and downstream transform contracts.
Actual timing presence is a separate typed question from correspondence.
WorTier::timing_evidence() returns Absent or a RecordedWorTiming carrying
the first real word-level bullet. Media validation uses this state directly.
Equal counts with no bullets are not timing evidence, while a real bullet
remains timing evidence even when counts drift. Alignment processing is not
required to observe a bullet already in the typed CHAT model. A %wor tier
cannot carry a trailing tier-level bullet: the grammar does not permit that
state, and WorTier cannot construct or serialize it.
Complete CHAT admission rejects a word bullet that ends before it starts
(E362). A zero-duration word bullet is legal, as are untimed words, and there is
no cross-word monotonicity, word-overlap or cross-speaker restriction (see the
leniency policy, Decision 10). A consumer that needs positive intervals, such as
assess_wor_timing_sequence, refuses them at its own boundary. Each word bullet
carries its own source span, so E362 is located inside its %wor tier.
The adaptive source-bound replacement plan (WordTimingPlan::PreferRetained)
discards only the word tiers that are actually wrong, and keeps the rest:
flowchart TD
L[Lower each %wor once] -->|own lowering failed| R1[Removed: OwnLowering]
L -->|clean| I[Move entry into the document]
I --> V{Validate the candidate}
V -->|valid| A[Admit; removed tiers decide the disposition]
V -->|errors inside word tiers| R2[Remove those tiers: LocatedValidation]
R2 --> V
V -->|errors in no word tier| R3[Remove every remaining tier: UnattributedValidation]
R3 --> V
V -->|rejected, no word tier left| F[Refuse: a retained fault]
A validation diagnostic belongs to a tier when its span lies inside that tier’s
source span; the tiers holding an error are removed with every diagnostic
inside them bound to the receipt (RemovedTier::cause), and the reduced
document is validated again. Errors located in no word tier remove every word
tier still retained, recorded as unattributed with the shared diagnostics; any
retained fault then still refuses. That fallback is the old remove-all
behaviour: when removal makes such an error go away, it shows only that the
error depended on word tiers, not which one, so the receipt says
“unattributed” rather than naming a cause. If a removed tier carried a recorded word
bullet, the next validation runs in the timing-regeneration phase, so a file
whose only timing was in removed tiers becomes a pending regeneration (E544
obligation), while one that keeps some timing is a complete replacement. A
preservation plan has no exemption of any kind.
Current membership policy
The canonical policy is FilteredLexicalV1. A typed
WorMainTierProjection is its single traversal owner. Both %wor generation
and timing binding consume that projection, so a policy edit cannot update one
path and leave the other behind.
| Main-tier content | Current membership |
|---|---|
| Regular word | Included |
Filler such as &-um | Included |
| Retraced regular word | Included |
| Original surface of a replacement | Included when otherwise eligible |
Phonological fragment such as &+w | Excluded |
Nonword such as &~gaga | Excluded |
xxx, yyy, or www | Excluded |
| Omission | Excluded |
| Separator or terminator | Excluded |
This policy is explicit because the meaning of one-to-one correspondence
depends on which main-tier items count. A future research policy must receive a
new name and separate evaluation. It must not silently change
FilteredLexicalV1.
Typed binding and correspondence states
Consumers call bind_wor_timing(main, wor). The data state is one of:
Missing: no%wortier exists. This is distinct from a present empty tier.Drifted: the selected main-tier slot count differs from the physical%worword-entry count. No positional slots are exposed.CountMatched: counts match under the named policy. This state permits a positional comparison but exposes no timing slots. Equal counts alone do not prove that a parsed legacy tier and the current main tier share origin.
Consumers pass CountMatched to corroborate_wor_timing. The next state is:
Uncorroborated: one or more%wordisplay tokens differ from the canonical display tokens generated from the current main-tier projection. The state exposes exhaustive mismatch diagnostics but no timing slots.Corroborated: every display token matches the canonical generated token at its count-matched position. Only this state exposes positional timing slots.
Each corroborated slot has:
- a borrowed typed main-tier
Wordand itscleaned_text, which remain the only lexical identity; - the exact corroborating
%worWord, exposed bywor_word()for source-bound display-token updates that preserve timing and metadata; Timed(WorRecordedInterval)when the corresponding%worentry has an inline bullet;Unalignedwhen the entry exists but has no inline bullet.
This transition detects same-count edits when they change at least one
canonical display token. It cannot establish immutable common origin: repeated
tokens can be exchanged invisibly, and CHAT does not carry a generation
identifier. The state is therefore named Corroborated, not Aligned or
Proven.
Selective transforms must finish corroboration before proposing %wor
changes. They derive replacement display text from the selected main-tier
word, not a second name search over %wor. Count drift or lexical drift
requires an explicit refusal/review path, not an independently guessed pairing.
The tiers/wor-drift.cha reference file witnesses both count-drift directions
and same-count lexical drift; these are stale sidecars in otherwise valid CHAT.
Temporal sequence transition
Lexical corroboration is necessary but not sufficient for algorithms that need a word-timing hull. A corroborated tier may still contain an untimed slot, a zero or backwards interval, or adjacent word intervals that overlap.
Consumers call assess_wor_timing_sequence(corroborated) for the next checked
transition. It returns one of:
Empty: the sidecar is present and corroborated, but the membership policy selected no words. There is no hull.Rejected: the binding contains one or moreUnalignedorNonPositiveIntervalissues. The diagnostic state exposes typed slot indices and numeric evidence, but no partial hull.Complete: every selected word has a positive interval. Only this state exposes borrowed main words paired with typed recorded intervals, a min/maxWorTimingHull, and every typed adjacency relation.
Each adjacency is Gap, Touching, Overlap, or BackwardStart. Overlap and
backwards starts remain visible evidence, but do not erase a hull that is still
mechanically defined by the recorded extrema. An algorithm that requires
common origin, acoustic accuracy, non-overlap, or monotonic starts must state
and enforce that later policy over additional evidence or the relation types.
The assessment transition is infallible because its control flow makes the
remaining construction failure unrepresentable. An empty binding returns
Empty. A nonempty binding assesses the first slot before it can construct a
complete accumulator. A first-slot failure starts a rejected accumulator; a
first complete slot is the required seed for the hull. No caller or internal
branch can construct a nonempty complete sequence without that seed.
WorSlotIndex, WorMediaOffsetMs, WorDurationMs, and WorTimingHull have
private constructors. A caller cannot mint an index for an unrelated tier,
present arithmetic as a recorded media coordinate, or label arbitrary offsets
as a hull that passed chatter’s sequence assessment. Recorded offsets and the
duration derived by subtraction remain different types.
The complete state does not return the original Bullet. Returning it would
reopen direct access to raw integer fields and let every consumer rebuild the
same loose arithmetic. WorRecordedInterval is the only timing surface after
binding.
This is deliberately stricter than ordinary CHAT validation. CHAT can retain legacy or partially aligned data. A timing-consuming algorithm needs an explicit admission contract and must not infer one from the fact that the file parsed.
The count types for the main sequence and the physical %wor sequence are
different newtypes. Callers cannot swap them accidentally. Constructors for
the proof states and counts are private.
The count-matched state is deliberately named for the fact it actually proves. It does not claim common origin and it cannot expose timing. The corroborated state is also deliberately limited: canonical display equality provides useful evidence against stale reuse, but serialized CHAT carries no immutable generation identity. Acoustic or common-origin qualification requires later evidence and a different state.
The binding borrows both typed sequences until correspondence is decided.
Corroboration retains main-tier words as lexical identity and copies the two
recorded media offsets into private-constructor coordinate types. It does not
clone a temporary generated %wor tier or reduce structured lexical identity
to a string.
The main-tier projection owns word and separator selection in source order.
Its constructor is private to MainTier::wor_projection(). Generation derives
the visible tier from this capability, and binding consumes the same capability
to obtain lexical slots. There is no independent counter or generator whose
agreement must be tested at runtime. Count drift between a parsed legacy tier
and its main tier remains a real Drifted data state.
flowchart LR
Main[Typed MainTier] --> Projection[WorMainTierProjection]
Projection --> Generated[Derived WorTier]
Projection --> Binding{Timing binding}
Parsed[Parsed legacy WorTier] --> Binding
Binding --> Missing
Binding --> Drifted
Binding --> CountMatched
CountMatched --> Correspondence{Canonical token correspondence}
Correspondence --> Uncorroborated
Correspondence --> Corroborated
Corroborated --> Sequence{Sequence assessment}
Sequence --> Empty
Sequence --> Rejected
Sequence --> Complete
Generation and parsing
For newly generated data, word timings are embedded on typed main-tier words.
MainTier::generate_wor_tier() derives the visible sidecar from those words.
It copies main-tier cleaned_text for display and copies the inline bullets for
timing.
For parsed legacy CHAT, the main tier and %wor are separate AST values. A
consumer must use the binding transition and then lexical corroboration before
recovering timing by position. Different display words produce an
Uncorroborated state. They do not replace main-tier lexical identity and do
not make the CHAT file invalid.
Validation versus evidence admission
A drifted %wor tier does not make a legacy CHAT file invalid. Editors may
change the main tier without immediately rerunning forced alignment, and older
corpora used different membership conventions.
Evidence-consuming algorithms have a stricter contract. They must refuse
Missing, Drifted, or Uncorroborated when their operation requires word
timing. They must also decide whether Unaligned slots, empty intervals,
nonmonotonic intervals, or timings outside the main bullet are acceptable for
that specific operation. Structural and lexical admission do not prove
acoustic accuracy.
Temporal completeness does not prove acoustic accuracy either. It establishes only coverage, positive duration, a min/max location hull, and explicit adjacency geometry. The aligner may still place a perfectly well-formed onset or offset too early or too late. Model score, boundary origin, human calibration, and downstream merge outcome belong to a later evidence layer.
Research boundary
The abandoned goal of timing every spoken main-tier item remains a legitimate research question. It is not a correction to current semantics until the membership question has been specified and evaluated. In particular, fragments, nonwords, untranscribed material, interactional sounds, retraces, and editorial replacements need explicit rules.
Confidence, acoustic quality, and provenance should remain typed internal
evidence attached to a binding or downstream decision. Chatter must not revive
public %xalign clutter as a side effect. A public %wor tier remains a
derived view unless TalkBank deliberately adopts a new visible format.
Alternative policies and acoustic qualification should be tested against immutable evidence artifacts before changing corpus output. The relevant questions include:
- whether the proposed policy reduces human correction time;
- whether every additional slot can receive defensible acoustic boundaries;
- whether changed timing improves the placement of manual transcripts merged onto automatic ones;
- whether confidence and provenance improve decisions without becoming public transcript clutter;
- whether a new policy can coexist with legacy
%wordata without ambiguous automatic reinterpretation.
Release boundary
Analysis projects that pin a released chatter tag must adopt the binding API
only after that chatter release is cut. They must not switch to a live path
dependency to test this code. Until then, they may reproduce current membership
with the released typed generator, but the first post-release change should
replace that local pairing with bind_wor_timing followed by
corroborate_wor_timing. Its regression evidence must show both that a
same-count token edit is refused and that %wor text never becomes lexical
authority.
Downstream code that currently derives a main-tier bullet by manually checking
every %wor word and taking the minimum and maximum timing should then use the
sequence transition, WorTimingHull, and typed adjacency relations. This
removes the repeated loose procedure while preserving the important rule that
one untimed word cannot claim a complete child-utterance span. Compatibility
with a downstream policy must be measured before replacement because the new
API reports overlap and backwards starts instead of silently ignoring them.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Memory and Ownership
Status: Current Last updated: 2026-03-24 01:32 EDT
This chapter documents the memory management and ownership patterns used across the TalkBank Rust crates. Understanding these decisions helps contributors make consistent choices when adding new code.
String Representation Strategy
CHAT corpora contain massive repetition, the same speaker codes, language codes, POS tags, and high-frequency words appear millions of times across files. The codebase uses three string types, chosen by expected cardinality and duplication:
flowchart LR
raw["Raw input (&str)"]
smol["SmolStr\n(inline ≤23 bytes)"]
arc["Arc<str>\n(interned, deduplicated)"]
string["String\n(owned, unique)"]
raw -->|short, low repetition| smol
raw -->|high repetition domain value| arc
raw -->|ephemeral/unique| string
| Type | When to use | Examples |
|---|---|---|
SmolStr | Short tokens, low duplication | Postcode text, tier content, event labels |
Arc<str> (interned) | High-cardinality domain symbols | Speaker codes, language codes, POS tags, stems |
String | Ephemeral or unique values | Error messages, temporary formatting |
String Interning
Location: talkbank-model/src/model/intern.rs
Five global process-local interners, each a DashMap<Arc<str>, Arc<str>> behind
OnceLock<StringInterner>:
| Interner | Pre-seeded values | Typical savings |
|---|---|---|
speaker_interner() | 30+ codes (CHI, MOT, FAT, …) | High, 3-letter codes repeat per utterance |
language_interner() | 45+ ISO 639-3 codes | Moderate, per-file |
pos_interner() | 60+ POS tags + UD relations | Very high, every %mor word |
stem_interner() | 200+ frequent English stems | High, function words dominate |
participant_interner() | 14 roles (Target_Child, …) | Low, per-file |
How it works:
- Fast path:
get()on DashMap, O(1)Arc::cloneif found - Slow path:
insert()new Arc if miss, deduplicates on future access - Thread-safe: DashMap uses shard-level locks, no global contention
- After initialization, reads are lock-free
Memory impact: 50-200 MB savings on large corpora (5-20% reduction). Arc::clone
is O(1) atomic increment vs String::clone O(n) copy.
Newtype Macros
Two macros generate domain-typed string wrappers:
string_newtype!: wrapsSmolStr. Used for generic CHAT text.interned_newtype!: wrapsArc<str>with automatic interning. Used for domain symbols.
// SmolStr-backed: no interning, inline small strings
string_newtype!(PostcodeText);
// Arc<str>-backed: interned via global interner
interned_newtype!(SpeakerCode, speaker_interner);
Ownership Model
ChatFile Lifecycle
flowchart TD
src["Source text (&str)"]
cst["tree-sitter CST\n(Tree, borrowed nodes)"]
model["ChatFile\n(owned AST)"]
cache["SQLite cache\n(validation result)"]
lsp["LSP server\n(per-document state)"]
json["JSON output\n(serde serialization)"]
cli["CLI output\n(CHAT text)"]
src -->|tree-sitter parse| cst
cst -->|CST-to-model conversion| model
model -->|validate + hash| cache
model -->|held in backend| lsp
model -->|to_json()| json
model -->|to_chat_string()| cli
- Parsing: tree-sitter
Treeowns the CST.Node<'a>values borrow fromTree, zero-copy traversal. The CST-to-model conversion copies data into ownedChatFilefields (SmolStr,Arc<str>). TheTreeis dropped after conversion. - Validation:
ChatFileis borrowed (&self) during validation. Errors are streamed to anErrorSink, no accumulation required. - LSP: Each open document holds an owned
ChatFilein the backend. Re-parsed on every edit via tree-sitter incremental parsing. - CLI batch: Each file is independently parsed → validated → reported → dropped. No cross-file state except the shared cache.
Arc Usage
Arc appears in three distinct roles:
| Role | Type | Why |
|---|---|---|
| String interning | Arc<str> in model types | O(1) clone for high-repetition domain values |
| Worker pool | Arc<WorkerGroup> in batchalign | RAII CheckedOutWorker::drop() needs group reference to return worker |
| Cache backend | Arc<dyn CacheBackend> in batchalign | Shared across async request handlers |
No Rc (single-threaded sharing not needed). No Cow<str> (SmolStr covers the
inline-small-string use case more naturally).
Interior Mutability
| Pattern | Where | What it protects |
|---|---|---|
RefCell<Parser> inside TreeSitterParser | talkbank-parser | Tree-sitter Parser needs &mut self but isn’t Sync. Callers create a TreeSitterParser and pass &TreeSitterParser everywhere. |
DashMap<Arc<str>, Arc<str>> | String interners | Concurrent interning during parallel parsing. Shard-level locks. |
OnceLock<StringInterner> | 5 global interners | Lazy init, lock-free after first access |
LazyLock<Regex> | All regex patterns workspace-wide | Compile-once, no per-call overhead |
std::sync::Mutex<VecDeque> | batchalign worker idle queue | Held < 10 μs for push/pop only |
tokio::sync::Mutex<HashMap> | batchalign job store | Short reads/writes, never held across .await |
Semaphore | Worker availability (batchalign) | Async signaling without holding locks during dispatch |
Rule: std::sync::Mutex for data accessed from sync code or held briefly.
tokio::sync::Mutex only when the lock must be held across .await points (which
we avoid when possible). DashMap when many threads read concurrently.
Collection Choices
| Collection | Where | Why not HashMap/Vec |
|---|---|---|
BTreeMap | All test/snapshot JSON output | Deterministic key ordering for reviewable diffs |
IndexMap | Participants, per-speaker results | Preserves encounter order (CHAT spec requires @Participants order) |
SmallVec<[T; N]> | Headers (N=2), tiers (N=3), features (N=4), token mappings (N=4) | Inline storage for common sizes; avoids heap for typical cases |
VecDeque | Worker idle queue (batchalign) | FIFO fair scheduling |
Dense Vec indexed by position | Retokenize word-to-token mapping | O(1) lookup, no hashing overhead, cache-friendly |
No LinkedList, BinaryHeap, or custom allocators.
Tree-Sitter Memory Model
Tree-sitter parsing is zero-copy for CST traversal:
// Node<'a> borrows from Tree, no allocation per node
fn process_node<'a>(node: Node<'a>, source: &str) -> ParseResult<...> {
for i in 0..node.child_count() {
let child: Node<'a> = node.child(i).unwrap(); // Stack-only, no heap
let text: &str = child.utf8_text(source.as_bytes())?; // Borrows source
// ... convert to owned model types ...
}
}
The tree-sitter parser consumes &str, produces a CST, and the Rust traversal
code constructs owned model types from CST nodes.
SQLite Memory-Mapped I/O
The validation cache uses SQLite with memory-mapped I/O for fast random access:
SqliteConnectOptions::new()
.journal_mode(SqliteJournalMode::Wal) // Concurrent reads during writes
.pragma("cache_size", "-8000") // 8 MB page cache
.pragma("mmap_size", "268435456") // 256 MB memory-mapped region
.synchronous(SqliteSynchronous::Normal) // Balanced durability
This configuration handles 95,000+ cached entries efficiently. The cache is never
deleted (use --force to refresh specific paths).
Manual Drop Implementations
Three types have custom Drop for resource cleanup:
| Type | Cleanup action | Why |
|---|---|---|
AuditReporter | Joins audit writer thread and flushes output | Audit mode owns file IO in a dedicated writer thread |
CheckedOutWorker | Returns worker to idle queue + releases semaphore permit | RAII pool resource management |
WorkerHandle | Sends SIGTERM/SIGKILL to child process | Process must be terminated when handle drops |
All drops are acyclic, no ordering dependencies between them.
Allocation Optimization Patterns
Rather than using an arena allocator (bumpalo was evaluated and removed, the data lifetimes don’t fit the “allocate many, free all at once” pattern), the codebase uses targeted optimizations:
| Pattern | Where | Savings |
|---|---|---|
| Scratch buffer reuse (clear + swap) | DP alignment row costs | ~50% fewer allocations in inner loop |
Flat table (vec![...; rows * cols]) | DP small-problem fallback | 1 allocation vs rows+1 |
| Dense Vec instead of HashMap | Retokenize word mapping | O(1) lookup, no hash overhead |
| SmallVec inline storage | Throughout | Avoids heap for 1-4 element collections |
SmolStr inline strings | All short CHAT tokens | No heap allocation for ≤23 byte strings |
See also: the batchalign3 book’s Arena Allocators page for the full evaluation of where arenas do and don’t help.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Algorithms and Data Structures
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
This chapter documents the key algorithms and data structure decisions across the TalkBank Rust crates.
CHAT AST Representation
The CHAT model is a tree of owned enums. The two central types are:
UtteranceContent: 24 variants covering all main-tier contentBracketedItem: 22 variants for content inside groups/brackets
flowchart TD
file["ChatFile"]
header["Headers\n(@Languages, @Participants, ...)"]
utt["Utterance"]
mc["MainContent\nVec<UtteranceContent>"]
dt["DependentTiers\n(%mor, %pho, %gra, ...)"]
file --> header
file --> utt
utt --> mc
utt --> dt
mc --> word["Word / AnnotatedWord / ReplacedWord"]
mc --> group["Group / PhoGroup / SinGroup / Quotation"]
mc --> marker["Pause / Separator / OverlapPoint / ..."]
group --> bi["BracketedContent\nVec<BracketedItem>"]
bi --> word2["Word / ReplacedWord / Separator"]
bi --> nested["Nested groups"]
Memory layout: Large variants (e.g., AnnotatedWord with scoped annotations)
are Boxed to keep the enum’s stack size bounded.
Content Walker
Location: talkbank-model/src/alignment/helpers/walk/
Closure-based recursive traversal centralizing the walk over all 24+22 variants:
pub fn for_each_leaf<'a>(
content: &'a [UtteranceContent],
domain: Option<AlignmentDomain>,
f: &mut impl FnMut(ContentLeaf<'a>),
)
Domain-aware gating:
Some(Mor): skips retrace groups (retrace words aren’t morphologically analyzed)Some(Pho | Sin): skips PhoGroup/SinGroup (treated as atomic by those tiers)None: recurses everything unconditionally
Both immutable (for_each_leaf) and mutable (for_each_leaf_mut) versions exist.
Used by talkbank-model, talkbank-transform word extraction, and other
typed CHAT traversals across the workspace.
Parsing Strategies
Tree-Sitter (Canonical Parser)
flowchart LR
src["Source .cha text"]
ts["tree-sitter C parser\n(generated from grammar.js)"]
cst["CST (Tree)"]
conv["Recursive descent\nover CST nodes"]
model["ChatFile (owned AST)"]
errors["ErrorSink\n(diagnostics)"]
src --> ts --> cst --> conv --> model
conv --> errors
- Grammar defined in
grammar/grammar.js(source of truth) parser.cis generated, never edit directly- CST-to-model conversion: recursive dispatch on node kind, skip
WHITESPACES, report unrecognized nodes viaErrorSink - Strict + catch-all pattern: Known header values get named grammar rules (syntax highlighting); unknown values hit a catch-all (flagged by validator)
Fragment Parsing
TreeSitterParser provides fragment methods for parsing individual CHAT
fragments (a word, a tier line) directly. Methods like
parser.parse_word_fragment(), parser.parse_main_tier_fragment(), etc.
are used when synthesizing CHAT from non-CHAT sources (ASR output, UD
annotations).
Structural Tier Alignment
Location: talkbank-model/src/alignment/traits.rs
Generic positional_align() pairs main-tier words with dependent-tier items by
position (O(n)). Traits: AlignableTier, TierAlignmentResult, AlignableContent.
%phoand%sinuse generic positional alignment%mor,%gra, domain-specific custom implementations- Mismatch diagnostics via
similarcrate (Patience diff algorithm, O(n log n))
%wor is not part of this structural alignment family. Timing consumers use
the checked Missing | Drifted | CountMatched transition documented in
%wor Timing Semantics.
Caching
The CHAT-core validation cache is documented separately in
Validation Cache. The
upstream batchalign3 project documents its own audio-task cache
(FA / UTR ASR / media conversion) separately.
Text Processing
Regex Compilation
All regex patterns use LazyLock<Regex> from std::sync, compiled once at
first use, lock-free thereafter. Never call Regex::new() inside functions or
loops.
Deterministic Output
BTreeMapfor all test/snapshot JSON (lexicographic key ordering)IndexMapfor participant/speaker ordering (preserves encounter order per spec)- Frequency results collected into
BTreeMap<NormalizedWord, Count>
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Setup
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
Getting a working checkout, and what you need installed for each surface you might touch. What to RUN once you are set up is in Developer Verification Checks, which owns that list.
Development is supported on Windows, macOS, and Linux. The commands below use Unix shell syntax; on Windows use PowerShell or Git Bash.
Prerequisites
Always:
- Rust via rustup. Do NOT install a version by hand:
rust-toolchain.tomlselects the current stable release and its components, and rustup honours it automatically. A new stable’s clippy lints are caught by the weeklyclippy-rolling.ymlrun and fixed in a focused commit. - just for the repo’s recipes. Not
strictly required, but every command in the contributing docs is a
justrecipe, and the recipes are the single owner of how each check is invoked.
Per surface, only if you touch it:
| You are changing | You also need |
|---|---|
the grammar (grammar/grammar.js) | Node.js at grammar/.nvmrc, then cd grammar && npm ci for the locked Tree-sitter CLI 0.27.0 |
| the grammar, so the typed traversal must be regenerated | a local checkout of tree-sitter-grammar-utils, which is not yet published (see Grammar Workflow) |
the re2c lexer (crates/talkbank-parser-re2c/src/lexer.re) | re2c at the exact version in re2c-version.toml, which provides the re2rust binary; just verify-vendored-lexer rejects drift |
| the book | just book-install-tools (installs mdBook and lychee into .tooling/) |
Nothing here needs a TalkBank corpus or any network service. The CHAT core builds and its tests pass on a fresh machine with only the “always” row.
Clone and build
git clone https://github.com/TalkBank/chatter.git
cd chatter
cargo build --workspace --locked
Then run the tests to confirm the checkout is sound:
just test # cargo test --workspace --tests, about a minute
Two Cargo workspaces
The repository has two INDEPENDENT Cargo workspaces. This trips people up
because --workspace from the root does not reach the second one, so a spec
change can be broken while every root gate is green.
1. The root workspace (Cargo.toml)
Every crate for parsing, model, validation, transform, CLI, LSP and desktop.
Plain cargo commands from the repo root operate here.
2. The spec workspace (spec/Cargo.toml)
Two member crates, spec/tools and spec/runtime-tools. Reach it with the
WORKSPACE manifest, not an individual crate’s:
cargo test --manifest-path spec/Cargo.toml --workspace # or: just test-spec
just test-spec is the same thing, and just gate runs it. What
the two crates are for, and why the split exists, is in
Spec Tooling.
The recipes
just --list
That is the authoritative catalog and it is worth reading once end to end: it
covers testing, both generators, the spec gates, formatting, the book, doc
dates, the vendored lexer, coverage, and the release commands. This page
deliberately does not reproduce it: a copy would drift and tell contributors
that recipes such as just test-spec, just spec-status,
just form-markers-gen, just symbols-gen, just verify-vendored-lexer and
just doc-dates do not exist.
Which recipes to run, when, and what each costs: Developer Verification Checks.
Pushing
just push # runs `just gate`, then pushes
just gate is the pre-push gate: everything CI runs that can run on one
machine, in one command. It takes 12-15 minutes. CI is a confirmation, never
the thing that finds your bug for you.
A green just test is not a green gate: the gate also runs doctests, the
spec/ workspace and the lints. If you find yourself assembling the gate by
hand from a list, that list is the bug.
There is no make verify and no Makefile; the just recipes are the only
entry points.
Editor setup
rust-analyzer works out of the box on the root workspace. If you are editing
under spec/, point your editor at spec/Cargo.toml as a second linked
project, or it will report the spec crates as not belonging to any workspace.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Grammar Workflow
Status: Current Last modified: 2026-10-07 (commit 5e895791)
The tree-sitter grammar at grammar/grammar.js is the formal definition of the CHAT format. Changes require careful validation.
The following diagram shows the complete regeneration pipeline. Every step must pass before committing a grammar change.
flowchart TD
edit(["Edit grammar/grammar.js"])
generate["tree-sitter generate\n→ src/parser.c\n→ src/node-types.json"]
traversal["regenerate the typed traversal\n→ generated_traversal.rs"]
grammar_test["tree-sitter test\n(corpus tests)"]
rust_test["cargo test -p talkbank-parser\n(CST-to-model conversion)"]
equiv["parser equivalence\n(corpus/reference/ files)"]
spec_check{"Grammar change\naffects spec examples?"}
test_gen["spec/tools generators\n→ grammar/test/corpus/generated/\n→ parser-tests generated tests\n→ validation fixture corpus"]
snapshot["observation snapshot\n(codes + roundtrip per spec example)\nadjudicate every diff"]
commit(["Commit"])
edit --> generate --> traversal --> grammar_test --> rust_test --> equiv --> spec_check
spec_check -->|Yes| test_gen --> snapshot
spec_check -->|No| snapshot
snapshot --> commit
Step-by-Step Procedure
1. Edit the Grammar
Modify grammar.js in the grammar/ directory. Key design principles:
- Explicit whitespace (no
extras) - Precedence annotations to resolve ambiguities
- Named rules for all semantically meaningful nodes
2. Generate the Parser
cd grammar
tree-sitter generate
This produces src/parser.c and src/node-types.json. Never edit these files by hand.
tree-sitter test does NOT detect a stale parser.c, so nothing downstream
can be trusted until this has run.
3. Regenerate the Typed Traversal
crates/talkbank-parser/src/generated_traversal.rs is the single generated
visitor the whole production parser dispatches through, produced from the
grammar’s JSON by tree-sitter-grammar-utils. A grammar change that alters
node types or their positions makes it stale.
From the Chatter checkout, use its owning recipe:
TSGU_DIR=/path/to/tree-sitter-grammar-utils just traversal-gen
just conformance-gen
TSGU_DIR selects the generator checkout; the recipe defaults to the tsgu
directory in the user’s home directory. It refuses a dirty generator checkout,
builds its generate_typed_traversal example, and reads edition and toolchain
from Chatter’s own manifests rather than a copied command. The output header
records the generator’s version and source identity.
The same recipe emits grammar/bindings/rust/tsgu_language_metadata.c through
the generator’s --metadata-bridge mode. This companion is generated, not
hand-maintained: the grammar build compiles it against its own parser.h and
namespaces its export. It supplies actual ABI-15 symbol/alias metadata for
source-bound nonmissing admission; it does not parse CHAT or C source text.
The generator runs rustfmt on its output, so no separate cargo fmt step is
needed. Never hand-edit the file; if the output is wrong, fix the generator as
a general change and regenerate.
Generated lint findings belong in the generator, not hand-edited output. Run
its generated-code Clippy probe and reconstruction regressions before consuming
a changed generator. Fallible extraction uses Result’s must-use contract;
source projections have explicit must-use annotations. Admission remains sealed
against caller-supplied projections, with a narrowly documented private-bound
allowance rather than a public escape hatch.
The audited compiled-metadata FFI boundary and grammar-specific admission
scaffolding also carry narrowly justified warning allowances in the generator’s
runtime source. Keep those allowances at their owner; do not suppress warnings
across handwritten parser consumers or remove source/language admission checks.
Borrowed uninhabited recovery arms do not dereference their payloads: they fail
explicitly if reached, without fabricating a recoverable CHAT state. Source-bound
accessors may be constant functions, but that does not establish new admission.
For a generator API change, first implement and verify the change in TSGU,
then commit that reviewed generator change locally so regeneration has clean
source identity. Regenerate Chatter, migrate the handwritten consumers of the
changed API, and exercise the affected spec/reference workflows. A local
generator commit is not publication. Grammar changes additionally require
tree-sitter generate before traversal regeneration; a generator-only change
does not require changing the grammar.
The generated module owns both CST shape and source-binding capabilities.
ParsedSource owns the tree/input association; its bind_typed transition
produces SourceBound<T> for supported generated wrappers. Consumers obtain
the node and its source from that binding, not separately supplied values.
Raw-node admission checks tree membership and canonical coordinates, because
independently editing a node copy can leave membership intact while shifting
its ranges. See Parsing for the current migration
boundary and recovery policies. Binding proves source identity, not syntactic
completeness or semantic validity; do not remove recovery states merely because
the finite corpus has not reached them.
The generated source-bound extraction path:
bound.extract() returns Result<SourceChildren, ReconstructionFault>; the
admitted immutable carrier’s generated
field_<minted_name>() accessors retain source identity in SourceField.
Structural projections (slot, iter, optional, and view) preserve that
association through positional slots, sequence groups, choices, extras, and
recovery payloads. SourceSlotView retains every recovery state and preserves
uninhabited payloads; a leaf’s read() admits its range without another root
membership search. Choice views preserve the already reconstructed variant.
SourceSlice::extract_header() retains association through supertype dispatch,
including Missing, Error, and Unexpected outcomes.
Selection records leaf, choice, optional, sequence, repeat and recovery decisions in one source-bound match plan. Extraction consumes those decisions; it must not independently rematch a mutable child cursor. Only successful construction commits the selected span and recovery sink. Flat repeat links avoid recursively owned suffixes. Runtime plan/range checks remain explicit: a producer fault is not an input recovery state and must not be replaced with fabricated children.
Concrete free extractors also return Result; ERROR-root recovery returns
Result<Option<_>, ReconstructionFault>, where Ok(None) means the node is not
an eligible recovery root. Supertype self-classification remains infallible.
Chatter reports reconstruction faults as E001 internal failure, preserving partial
evidence but refusing a validity verdict. The conformance inventory explicitly
fails on a producer fault rather than silently omitting that node.
The transitional typed-node-plus-source text reader also reports E001 when a node range is outside the supplied source or bisects UTF-8. Those failures describe a broken API pairing, not malformed CHAT. Its successful range check still proves only readability, not source identity: use the source-bound API for new callers. Do not treat equal-width foreign text as an admitted pairing. Likewise, a present CA token with empty text or a symbol outside its shared generated registry is a producer/source-association failure. Actual MISSING tokens retain their existing recovery handling.
Document-root classification retains associated carriers for both complete
documents and reconstructed ERROR roots. The latter uses
SourceSlice::extract_full_document_from_error_recovery(). Line and header
dispatch preserve association from those carriers; wrapped-header fragments
enter the same dispatcher with their already admitted source slice.
Participant lowering accepts a SourceBound<ParticipantsHeaderNode> and carries
association through its contents, repeated groups, speaker, name, and role.
Its text reads accept no separately supplied source string.
Calling the free extract_participant(bound.node()) returns a fallible ordinary
carrier. Other header families retain their transitional lowering APIs behind
the associated dispatcher; their checked decoding remains necessary until they
also migrate. Do not infer that every parser region is source-bound.
Do not replace it with raw kind-based child walks or repeated bind_typed
membership searches in an inner loop. Canonical descent avoids those searches,
but still reports invalid UTF-8 ranges from the runtime: source identity alone
does not prove every descendant’s range is readable. Keep that refusal at the
admission boundary and retain all genuine recovery states. A bound root whose
children are immediately separated from their source is not a completed
consumer migration.
The staleness guard proves less than it looks.
generated_traversal_is_current recomputes the digests of grammar.json and
node-types.json, so it catches a forgotten regeneration after a GRAMMAR
change. Its inputs are those two files, so it cannot see the generator at all:
a module emitted by an older backend passes indefinitely, and the guard is not
wrong to pass it. It is answering a different question from the one its name
invites you to ask.
Which generator wrote the file is answered by the file, in its own header comment. A bare semver does not identify a build, which is why the generator stamps its source commit beside the version. When the question is which backend produced the module, read that header rather than trusting a green suite.
After traversal regeneration, run just conformance-gen. It derives the test
inventory from the current traversal source. Its generator-only build omits
the default conformance feature so an outdated generated consumer cannot
prevent its own regeneration. Normal test builds keep that feature enabled;
the integration tests still require conformance support.
4. Run Grammar Tests
tree-sitter test
Every test under grammar/test/corpus/ must pass. Tests live there
and are partially auto-generated from specs (primarily via
just spec-gen).
5. Run Parser Tests
cargo test -p talkbank-parser
This verifies the Rust parser wrapper handles all CST nodes correctly.
6. Run Parser Equivalence
cargo test -p talkbank-parser-re2c --test integration equivalence_reference_corpus
Every file in the reference corpus must parse correctly. Each .cha file is its own test, so failures are reported per file.
7. Regenerate Spec Tests
If the grammar change affects any spec examples:
just spec-gen
just spec-gen # every artifact derived from spec/
just spec-check # or: is the committed copy current?
This regenerates tree-sitter corpus tests and other generated outputs that still depend on the spec pipeline.
Do this when the grammar change actually affects generated artifacts.
8. Adjudicate the observation snapshot
just regen rewrites spec/observations/example-diagnostics.json, which
records for every spec example the codes each stage emitted and whether the
parsed model serializes back byte-exact. Its currency test keeps it honest,
but the test is satisfied by any regenerated file, so the gate here is human:
read the diff. Every changed entry is either INTENDED (the behaviour change
was the point; commit the regenerated snapshot in the same change) or
UNINTENDED (a regression; fix the code, never the snapshot). A construct the
suite does not exercise is a missing spec example, and adding one is part of
the change, not a follow-up.
The reference corpus is a regression signal, NOT a validity authority
corpus/reference/ must stay green, but it is not “the ultimate arbiter of
correctness”, and a single failure is not a reason to revert immediately:
acting on that would entrench bad data.
The corpus is SYNTHESIZED. When a change makes it reject a file, adjudicate the
FILE against the real authorities (spec/, the grammar, and real corpus data)
and fix the data, or move it to spec/errors/ if the construct is genuinely
invalid. Weakening the parser to keep a reference file green is the one
response that is always wrong. The roundtrip gate stays green either way.
Common Patterns
Adding a New Token
- Define the token in
grammar.js - Add handling in the Rust tier parser (match on the new node kind)
- Add a spec construct example
- Run the relevant generation and verification steps
For small, isolated syntax additions, the grammar workflow should stay local:
- one grammar change
- one grammar corpus example
- one full-file fixture if needed
Changing a Rule
- Modify the rule in
grammar.js tree-sitter generate && tree-sitter test- Update Rust parser if CST node structure changed
- Update spec examples if the expected CST changed
- Run the current local verification sweep from
contributing/dev-checks.md
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Spec Workflow
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
How to change spec/ and leave the repository consistent. For what the fields
MEAN, read Spec System first; this page is
the procedure.
Every command here is written out. If a step here disagrees with what the tools do, the tools are right and this page is a bug.
Before and after any spec change
just spec-status # what state the spec system is in, derived from the gates
Run it before you start, so you know what “unchanged” looks like, and again at the end. A change that moves the “deferred” or “failing” counts in the wrong direction is worth a second look.
Before treating deferred examples as implementation work, inspect their live
claim review with cargo run --manifest-path spec/Cargo.toml --bin spec_status -- --deferred.
Legal controls and subsumption claims remain regression obligations, not a request
to recreate a diagnostic that no longer applies. Resolve contradicted claims from policy and
source evidence; do not flip registry status merely to reduce the deferred count.
Adding a construct spec
A construct spec is a VALID fragment plus the tree it must parse to.
1. Write the file under the right spec/constructs/ subdirectory
(header/, main_tier/, tiers/, utterance/, word/):
# my_example
Description of what this example demonstrates.
## Input
```utterance
*CHI: hello world .
```
## Expected CST
```cst
(utterance
(main_tier
...))
```
## Metadata
- **Level**: utterance
- **Category**: main_tier
The fence label (utterance here) names a template in spec/tools/templates/
that wraps the fragment into a full CHAT file. If no template matches, create
one; the generator fails rather than guessing.
2. Get the real CST rather than writing one by hand:
cd grammar && tree-sitter parse <a file containing your input>
Copy the tree, dropping byte positions and field names.
3. Regenerate and verify (see “Regenerating” below).
Adding an error spec
An error spec is INVALID CHAT plus the codes it must produce.
1. Write the file in spec/errors/, named E###_<slug>.md. Everything
declared goes in +++ TOML frontmatter; the prose goes in the body.
+++
code = 'E301'
name = 'Empty speaker code'
[[example]]
source = 'E3xx_main_tier_errors/E301_empty_speaker.cha'
level = 'utterance'
claim = 'violates'
chat = '''
@UTF8
@Begin
@Languages: eng
@Participants: CHI Target_Child
@ID: eng|corpus|CHI|||||Target_Child|||
*: hello .
@End
'''
+++
## Description
Empty speaker code.
A misspelled or unrecognised key is a LOAD ERROR, so you find out from
just spec-check rather than from a field that silently did nothing.
Four things decide whether your spec asserts anything, and each is easy to get wrong. They are covered in full in Spec System; in short:
claimis the field that asserts, and it is REQUIRED.violates(the spec’s code must appear),legal(it must not), orsubsumed_by <code(s)>(the targets appear and the spec’s code does not). Extra emitted codes still pass; the exact per-stage sets are the snapshot’s business.- There is no
layerfield. Which stage catches a rule is observed, not declared: every example is a fixture whose runner checks both stages, and the per-stage record lives in the observation snapshot. statusandkindare NOT yours to declare. They are facts about the CODE, and they live inspec/codes/error-codes.toml, one entry per code. A spec naming a code that file does not declare does not load, andstatus = 'not_implemented'THERE still defers every example of that code and#[ignore]s its generated tests. Writing either key in a spec file is a load error naming the key. The separate deferred-spec regression check still verifies explicitlegalandsubsumed_byclaims against the live parser and validator. Deferring implementation of a code does not excuse a stale claim about accepted input or the alternative diagnostic that rejects it. Plannedviolatesexamples remain deferred; satisfying a subsumption claim does not implement the deferred code or change its registry status. There is no default forstatus: an invented answer to “is this rule live” is the kind of wrong value nothing notices, and a per-file copy of a per-code fact is the kind that several files could disagree about.source’s stem names the transcript, which is what rules about the file’s own name (E531) compare against.
Write the failing case first. A new error spec should fail before the rule exists; that is what proves the fixture actually triggers it.
A brand-new code needs one bootstrap step. The ErrorCode enum is
generated by spec_gen, and spec_gen (like schema-gen, the first step of
just regen) links against talkbank-model. So a validation rule that names
the new variant cannot compile until the enum exists, and neither command can
produce the enum while the rule is in place. Add the registry entry and the
spec first, run just spec-gen with the rule not yet referencing the variant,
then write the rule and run just regen in full; the second run rebuilds the
observation snapshot against the rule’s real behaviour.
Regenerating
Mutation candidates are not golden tests
For source-bound main-tier terminator deletions, use:
just spec-mutation-candidates corpus/reference/content/terminators-standard.cha
This runtime command admits one seed only after diagnostic-free parsing and
alignment-aware validation, including rules about its actual transcript stem.
Warnings also refuse admission. Add --strict-linkers to select the optional
cross-utterance rules; admission does not claim validity under unselected rules.
It produces one candidate per present typed main-tier terminator, verifies that
the model’s span matches the complete token, and preserves all other bytes.
Missing optional terminators produce no candidate. It never rewrites the model
or guesses a token from line endings.
Output is JSON on stdout, containing the exact seed, transcript identity, selected linker policy, and each candidate’s deleted byte range, removed text, and resulting CHAT. Every candidate is explicitly unreviewed, with no expected diagnostic. No output files are created by the command. Seed or span rejection occurs before candidate output; stdout I/O can still fail partway through a write. Retain the seed record when saving candidates so provenance does not depend on an input path continuing to hold the same bytes.
The API lives in spec-runtime-tools::mutation, where live parser/model
dependencies belong. AdmittedSeed privately owns its model and borrows its
immutable source; candidates can only be constructed from that model’s checked
spans. This is AST-source association, not an assertion that the generated CST
API has already migrated every mutation family. Only terminator deletion is
implemented here; the legacy mutators below are not migrated.
just spec-perturb produces unreviewed candidates, not diagnostic oracles.
Its JSON records have assessment: "unreviewed" and no expected_error field.
The removed diagnostic table had drifted from canonical rules; a mutation’s
name cannot establish which diagnostics a particular input should produce.
The legacy text mutators do not admit a valid seed through the parser/model. They can affect continuation lines, normalize final newlines, select a speaker code already declared by a seed, or introduce multiple defects. Do not use their outputs directly as goldens or treat a syntax-error census as validity. Before running them, review output-name collisions and the intended input scope.
Candidate output has an explicit planning phase: duplicate destinations
and existing output paths are refused before candidate writes begin. Execution
uses exclusive file creation as well, so a file created after preflight cannot
be overwritten. This protects output, not seed validity. A mid-run I/O failure
may leave earlier newly created candidates; preserve and review them rather
than assuming the batch was atomic or automatically deleting them. The legacy
--mine summary path is separate and is not covered by this candidate contract.
For promotion, retain the valid control, describe the precise mutation and its
source relationship, justify a claim from the specification or a recorded
ruling, then add the reviewed example to spec/errors/ and regenerate through
the normal owner. Review observations separately from expectations. A changed
parser output is evidence to adjudicate, not permission to rewrite the claim.
Generate owned artifacts
One command, from anywhere in the checkout:
just spec-gen # rewrite every generated artifact from the specs
just spec-check # or ask whether the committed copies are current
It regenerates every artifact in the registry, in dependency order (the
observation snapshot first, since the tree-sitter corpus derives its
membership from it); the generated artifact
table included in the
spec-system chapter is the live list. There is nothing to choose and no path to type:
every destination is a constant in spec/tools/src/artifacts.rs, so a
generator cannot be aimed at the wrong tree.
just spec-check writes nothing and is exactly what the
every_generated_artifact_is_current gate runs, so a green check means a green
gate.
The published error-reference pages under docs/errors/ are part of
spec-gen like every other artifact, and spec-check gates them.
Never hand-edit anything under a generated/ directory. An artifact that owns
its directory wipes it wholesale and refuses to clear one lacking its
.generated-output-dir marker.
Verifying
just spec-status # the derived summary
cargo test --manifest-path spec/Cargo.toml --workspace # every spec-side gate
just test # the main workspace
If your change touched the grammar, follow the full
Grammar Workflow as well: a
grammar.js edit needs tree-sitter generate before any parser behaviour can
be trusted.
Updating a registry
Two closed vocabularies live under spec/, each generating every site that
names it. Neither is edited anywhere but its registry.
just symbols-gen # spec/symbols/symbol_registry.json
just form-markers-gen # spec/form_markers/form_marker_registry.json
flowchart TD
registry["Edit the registry JSON"]
gen["Run its generator\n(loading validates; there is no separate check step)"]
fmt["Generator runs rustfmt on Rust output"]
gate["Drift gate compares committed output\nagainst what the generator produces"]
registry --> gen --> fmt --> gate
The generators format their own Rust output deliberately: otherwise just fmt
and the generator each rewrite the same bytes and the drift gate fails forever,
with both sides correct.
Each registry’s README covers its authorities and the follow-ups its generator
cannot do:
spec/symbols/README.md,
spec/form_markers/README.md.
Common mistakes
- Editing generated files. Change the spec or the registry, then regenerate.
- Wishing for an example that asserts nothing. There is no such state:
claimis required, and an example that cannot honestly sayviolatessayssubsumed_by(the worklist) orlegal(the boundary). - Flipping
statustoimplementedwithout regenerating. The fixture manifest still carries the old status, so the runner keeps skipping what you just enabled. - Regenerating reflexively. Regeneration is for artifacts that genuinely changed, not a substitute for deciding what the change needs.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Correctness Architecture
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
This is the target design of chatter’s correctness machinery, written for the maintainer who inherits it. It is not a patch list and not a description of the tree as it stands today. Where the current tree disagrees with this page, the tree is what has to move.
The measurements below are design evidence, not a completion report. For the current measurement boundary, use Coverage scope and completeness. Uncovered code is an investigation target, never automatic permission to delete recovery, constructor, wire-format or internal-failure handling.
Read Testing for what the layers are called today and Spec System for what the spec fields mean. This page says what the layers are FOR, which of them may be deleted, and how a successor knows the suite is complete rather than merely green.
Every count on this page carries the command that produced it. The commands are collected in Appendix A so that a number which has drifted can be re-derived rather than believed.
Properties of the gate mechanism and the measured suite
These are properties of the current machinery, stated as design a successor must preserve. Counts are re-derived with the commands in Appendix A.
The whole suite can be run. Chatter’s whole suite runs in well under a minute, so the workspace guard permits running tests freely. A guard that prevents measurement prevents the work it protects.
Dead snapshots are found with two independent witnesses. A snapshot is dead
only if it is unreferenced by a run AND no test function of that name exists
anywhere, or it is an exact duplicate at another path, or its crate prefix names
a target that no longer exists. A witness a different function can satisfy is
not a witness (fn pho_tier matches the accessor pho_tier(). The
snapshot-hygiene gate is part of gate.
Undemonstrated codes are read from the registry. spec/codes/error-codes.toml
carries a status for every code, and three of its values legitimately have no
example: not_implemented, deprecated and unreachable_from_chat. A code is
undemonstrated only when it has none of those and no example. Production code
names a rule by its ErrorCode VARIANT, not by its number, and the number is
what a maintainer writes in comments; any scan for “is this code applied” must
use the registry’s variant field over non-comment code.
The fabricated-AST population is two populations. Constructions that pass
the same string twice are Word::simple, correct by construction. Those that
pass two different string literals state a raw text carrying CHAT markers and a
cleaned text without them independently, with nothing forcing the second to be
what cleaning the first produces; and the constructor stores the cleaned string
as a single flat text element, where a parse of the same word would produce
structured content naming the marker. Such tests may assert on shapes the parser
cannot produce, which is this page’s central claim. They cannot be fixed in
place, because the crate cannot parse; they move to one that can. The
remainder is concentrated in talkbank-model, the crate without a parser
dependency, as a consequence rather than a coincidence. The ratchet that holds
the count excludes comments and string literals, because prose about the hazard
must not score as the hazard, and a measure that errs toward undercounting
would mask a real addition.
The grammar corpus is self-certifying where it is generated. No construct
spec declares an expected tree, and the branch that could set one is
unreachable, since it fires only for a chat-file or document input fence
that no spec uses. The field the corpus test asserts is full_cst, the
whole-document tree, so every GENERATED corpus case expects exactly what the
parser produced when the case was written: a wrong grammar rule regenerates a
wrong expectation and passes. The cases under grammar/test/corpus/manual/ are
hand-authored and no generator touches them, so “self-certifying” is true of
the generated tree, not of the corpus.
The human-authored expectation is disconnected, not missing. Almost every
construct spec carries a fragment cst block and nothing asserts one: the
accessor is reachable only through a dead branch, the one function that would
compare it has no caller, and the sole gate on the block is a paren-balance
check whose own doc comment says the block is read only by humans. Many of those
blocks name node types the grammar does not have, and many contain a literal
... ellipsis. So “make the corpus assert the fragment the human wrote” is not
available as stated; the options are to repair the blocks first, or to make
full_cst required and derive it once under review, which buys review rather
than authorship. Either way the change is in
spec/tools/src/output/tree_sitter.rs around the substitution, and it
regenerates every generated case.
How the gate-probe mechanism is built
Every gate must be able to fail, and the mechanism that proves it is itself built so that its two advertised properties are not reachable around.
- A clean verdict cannot be composed by its author. A tuple variant of a
public enum is a public constructor of whatever it holds, so
Outcomecarries noClean(String, Examined)that would pair any summary with any witness. A clean verdict is produced only throughReadTree::clean, which requires theExaminedwitness. An absent directory mints a different witness from an empty one, so a gate whose whole scope was removed cannot report clean over nothing. - A probe suite always contains a real planted violation.
ProbeSuite’s constructor takes the first refusal as arguments and adds the control probe itself, so a suite with no probes and a suite of only the control (which satisfies “a gate that rejects everything” because a control IS aMustPass) are unconstructible, and there are no runtime checks for them. - A gate’s tier is the caller’s. Each gate
Preconditionis compared at a tier:Precondition::atreturnsUnmet::{Skip, Fail}rather than anifon a bool inside a test. The tier is a private recipe ARGUMENT (_test tier), sojust testruns atinner-loop(a missing input is a skip) andjust gateruns strict (a missing input is a failure); a gate that can skip itself at the pre-push gate is the state the mechanism exists to forbid. - A gate cannot open a second tree.
Tree::liveis public andpub(crate)is no barrier inside the crate, and no type expresses “do not open a second tree”.gate_disciplineis a registered gate that reads thefn checkbody of every file declaringimpl Gate forand refuses the calls that reach a checkout directly, scoped to those bodies so a helper in the same file may still build its own tree. - A walk failure is plantable. A directory enumeration that fails is a
TreeEditset, so a walking gate proves its walk-failure half like any other rule, instead of declaring it unprobeable. - Ratchets are gates, not scripts. A check written as Python under
scripts/is at the wrong altitude;scripts/lint/holds no ratchet. The fabricated-AST gate counts neither comments nor string literals, and the demonstration gate decides whether a file IS a code’s spec file rather than merely documents the code (E202_missing_form_typedocuments E202; it is not E202’s file). - Fragment entry points rebase diagnostics, not sinks.
rebasemoves the model; a diagnostic already handed to a sink is past moving, so the re2c fragment entry points do not pass the caller’s raw sink into fallible%gralowering.
What the mechanism cannot reach:
- A probe’s plant and its expected message are both its author’s invention, so a green suite certifies the author’s understanding of the rule, not the rule.
- The unproven-rule record is two hand-written lists reconciled by a test, which
catches drift between two files and is NOT the ratchet-against-reality that
UNPROTECTEDis. Deriving it needs a per-gate vocabulary of rule identifiers that a probe and an unproven entry each name, so the unproven set becomesrules() - probed(). - The mechanism is general repository infrastructure living in a CHAT
test-support crate, which is why
talkbank-parser-re2c’s parity gate cannot register inALLand keeps a degraded two-case copy of the shape. - A speedup of one piece is not a speedup of the loop. Parallelising the probe
run made the standalone binary faster and
just testslower (cargo testalready runs test binaries in parallel), and a shared read cache showed the reads were never the cost; time the loop, not the piece. The remaining lever is memoizing the derived per-file artifact each hygiene probe rebuilds (the blanked source), which needs a key a planted file invalidates.
The thesis
- The spec system is the single normative source of what CHAT is and what
chatter must do. A claim about correctness that is not written in
spec/is not a claim chatter makes. - Everything else is derived from it, or is a property no example can express, or is deleted. There is no fourth category. A test that is none of those three is volume, and volume is not evidence.
- Coverage is the completeness proof. An unreached branch is either a missing spec example or dead code. Those are the only two verdicts, and each has a named consequence.
- Nothing depends on private data. A successor who can read only this repository can run every gate, add every kind of evidence, and adjudicate every failure. The production corpus is read for evidence about which paths matter; its findings land as committed fixtures, and then it is not needed again.
The answers, in one screen
What proves what. Six evidence classes, and nothing outside them survives: spec examples (normative), comparisons (experimental backend and adjudicated CHECK fixtures), properties (universally quantified), boundary tests (subprocess, stdio, cross-process), algorithm units (behaviour no type can hold), and derived artifacts (which prove nothing and are the mechanism). Section What proves what is the table.
Where to add evidence when you fix a bug. A wrong verdict on CHAT input is
a [[example]] in spec/errors/. A wrong parse shape is a construct spec with
a model claim. An invariant over all inputs is a property test with a real case
budget. A bug that only appears through the CLI, the LSP or the desktop app is
a boundary test in that crate. A disagreement between the two parsers shrinks
KNOWN_DIVERGENCES. See
Where evidence goes.
What you may delete. 335 orphaned snapshot files, a duplicated roundtrip harness, a second error corpus at the repository root, a hand-mirrored CLI command list, 137 rotted CST blocks, and roughly 210 unit tests that a type change makes unwritable. Sizes and receipts in What is deleted.
How you know the suite is complete rather than green. Three questions, three instruments: did anything RUN the code (coverage attribution), did anything OBSERVE what it did (mutation, scoped), and is what it does RIGHT (the oracles, and human adjudication against the format authority). The first two are gated. The third cannot be, and What the criterion cannot cover says so plainly.
The one architectural fact that explains the current shape
talkbank-model cannot parse a CHAT file.
Its Cargo.toml declares no parser dependency; its dev-dependencies are
insta, proptest and tokio. It holds 671 of the repository’s 2,868 test
attributes, 23% of the suite, and not one of those tests can turn CHAT text
into a model. Every one of them fabricates the AST it then judges. Repository
wide there are 288 new_unchecked call sites and 403 Span::DUMMY
constructions, the large majority of both inside that crate.
Span::DUMMY is Span { start: 0, end: 0 }, which is also the legal
zero-width position at the first byte of a file. The type already carries a
long comment about this, opening with KNOWN HAZARD, for maintainers and
conceding that “the VALUE is still overloaded, and that part is not fixed”.
That comment is a receipt: prose is gated by nothing, so a paragraph explaining
why something is safe is a work item with an address, not a mitigation.
Three consequences follow from that single fact, and together they explain the shape of everything else:
- The spec corpus cannot reach the validation rules. 436 spec examples are
lowered into real
.chafiles and run through both stages, and they are the best evidence in the tree. But a rule that only fires on a shape the parser never produces cannot be demonstrated by any file, so those rules were demonstrated by hand-built models instead, in the crate that cannot parse.spec/errors/E232.mdsays this in its own words: the parser never constructs such a word from real input. - The test states the input twice. A hand-built test writes the raw text in a comment or a string, and the parsed structure separately, and nothing forces the two to agree. A test that fabricates its own input cannot be wrong about the format, only about itself.
- Volume accumulated as compensation. 679 committed snapshot files totalling
4.36 MB, of which 335 name a reference-corpus stem that does not exist; a roundtrip gate
implemented twice over the same 107 files; and a second error corpus at
tests/error_corpus/, 26 tracked files, whose generator walked one parent too many and had been writing a full corpus BESIDE the repository, touching nothing tracked. Confirmed by the 66 files it had left there. Fixed 2026-09-07; the duplication with the spec-derived corpus remains.
The redesign therefore begins with a type change, not a test cull. Make the AST constructible only from a parse product, and the 671-test family stops compiling. It does not get better tests; it stops being writable, and it has to move onto spec-derived input, which is the permanent deliverable anyway.
What proves what: six classes, and nothing else
| Class | What only this class can prove | Where | Size today | Who may add |
|---|---|---|---|---|
| SPEC (normative) | That a stated input violates, or does not violate, a stated rule. The only artifact a successor with no corpus can read and act on. | spec/errors/ (223 specs, 436 examples), spec/constructs/ (138 specs) | 436 lowered fixtures, one data-driven runner | Anyone. This is the default destination for new evidence. |
| COMPARISON | A disagreement requiring independent policy adjudication; neither implementation is an oracle. | Authoritative CHECK assessment; experimental re2c comparisons | Derived by the owning reports | Add independently justified regression evidence, not diagnostic mimicry. |
| PROPERTY | A statement quantified over all inputs: never panics, roundtrip, span arithmetic, cleaned-text invariants. The spec format has no quantifier and never will. | talkbank-parser-tests/tests/integration/property_tests/ | 27 files, 18 proptest! blocks | Anyone, at a real case budget (see below). |
| BOUNDARY | Behaviour of the actual seam: argv, exit codes, stdout contracts, LSP stdio lifecycle, Tauri async runtime, cross-process cache. No type of ours reaches the outside world. | crates/chatter/tests/integration (269 attrs, 136 spawns), talkbank-lsp stdio tests, apps/chatter-desktop/src-tauri/tests | 147 of 2,868 attrs (5.1%) reach a subprocess or runtime | Anyone. This class should GROW. |
| ALGORITHM UNIT | Behaviour no signature describes and no type can hold: number-to-words, POS mapping. | talkbank-transform/src/num_words/, clan_ud_mapping.rs | 65 attrs | Anyone, labelled in source with the category. |
| DERIVED ARTIFACT | Nothing. These are the mechanism, not the evidence: lowered fixtures, generated tests, the observation snapshot, the corpus sexp pins. | the seven registry artifacts | about 880 files written by one command | Nobody by hand. A generator owns every byte. |
There is a seventh class, and it is transitional by design:
| Class | Rule |
|---|---|
| PRODUCTION ATTESTATION | Retained production evidence supplies candidates for independently specified rules and minimal sanitized fixtures. Ordinary correctness tests consume the committed finite corpus without $TALKBANK_DATA. Production differential runs are not a release or CI requirement; completing the spec/reference coverage removes the need for that heavyweight process. An occurrence in production does not establish validity. |
The rule that makes this a design and not a taxonomy: a test that is not in one of these classes is deleted. Not deprecated, not annotated, deleted. If deleting it feels wrong, it belongs to a class and the class is the place to say so, in one line, in the source.
The normative core
What a spec example can say today
| Property | How |
|---|---|
| “this input violates rule X” | claim = 'violates' |
| “this input does NOT violate rule X” | claim = 'legal' (96 examples) |
| “this violates X but chatter reports Y today” | claim = { subsumed_by = [...] } (42 examples), positive and negative halves both checked |
| “the grammar produces exactly this tree” | a tree-sitter corpus case |
| “this parses with zero diagnostics” | a construct spec |
| “which stage catches this” | observed per example in spec/observations/example-diagnostics.json |
| “this input roundtrips byte-exact” | the snapshot’s roundtrip field, byte-gated, so a flip fails the currency gate |
The meaning of a claim has exactly one owner, Claim::satisfied_by, and both
runners call it. Keep that. It is the single best structural property of the
system, alongside the artifact registry.
What it cannot say, and what the redesign adds
It cannot state a property over all inputs, an idempotence, a preservation invariant other than the whole-file roundtrip, a performance bound, a claim about the parsed MODEL, a claim about the exact code SET, or anything about the CLI, LSP, transform layer or desktop app. Most of those stay outside forever; that is what the PROPERTY and BOUNDARY classes are for.
Three of them must move inside, because without them the spec cannot carry the completeness criterion:
1. The exact per-stage code set becomes claimable. Today, extra codes always pass: the claim is set membership, so “E316 fired instead of the rule you meant” is a recorded observation rather than a failing state. The snapshot already holds the exact per-stage sets. Promote them from observation to claim wherever a human has adjudicated them, and leave the rest observed. This is the smallest change that makes a spec example a statement about chatter’s behaviour rather than about one bit of it.
2. A construct spec claims a MODEL, not a CST. The authored
## Expected CST block is dead data and has rotted: 137 of the 138 construct
examples carry one, and 41 of them name at least one node type the grammar does
not have (25 distinct dead names, led by initial_word_segment in 18 files).
The generator ignores the block entirely and substitutes the parser’s own
to_sexp(), so the committed expectation is a recording of what the parser did.
Two repairs were available and they conflict. Making the authored CST normative is the wrong one: a CST is a statement about the grammar’s internal node names, which change legitimately under refactoring, and 41 rotted blocks are the receipt for exactly that. The model is the published contract; it has a JSON schema, and that schema is already gated. So:
- the
## Expected CSTblock, the unread## Metadatasection,update_cst(a library function that writes into a human-authored spec file, bypassing the.human-authoredmarker, one call away from being re-enabled) and the second format reference are DELETED; - a construct spec gains a declared model projection, which the generated test asserts, replacing the 138 bodies that today parse a string and discard the result;
- the sexp corpus stays and is relabelled honestly as a derived regression pin. It catches an accidental grammar change at the node level, which nothing else does. It is not a specification and must stop being described as one.
3. An example and an emit site name the same rule. Codes are emitted from
many branches: 670 ErrorCode:: references over 214 distinct variants, 133 of
them referenced more than once, and ErrorCode::TreeParsingError (E316) alone
appearing 83 times. A per-code example demonstrates that a code CAN fire, never
that a given site can. Until an emit site and an example can name the same
identifier, “every reachable branch” is checkable only by instrumentation, never
by the spec. Adding that identifier is what lets the two meet.
Two smaller repairs that the criterion depends on
- Derive
statusfrom the snapshot. It is authored today, and the format reference admits it. The system already knows, per example, whether a code fires. Leave only the genuine adjudications (deprecated,unreachable_from_chat) to a human. - Make the drop-outs visible. 117 examples qualify for the tree-sitter error
corpus and 71 files exist; the 46 that fall out are announced as generation-time
stderr and leave no committed trace. They become a committed, gated list with a
reason per entry, in the shape
node_coverage.rsalready uses, where an entry that becomes covered FAILS so the list cannot rot.
The completeness criterion
Coverage target and finite evidence
The target is 100% line, region and branch coverage of handwritten parser, model, validation and transform code. Tree-sitter has first priority; re2c remains experimental and lower priority. Generated code is reported separately. Uncovered executable behavior remains work: do not exclude it or fabricate an impossible model to raise the percentage. A path removed through a type invariant needs evidence that the invariant holds at its producer and that no consumer bypass remains.
Coverage does not establish semantic completeness. Pair independently authored spec expectations with legal boundaries, invalid witnesses, reference-corpus combinations, and adjudicated CHECK behavior. Preserve constructor and wire boundary tests where those APIs admit inputs that CHAT parsing cannot produce.
Reference-corpus traversal, roundtrip and backend comparison tests admit their
inputs through talkbank_parser_tests::chat_corpus::ChatCorpus. Its constructor
requires a nonempty recursive file inventory and reads every source before
returning. Missing directories, walk errors and unreadable sources fail the
measurement; the type does not certify CHAT validity. Directory enumeration
must never silently shrink the measured population.
Generated validation fixtures also preserve the example’s authored transcript
name in a required transcript_name manifest field. Their storage names (such
as W109_1.cha) must not become validation context. Anonymous examples remain
anonymous; named examples retain their exact Unicode spelling. Both the spec
runner and the generated-corpus runner therefore exercise filename-sensitive
rules with the same input identity. Older manifests require regeneration.
For a manifest-only generator change, regenerate its owning artifact without rebuilding unrelated generated outputs:
cargo run --manifest-path spec/Cargo.toml --bin spec_gen -- \
--artifact-root crates/talkbank-parser-tests/tests/error_corpus/validation_errors
--artifact-root selects an existing artifact owner and rejects unknown roots
before writing. It does not regenerate dependencies; use full generation when
the changed inputs affect other artifacts too.
Stated mechanically
For every branch in hand-written, non-generated code, the attribution artifact carries a row, and every row carries a verdict from a closed set. A row with no verdict fails the gate. A row whose verdict says it is unreachable and which is then covered fails the gate. The count of rows awaiting a spec example may only go down.
The gate is the VERDICT, not a percentage. That distinction is the whole design. A coverage percentage as a gate is a number to be gamed; a verdict is a statement a human made, with a reason, that a later measurement can contradict.
The instrument
# Include integration fixtures and report production dependencies, not just
# the fixture crate. Branch coverage requires nightly.
cargo +nightly llvm-cov --branch --workspace --exclude chatter-desktop --tests \
--json --output-path <out>.json -- --skip gates:: \
--skip conformance_inventory_current:: --skip generated_traversal_current::
This measures the functional suite, excluding gate modules and generated-source currency checks. The desktop runtime and the separate spec workspace are not measured by this command. Record these limits with the report; a failed run is not a complete baseline. The export also includes inline test code, so a source directory’s raw percentage is not automatically production-only coverage.
The region worklist’s JSON field is uncovered_region_starts, because the
inventory counts uncovered region starts, not branches. The separate
--branches-json <path> output records independent uncovered true/false
outcomes from LLVM’s branch records. Both inventories union source positions
across instantiations. They are not LLVM’s summary percentages: LLVM merges
region summary counts by taking the maximum per instantiation group, so
complementary instantiations can cover more distinct positions than its
summary reports. See LLVM’s
RegionCoverageInfo::merge.
The tool reconstructs that summary separately to check its arithmetic.
Unexplained disagreements remain explicit exclusions and cause a nonzero exit;
excluding every residual row must never produce a completeness claim. Reports
retain the caller’s --produced-by command beside their rows.
For conservative inline-test attribution, the parser-test example
coverage_source_ranges reads a JSON array of repository-relative Rust paths
on stdin and emits source-hashed test ranges. It parses explicit #[test]
functions and #[cfg(test)] inline modules/functions/impls with syn, not
regular expressions or mangled-name guesses. Pass its output to the worklist
with --test-ranges. Stale source hashes are rejected. Other cfg expressions
and macro-generated code remain production or unclassified; this does not
claim to reproduce rustc’s expansion or establish parse-backed evidence.
The inventories honor the export’s files selection: LLVM can retain function
records for test files absent from its file report, and those must not silently
re-enter the reported scope.
Region coverage is a good proxy for match-arm coverage, because each arm body is
its own region. It is a bad proxy for two shapes that are everywhere in a
validator: if cond { emit(...) } with no else, where the not-taken path is
not a region at all, and a && b, where short-circuit operands get no separate
region. “The rule never declined to fire” is precisely the state this redesign
wants to detect, so the criterion runs on --branch and the toolchain choice is
a measurement decision, not a CI decision.
Generated code is excluded, and the list of what is generated has ONE owner. It
does not have one today: mutants.toml excludes exactly one generated file with
a header explaining why (a mutant there indicts the generator), and that
reasoning applies verbatim to ten other committed generated files it does not
exclude. Derive both the coverage --ignore-filename-regex and
mutants.toml’s exclude_globs from the artifact registry, so a new generated
artifact cannot be excluded from one instrument and measured by the other.
The exclusion is not cosmetic. crates/talkbank-parser/src/generated_traversal.rs
alone is 45,284 lines, 24,549 instrumented lines and 3,924 branches, more than
double the entire hand-written parser, at 18.1% line and 19.4% branch coverage.
Including it would drag the parser’s reported figure down by about ten points
and every finding in the report would be a statement about the generator.
The row
One JSON file, one row per branch, plus a rendered page:
{
"file": "crates/talkbank-model/src/validation/utterance/spacing.rs",
"line": 214, "column": 21,
"region_kind": "branch",
"function": "...::check_separator_spacing",
"function_regions_uncovered": 1,
"function_region_total": 31,
"covered_by": [],
"witnesses": [],
"verdict": "NEEDS_SPEC_EXAMPLE",
"verdict_detail": "no example whose separator is trailing"
}
function comes from the export’s function-to-region nesting, so it is an exact
lookup rather than a line-range guess, and it resolves correctly through the
include!d generated test bodies. function_regions_uncovered over
function_region_total helps prioritize investigation; it cannot decide
deletion versus a missing fixture. A ratio of 1.0 says only that this measured
population did not execute the function. Review its public boundary, supported
configuration and producer invariants before assigning a disposition.
The verdict set is closed:
| Verdict | Meaning | Consequence |
|---|---|---|
REACHED_BY_SPEC | a spec example reaches it | none, this is the goal |
NEEDS_SPEC_EXAMPLE | reachable from CHAT, nothing reaches it | write the example; this count ratchets down |
NEEDS_PROPERTY | quantified, no single example expresses it | write the property |
UNREACHABLE_FROM_CHAT | no CHAT input reaches it | Classify supported API, recovery and tool-failure boundaries separately; retain needed handling and its boundary evidence. CHAT unreachability alone does not justify deletion. |
UNREACHABLE_BY_TYPE | the producer and all admitted consumers forbid the state | Remove the redundant arm only after checking construction, deserialization and mutation bypasses. |
REACHED_ONLY_BY_WILD_DATA | only the production corpus reaches it | synthesize a fixture, then re-verdict |
COVERED_ONLY_BY_FABRICATION | reached only by hand-built inputs | Not canonical CHAT evidence; convert semantic examples to parsed fixtures, but preserve legitimate constructor, wire and fault-boundary tests separately. |
OUT_OF_SCOPE_GENERATED | generated code | excluded by the registry, never hand-marked |
Coverage has a PROVENANCE, and fabricated coverage counts as uncovered
The maintainer, 2026-09-08: “Fake useless tests that fabricate are particularly dangerous and could inflate coverage numbers.” This is the sharpest constraint on the whole criterion and it was missing from it.
A test that fabricates its input still EXECUTES the code beneath it, so it
covers branches. chatter has hundreds: talkbank-model declares no parser
dependency, so every test in it hand-builds the AST it then judges. Coverage
bought that way is worse than no coverage, because it converts “nobody has shown
this rule firing on a real CHAT file” into a green number, and the number is the
thing a successor will trust.
So a covered region is not a fact until you know WHAT covered it:
- Parse-backed: reached by a spec example, a fixture or a reference file, through a real parse. This is coverage in the sense the criterion means.
- Fabrication-backed: reached only by a hand-built model. The rule ran; nothing showed it firing on CHAT. Counts as UNCOVERED for the criterion, and the test is a deletion or conversion candidate rather than an asset.
The split is measurable, and scripts/coverage_attribution.py measures it, but
only with THREE exports: the full suite, the fabricating tests alone, and the
suite with the fabricating tests excluded from the run but not from the report
(--exclude-from-test). Two exports give an upper bound only, because the full
run is a superset of the fabricating one and no subtraction between them
isolates anything. Saying so is the point: the bound is honest and the exact
figure has a price.
Pruning is on the table, and it is half the work
The maintainer, 2026-09-08: “Make sure that radical reorganization and pruning of tests is on the table.” The criterion above reads as a gap-filling machine, and read that way it can only make the suite bigger. It is equally a DELETION instrument, and the deletions are the cheaper half:
- A function with a ratio of 1.00 was not run by the measured population. Remove it only if its supported ownership and producer paths prove it redundant; otherwise classify the missing workflow or boundary evidence.
- A region reached only through hand-built input does not establish CHAT coverage. Use parsed evidence for CHAT behavior, while retaining necessary constructor, serialization and fault-injection contracts separately.
- A test whose parse-backed coverage is a subset of another’s, and which pins no policy of its own, is redundant. The suite is not better for having it.
- Type-enforced impossibility can remove a redundant implementation check. Keep tests of the constructor, policy and external boundary that establish the invariant; failure to reach a branch through CHAT is not that proof.
The standing rule that a test a type could obsolete should not exist is the same instruction from the other end. Test count is not the goal. The canonical coverage target remains 100%, without exclusions that conceal unfinished work.
The repository already trusts this exact shape twice: node_coverage.rs splits
its exclusions into INVALID_BY_CONSTRUCTION and NOT_YET_IN_CORPUS and argues
at length that an exclusion has a KIND which implies a check a flat list could
not express; construct_coverage.rs does the same for parent-child pairs, where
an entry that becomes covered fails. Follow both, in both directions.
Where it stands today, honestly
Measured over the WHOLE SUITE, every crate instrumented, generated files excluded, and split by what BOUGHT the coverage (values drift; re-derive them with the appendix commands before quoting any). The command set is in the appendix; the third run is what makes the split exact rather than bounded.
| Tree | Regions | Reported | Parse-backed | Fake |
|---|---|---|---|---|
talkbank-model/src/validation | 9,455 | 89.2% | 47.7% | 3,923 |
talkbank-model/src | 38,881 | 83.8% | 43.0% | 15,859 |
talkbank-parser/src | 21,229 | 69.3% | 68.8% | 110 |
talkbank-transform/src | 13,506 | 89.5% | 89.5% | 0 |
Read the third column, not the second. The validator reports 89.2% and less
than half of it is backed by parsing CHAT; for the model as a whole, 15,859
covered regions are reached only by tests that hand-build the AST they judge.
The parser and the transforms are almost entirely real, and the reason is
structural rather than cultural: those crates depend on a parser and
talkbank-model does not, so its tests cannot parse even when their authors
would prefer to.
That is the single largest fact about this repository’s test suite: the reported number is not the one that measures parsing-backed coverage.
The rows below are the --lib figures that the two-crate command produces:
| Tree | Lines | Regions | Branches | Functions |
|---|---|---|---|---|
talkbank-model/src/validation/ | 3121/5420 (57.6%) | 59.1% | 263/680 (38.7%) | 292/393 (74.3%) |
talkbank-parser/src/, hand-written | 4281/12295 (34.8%) | 34.2% | 478/1484 (32.2%) | 369/691 (53.4%) |
These are FLOORS, not the suite’s coverage. They were produced by --lib
runs of two crates. The spec-derived evidence (436 error fixtures, 138 construct
tests, 107 reference files) lives in talkbank-parser-tests integration
binaries, which were not in these runs; talkbank-transform, talkbank-lsp and
chatter were not measured at all. Do not quote these as chatter’s coverage.
The real figure needs
cargo +nightly llvm-cov --branch -p talkbank-parser-tests --tests, and the
first job of the attribution harness is to produce it.
The uncovered-branch report already exists in draft form: 501 rows for the validator and 1,078 for the parser. Their split is the useful part:
| Tree | neither side taken | false-only | true-only |
|---|---|---|---|
| validation | 401 | 51 | 49 |
| parser | 698 | 200 | 180 |
The 100 partial rows in the validator are the decidable ones, where the guard
fired but never declined or the reverse, and they are the immediate
NEEDS_SPEC_EXAMPLE worklist. The worst files by uncovered regions were
validation/utterance/phon_xtier.rs (333), validation/retrace/rendering/bracketed.rs
(199, 0% branch), validation/header/structure.rs (185),
validation/retrace/rendering/utterance.rs (172, 0%) and
validation/header/checkers.rs (161, 0%). Five functions had a ratio of 1.00
over more than 70 regions each, which means five functions nothing ran. The
two retrace/rendering files were a second serializer of the main tier from
which a marker’s location could be recovered, wrong on non-canonical spacing;
the parser records the marker’s own span (Retrace::marker_span) instead.
The precondition, and the deletion engine it unlocks
validation_errors_detected runs all 436 fixtures inside ONE test function. As
a class it attributes fine; per fixture it cannot, and “which spec example
reaches this branch” is the question the criterion asks. Converting it to one
rstest case per manifest fixture is about twenty lines, and the pattern
already exists in the same crate (reference_corpus_parses.rs uses
#[rstest] #[files(...)]).
That change unlocks the most valuable output of the whole exercise, which is not the uncovered list at all. With per-fixture attribution you run a greedy set cover over the 436 error fixtures and the 107 reference files and rank each by MARGINAL branches contributed. Every fixture contributing zero marginal branches is a deletion candidate. That is “volume is not evidence” made mechanical, and it is the pressure that stops the spec corpus growing without improving.
The six evidence classes run as six passes over ONE instrumented build:
cargo llvm-cov show-env, one --no-run build, then each class’s binaries with
its own LLVM_PROFILE_FILE and libtest filter, then one merge and export per
class. Because every class runs the same binaries, the region tables are
identical and the six exports join on region identity. The join result is one
bitmask per branch, and that bitmask IS the attribution.
One honesty note on the property class: proptest is nondeterministic across runs unless seeded. Run it with a pinned seed and case count and label the column with the seed, so the claim is falsifiable rather than a lucky draw.
What the criterion cannot cover
A claim of total coverage is the failure this exercise exists to correct, so this section is not a caveat, it is part of the design.
A region executed is not a region observed. This is the sharpest limit here and it is structural. The primary spec-derived gate over the entire validator asserts set membership of stringified codes; nothing reads a span, a message, a severity or a multiplicity. So across the validator’s functions and all 436 fixtures, a mutation that moves a diagnostic’s span from the offending word to the whole utterance, or emits it twice, or flips Error to Warning, leaves every branch covered and the suite green. The parser’s output is byte-pinned; the validator’s output is set-membership-pinned. That asymmetry is where the mutation budget goes:
cargo mutants -p talkbank-model --file 'src/validation/**' --timeout 180
Its survivor list is directly actionable: a mutant that makes a guard
unconditionally true turns a rule into “always fires”, and it survives exactly
when the code has no legal example. The survivors are therefore a worklist of
codes lacking a negative example, and the fix is a spec example rather than a
test. Claim::satisfied_by itself is the single highest-value mutation target
in the tree, three arms judging 436 fixtures, and nothing mutates it today
because it lives in the excluded spec workspace.
Coverage cannot see a false positive on an input nobody wrote. legal is
per example. There is no statement anywhere of the form “this rule fires on no
reference-corpus file”, and no coverage measurement can produce one.
Branch coverage is condition-level, not MC/DC. a && b yields two
conditions, not four combinations.
Coverage cannot distinguish an interesting value from a degenerate one. In a parser this bites at length and index boundaries: a span-arithmetic branch reached by a one-word utterance says nothing about an empty one.
Coverage cannot see the spec being wrong. This is the deep one, and the objection below is built on it. A defect both parsers share is invisible to the differential; a rule that encodes a wrong understanding of CHAT is invisible to everything in this repository. There is no gold reference anywhere: not the reference corpus, not CLAN CHECK, not a recorded verdict. Every comparison against a human artifact is AGREEMENT, not accuracy, and its ceiling is that artifact’s own reliability.
Coverage says nothing about the desktop or the LSP beyond the fact that their code was executed, which is why the boundary class exists and why it is the one class instructed to grow.
Type changes that delete tests
The compiler is the best failing test that exists: red before the change and green after, at every call site including the ones no test enumerates, failing at the mistake instead of later. Each row below is a change that removes a possible wrong value AND deletes tests, with the count.
| # | Change | Made unrepresentable | Tests deleted | What breaks |
|---|---|---|---|---|
| 1 | Word gets a structured body (segments joined by compound and clitic boundaries, at most one primary stress per segment, shortenings as balanced pairs); Word::new_unchecked becomes #[cfg(test)] and the fields go private | leading and trailing +, ++, misplaced stress and lengthening, secondary stress without primary, empty content, unbalanced shortening | 36 (34 in validation/word/{tests,snapshot_tests}.rs, 2 parser regressions) | 147 fabricating test sites; both parsers’ word builders move from mutate-then-patch to one fallible constructor; 11 error codes leave the model layer and become parse diagnostics; the WordContents wire format changes |
| 2 | Fragment offsets become three newtypes (DocumentOffset, SyntheticOffset, FragmentOffset) with the wrapper owning the only conversion, and a clamp that returns AsGiven or ClampedTo | rebasing a span into the wrong space, omitting the rebase, a silent clamp | 16 (all of context_public_api.rs) | four public parse_*_fragment signatures, and their LSP and CLI callers |
| 3 | The header preamble becomes typed slots plus nested GemScopes built once at the boundary | duplicate single-only headers, missing required headers, wrong header order, unmatched and mismatched gems | 35 (24 + 11 in validation/header/) | ChatFile::validate* takes a preamble instead of scraping lines, which forces row 4 |
| 4 | ChatFile drops its four cached header-derived fields and pub lines | a cached languages/options drifting from the lines they came from; a JSON roundtrip that omits them silently yielding ca_mode = false on a CA file | few directly | 23 ChatFile::new sites and every consumer reading the cached fields; deliberate wire-format change |
| 5 | GrammaticalRelation stores SemanticWordIndex1 and GraHeadRef, and Vec<_> becomes a validated DependencyForest | index 0, dangling head, cycles, multiple roots, non-sequential indices | 8-12 | 98 GrammaticalRelation::new sites. Receipt for the shape: the typed index_as_semantic() accessor is called ZERO times today while rel.head is read raw 29 times, which is a proof type built by its consumers and never by its producer |
| 6 | ParseHealth gains one Alignment enum with fn tiers() and an AlignPermit, replacing twenty hand-written mirrored predicates | forgetting an alignment gate; ParseHealthState::default() meaning “unknown” invisibly | 6 | 18 call sites |
| 7 | ValidationContext gets one FieldContext and one DeclaredLanguages, and the byte-duplicate get_other_language / is_tertiary_language pair is deleted | three of four field Options set and the fourth not; primary/secondary/tertiary decoded by list position; a dropped @Options: CA validating as non-CA | ~8 | the six helpers that rebuild the pair by hand |
| 8 | Span gains provenance (Source(SourceSpan) or Synthesized), with no Default and no numeric sentinel; SourceLocation’s two late-filled Options become a phase type fn locate(SourceIndex) -> LocatedError | a fabricated span indistinguishable from offset 0; a validator branching on is_dummy() and getting a wrong answer | few directly | 403 construction sites, overwhelmingly test scaffolding that row 1 removes anyway; 37 is_dummy() branches become decidable |
| 9 | Finish the typed traversal migration: no node.kind() string comparison outside the generated carriers | a forgotten node kind compiling cleanly and dropping a subtree | 87 (30 characterization and visitor files) | 262 .kind() sites; tokens.rs and its 17 tests go too if the grammar’s coarsened token(...) rules are un-coarsened |
That is roughly 210 tests deleted with a compiler receipt, before row 9, and about 300 with it. None of them is culled: each stops being writable.
The counterexample that keeps this honest. 65 tests in num2text.rs,
num2chinese.rs and clan_ud_mapping.rs are genuine behaviour tests of
algorithms no signature describes and no type can hold. Label them in source
with their class so the next reviewer does not spend a pass re-deciding. Not
every test is a missing type, and a design that cannot say which ones are not is
not a design.
What is deleted, with sizes
| What | Size | Why it goes |
|---|---|---|
| Reference-corpus whole-file JSON snapshots, orphaned half | 335 files, 2,277,490 bytes | They correspond to no live corpus stem. Nothing runs cargo insta test --unreferenced, which is why 76% of that directory can be orphaned undetected. |
| Orphaned insta snapshots from deleted test targets | 54 files, 53 KiB, in crates/talkbank-parser/tests/snapshots | Four prefixes, none of which exists as a test target in that crate. No test reads or writes them. |
| Reference-corpus whole-file JSON snapshots, live half | 104 files | Replaced, not merely deleted: see the decision below. |
| Duplicate roundtrip harness | direct_parser_roundtrip_corpus.rs, 107 cases plus a loop test | A verbatim duplicate of roundtrip_reference_corpus.rs over the same 107 files with the same parser; its DIRECT_PARSER_SKIP list is empty and its name refers to an integration that has happened. Two harnesses proving one property. |
| The repository-root error corpus | 26 tracked files under tests/error_corpus/, plus generate_error_corpus.rs | A second error corpus at a second root. Its generator refuses a root that does not contain the manifest directory. Four spec files still declare a source pointing at it. The 436-fixture spec corpus supersedes it; migrate the 19 parse-error cases into spec/errors/ and fix the four source lines. |
| CLI command-surface manifest | command_surface_manifest.rs + SURFACE_GROUPS, 5 tests | A hand-written second copy of a list clap already derives, compared by parsing help text. Its own comment records the failure: a wrapped description line produced a phantom command. Derive from clap’s command tree; keep the per-family coverage EXPECTATION as an attribute beside the command definition. |
| Duplicate word-validation snapshots | validation/word/snapshot_tests.rs, 7 tests | Same helper, same fixtures, same codes as tests.rs, asserted through insta instead. A second assertion of one behaviour. |
| Orphaned golden lists at the repository root | golden_words_featured.txt, golden_words_minimal.txt | Written to the CWD by an audit binary, tracked, read by nothing, disagreeing with both the crate copies and the book page that documents their counts. Three-way drift with no owner. |
| Authored construct CST blocks | 137 blocks, 41 of them naming node types the grammar does not have, plus ## Metadata, update_cst, and the second format reference | Dead data that has rotted. See the construct-claim decision above. |
| Characterization and visitor suites | 30 files, 87 tests | They pin a migration’s endpoint, captured by running the pre-migration parser, and one of them opens with a paragraph arguing it is not scaffolding while the next names its migration task. Delete AFTER row 9 of the type table, never before: they are load-bearing until the last .kind() comparison is gone. |
| The 20-case property tests | 12 of 18 proptest! blocks | Not deleted outright: either restore a real budget (256+) on the roundtrip and span properties and accept the runtime, or delete those blocks and keep their committed regression seeds as ordinary examples. A property at 20 cases over a 1-8 lowercase-letter alphabet is an example test with extra machinery. |
The one contested deletion, decided. The 4 MB reference-snapshot population
is 92% of all committed snapshot bytes and its diffs are unreadable, so a real
regression and a field rename look identical and both get accepted. That argues
for deletion. Against it: those snapshots are what kills parser mutants, and
deleting them would leave the parser’s structural output pinned by nothing while
the validator is already weakly pinned. Both are right, so the answer is to keep
the DETECTOR and drop the FORM: one derived artifact under the registry, one row
per reference file, carrying a hash of the canonical model JSON plus a
structural summary (counts by model node kind, diagnostic codes, roundtrip
byte-exactness), reviewed as a diff the way example-diagnostics.json already
is. Detection is unchanged, because any model change flips the hash. The honest
cost: when only the hash moves and the summary does not, seeing WHAT changed
means regenerating the full JSON against the parent commit. That is a two-command
operation, and it is the price of a review artifact a human will actually read.
Not deleted, wired. 39 JavaScript tests in the desktop app are invoked by no
workflow and no justfile recipe. Tests that no gate runs are worse than none,
because their presence reads as coverage. Either npm run test:unit joins the
gate or the files go. The Rust Tauri bridge tests stay and grow: they are the
specific hole a four-week desktop validation outage went through.
Migration order
Each step leaves the repository working and gated. Each names what it REMOVES, because a step that only adds is suspect.
Step 0. Make the instrument runnable.
Add llvm-tools to rust-toolchain.toml’s components, beside the existing
targets entry and for the same reason the comment there already gives: rustup
installs it, rather than a runbook asking a human to remember. Derive the
generated-file exclusion list from the artifact registry and consume it from both
the coverage invocation and mutants.toml.
Removes: the second, hand-maintained notion of “which files are generated”
(one file listed where eleven qualify), and one manual prerequisite from the
coverage recipe.
Step 1. Delete the dead weight and install the hygiene that stops it
returning.
335 orphan snapshots, the 26-file root corpus and its misdirected generator, the
duplicate roundtrip harness, the two root golden files and the audit binary that
writes to the CWD, the command-surface manifest, the duplicate word snapshots.
In the same commit: cargo insta test --unreferenced=reject joins just gate,
and a ratchet counts new_unchecked and Span::DUMMY construction in test code
and refuses an increase.
Removes: about 2.4 MB and roughly 470 files and cases, plus the ability for
any of it to come back. The ratchet is what makes every later step provable:
the number can only fall.
Step 2. Attribution.
Convert validation_errors_detected to one rstest case per manifest fixture.
Build the six-class attribution pass in xtask, emitting the row artifact with
every verdict empty. Publish the real coverage figure for the full suite, which
this page cannot state today.
Removes: the single opaque test that did 436 fixtures’ work anonymously, and
the state of not knowing what covers what. This is the one step that mostly
adds, and it earns that by producing the deletion list every later step consumes:
the marginal-contribution ranking of all 543 committed fixtures.
Step 3. Read the corpus once.
Run the attribution pass with the #[ignore]d production-corpus class included.
Record the branches reachable ONLY by it. Synthesize a fixture for each, seeded
from corpus/reference/ with the existing perturbation generator, which needs a
valid seed rather than a large corpus.
Removes: the 22 ignored corpus tests, the $TALKBANK_DATA dependency, and the
last route by which evidence is produced and discarded. After this step a
successor who cannot read the corpus inherits everything it taught us.
Step 4. The Word body.
Type table row 1. Rewrite the eleven affected specs’ claims from “implemented,
unreachable by this fixture” to parse-level violates examples, which is the
honest version of what they already say. Re-baseline the differential.
Removes: 36 tests, 11 model-layer emitters, a character-by-character rescan of
raw text inside the CHAT core, and 147 fabricating call sites.
Step 5. Coordinates, provenance and the preamble.
Type table rows 2, 3, 4, 8. Do them together: rows 3 and 4 are one change seen
from two sides, and row 8’s value is only realised once row 1 has removed the
scaffolding that constructs sentinels.
Removes: 51 tests, four cached fields that no longer have a source to drift
from, four Option fields on the validation context, one byte-duplicate function
pair, and the class of defect in which a fabricated span is indistinguishable
from a measured one.
Step 6. Construct claims.
Give a construct spec a model projection; make the generated test assert it;
delete the authored CST blocks, ## Metadata, update_cst and the second
format reference; relabel the sexp corpus as a derived pin in the registry’s own
vocabulary. Promote the exact per-stage code set to a claim where adjudicated,
and derive status from the observation snapshot.
Removes: 137 rotted blocks, one function that writes into a human-authored
directory, one duplicated format reference, one authored field the system can
observe, and 138 assertions that a string did not crash the parser.
Step 7. Finish the traversal migration.
Type table row 9, then delete the characterization and visitor suites.
Removes: 87 tests, 30 files, 262 string comparisons against node kinds, and,
if the grammar’s coarsened tokens are opened up, tokens.rs and its 17 hand
written string parsers.
Step 8. Close the criterion. Every row carries a verdict; the verdict list becomes the gate, ratcheting in both directions. Run the scoped mutation pass over the validator against a suite whose coverage is now attributable, and treat its survivors as the negative example worklist. Removes: the last unattributed branches, and the possibility of a green suite nobody can account for.
Step 9, standing. Stop committing what a generator can lower at test time, or exclude the lowered fixtures from review diffs. 436 committed fixtures are noise in every diff that touches a spec. This is a judgement about review ergonomics, not correctness, which is why it is last and optional.
Where evidence goes when you fix a bug
| The bug is | The evidence is | Where |
|---|---|---|
| a wrong verdict on some CHAT input | a spec example with a claim | spec/errors/E###_*.md, then just spec-gen, then adjudicate the observation diff as INTENDED |
| a false positive (we reject valid CHAT) | a legal example, which is the negative half nothing else can state | same |
| a wrong parse SHAPE | a construct spec with a model claim | spec/constructs/<area>/ |
| an invariant that holds for all inputs | a property, at a real case budget, seeded | property_tests/ |
| only reachable through the CLI, LSP or desktop | a boundary test in that crate | crates/chatter/tests/integration, talkbank-lsp, src-tauri/tests |
| one parser disagreeing with the other | a shrunk KNOWN_DIVERGENCES baseline | talkbank-parser-re2c |
| a CLAN CHECK adjudication | a manifest row plus a fixture, and the CLAN identity it was verified against | check_parity/manifest.json |
| an algorithm producing the wrong output | a unit test, labelled with its class | beside the algorithm |
Never: a new hand-built AST in talkbank-model. Never: a new whole-file
JSON snapshot. Never: a second corpus, a second manifest, or a second copy
of a list a generator owns.
And one rule for reading a failure before writing anything: when chatter
rejects a file, the working assumption is that chatter has a defect, not that
the data does, until the construct is shown to genuinely fail to make sense.
Editing data to satisfy a validator is how a parser bug becomes permanent.
How you know the suite is complete rather than green
Three questions, and each has exactly one instrument:
- Did anything RUN the code? The attribution artifact. Gated: no row
without a verdict, no covered row claiming to be unreachable, the
NEEDS_SPEC_EXAMPLEcount only falls. - Did anything OBSERVE what the code did? Mutation, scoped to the validator
and to
Claim::satisfied_by. Not gated per push (it is a report about the suite, not about a commit), run deliberately, its survivors triaged into spec examples. - Is what the code does RIGHT? The two oracles narrow it: an independent
second parser catches a defect in one implementation, and the CLAN ledger
catches drift against a recorded external verdict. Neither can catch a defect
both sides share. Beyond them the answer is human adjudication against the
format authority, recorded in the spec with a
sourceand anotes, and it is not gateable.
A suite that answers 1 and 2 is not blind. Nothing in this repository can make it right. Say that out loud in every report; a claim of total coverage is the failure this architecture exists to correct.
The strongest objection, and the honest answer
The objection. Coverage cannot see whether the spec is right, so a
completeness criterion built on coverage measures our own arithmetic rather than
the language. This repository could reach 100% branch coverage entirely from
spec examples that encode a wrong understanding of CHAT, and every gate would be
green while the tool is wrong about the format. Worse, the design deliberately
demotes the only three instruments that could catch that: the production corpus
is read once and then discarded, the CLAN grounding half is #[ignore]d and
depends on a binary a successor may not have, and human adjudication with the
format authority is explicitly outside every gate. The design therefore optimises
for a property it can measure (branches touched by committed fixtures) and
retreats from the property that matters (agreement with CHAT as it is actually
practised).
The answer, in four parts, and the first is a concession.
-
Conceded, without qualification. The criterion proves the suite is not blind. It never proves chatter is right. That is why the gate is a verdict and not a percentage, why “there is no gold” is stated in the limits section rather than buried, and why every comparison against a human artifact is called agreement rather than accuracy. A design whose strongest claim is “no branch is unaccounted for” is a smaller claim than “chatter is correct”, and it is the largest claim any repository-internal gate can support.
-
The instruments are demoted, not removed, and demotion is what makes them survive. A gate that needs a private corpus is a gate that a successor cannot run, which means in practice it is a gate nobody runs, which is strictly worse than an evidence pass with a committed output. Step 3 converts every branch the corpus alone can reach into a committed fixture: the corpus is spent once and its findings are permanent. The CLAN half keeps its grounding test and gains a recorded CLAN identity in the manifest, so “last verified against” is visible rather than assumed. The differential baselines shrink as ratchets rather than sitting as hand-typed constants. In each case the instrument’s OUTPUT enters the repository, which is the only form in which it outlives the person who ran it.
-
A wrong belief gets an address. The alternative on offer is not a suite that knows CHAT better; it is the current suite, which encodes the same beliefs implicitly across 2,868 tests, unaddressably, half of them over ASTs no parser produces. Under this design a belief about CHAT is one spec file with a code, a description, a rule statement, examples with claims, and a
source. When it turns out to be wrong, a successor changes one file and watches the observation snapshot show every behavioural consequence at once. That is not correctness, but it is the precondition for correcting anything. -
What genuinely remains uncovered. Nothing prevents a maintainer from writing shallow examples that touch branches and assert little, and no gate can detect intent. Three pressures bound it and none eliminates it: mutation kills a shallow example, because a fixture that touches a branch without observing its output lets the mutant survive; marginal-contribution ranking deletes fixtures that add no branches, so the corpus is pushed down as well as up; and the verdict is a written adjudication with a reason, which a later measurement can contradict in public. The residual risk is real and permanent: correctness against CHAT as practised is a question about the world, and the honest posture is to keep saying so rather than to let a green gate imply otherwise.
Two runners-up, answered briefly.
“Forcing validation tests through the parser couples two crates, so a parser defect can now mask a validator defect.” True, and intended. A validation rule that can only be triggered by an AST no parser produces is not a rule about CHAT through that parser alone. Independent authored expectations, legal and invalid controls, and separately admitted constructor/wire boundaries limit the masking risk; a second parser is not a required correctness oracle. Where no CHAT input reaches a guard, record that fact and assess all supported producers. Preserve genuine recovery and internal-failure handling until producer invariants justify removal. Do not invent CHAT fixtures to exercise tool failures or delete those failures merely to improve coverage.
“Branch coverage as a gate is Goodhart bait.” It would be, as a percentage. It is gated as a verdict per row, and the percentage is deliberately not a gate anywhere in this design. The number appears in this document exactly once, as a floor, with the reason it is a floor.
Appendix A: how every number here was produced
Run from the repository root. No number in this document was written from memory.
# Test attributes, repo-wide and per crate
rg -c '^\s*#\[(test|tokio::test)\]' $(git ls-files '*.rs') | awk -F: '{s+=$2} END{print s}'
rg -c '^\s*#\[(test|tokio::test)\]' $(git ls-files 'crates/talkbank-model') | awk -F: '{s+=$2} END{print s}'
# Fabricated-AST construction
rg -c 'new_unchecked' $(git ls-files '*.rs') | awk -F: '{s+=$2} END{print s}'
rg -c 'Span::DUMMY' $(git ls-files '*.rs') | awk -F: '{s+=$2} END{print s}'
# Construct specs, and the corpus cases derived from them. The `cst` fence is
# written both as ```cst and as ``` cst, with a space; a scan anchored on the
# first form misses the second and under-counts the blocks.
# The three coverage runs. The THIRD is what makes the fabrication split exact:
# it excludes the fabricating package from the RUN but not from the REPORT, so
# the regions only those tests reach are `full - without`.
cargo +nightly llvm-cov --branch --workspace --tests \
--json --output-path /tmp/cov-all.json
cargo +nightly llvm-cov --branch -p talkbank-model -p talkbank-parser --lib \
--json --output-path /tmp/cov-fab.json
cargo +nightly llvm-cov --branch --workspace --tests \
--exclude-from-test talkbank-model --json --output-path /tmp/cov-nofab.json
just coverage-attribution /tmp/cov-all.json crates/talkbank-model/src/validation
git ls-files 'spec/constructs/**/*.md' | wc -l
rg -N -c '^={80}$' grammar/test/corpus/generated/ -g '*.txt' | awk -F: '{s+=$2} END{print s/2}'
rg -N -c '^={3,}$' grammar/test/corpus/manual/ -g '*.txt' | awk -F: '{s+=$2} END{print s/2}'
# Committed snapshots
git ls-files '*.snap' | wc -l
git ls-files '*.snap' | xargs wc -c | tail -1
git ls-files '*.snap' | sed 's|/[^/]*$||' | sort | uniq -c | sort -rn
# Spec system
git ls-files 'spec/errors/*.md' | wc -l
rg -c '^\[\[example\]\]' $(git ls-files 'spec/errors/*.md') | awk -F: '{s+=$2} END{print s}'
git ls-files 'spec/constructs/**/*.md' | wc -l
git ls-files 'crates/talkbank-parser-tests/tests/error_corpus/validation_errors/*.cha' | wc -l
git ls-files 'corpus/reference/**/*.cha' | wc -l
# Error-code multiplicity (why per-code examples cannot prove per-branch coverage)
rg -o 'ErrorCode::[A-Z][A-Za-z0-9]*' crates/talkbank-model/src crates/talkbank-parser/src \
--no-filename --glob '!*generated*' | wc -l
rg -o 'ErrorCode::[A-Z][A-Za-z0-9]*' crates/talkbank-model/src crates/talkbank-parser/src \
--no-filename --glob '!*generated*' | sort -u | wc -l
# Generated code, for the exclusion list
rg -l --glob '*.rs' -e '@generated' -e 'DO NOT EDIT' -e 'do not edit' crates/ spec/ apps/
wc -l crates/talkbank-parser/src/generated_traversal.rs
# Branch coverage (nightly; stable gives regions only)
cargo +nightly llvm-cov --branch -p talkbank-model --lib --json --output-path model.json
cargo +nightly llvm-cov --branch -p talkbank-parser --lib --json --output-path parser.json
The coverage figures in this document came from those two commands with
generated files excluded in post-processing, and they are floors for the reasons
given in
Where it stands today. The construct-CST rot
count (41 of 137 blocks naming node types the grammar lacks) was produced by
checking each authored block’s node names against grammar/src/node-types.json;
the snapshot-orphan count by matching
talkbank_parser_tests__snapshot__<name>.snap against the tracked reference
corpus stems. Both are one-off scans; re-derive them rather than trusting these
numbers after any spec or corpus change.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Testing
Status: Current Last modified: 2026-10-07 (commit 5e895791)
What the test layers are and which one to reach for. The commands to run routinely, and what each costs, are in Developer Verification Checks; how they relate to CI is in Testing and Quality Gates.
Canonical output boundary contracts
The authored E370_split_retrace specimens pair complete repetitions and
corrections with deliberate turn splits. Repair contracts exercise all three
typed RetraceJoinScope policies: exact repetition, mismatching or short
prefixes, replacement targets, correction opt-in, intervening headers, and
three-turn repetition chains.
Successful matching joins must equal the independently authored unsplit model;
refusals preserve the original model, and a second repair is a no-op. A repair
policy is not evidence of speaker intent or a waiver of the input’s E370.
The repair scan owns the untouched source suffix and the resulting line prefix.
Only adjacent utterances can enter admission. An EligibleJoin exclusively
borrows the target and owns the successor until consumed, so mutation never
looks up raw indices or rechecks endpoints. Refusal returns the untouched
successor. Header barriers and chain behavior remain policy contracts; the
types prevent endpoint invalidation, not mistaken repair-policy choices.
The canonical test harness enables both model async and channels features.
Channel transport controls use legal, parse-invalid and validation-invalid
specimens: connected delivery must preserve ordered diagnostics, and a dropped
receiver must not change the validated proof/refusal. Sender ownership closes
the connected stream before it is drained. These transport contracts do not
certify every feature combination or complete macro/source attribution.
The single MOR-word fragment boundary rejects source-bound slices containing
post-clitics or multiple morphology items from the valid mor-gra reference.
An admitted single-word projection must contain exactly one item and no
post-clitics; it cannot silently discard either. The refusal test verifies
diagnostic multiplicity, code and caller-coordinate rebasing. Existing reference
item contracts retain acceptance of each independently selected MOR word.
GRA-relation, PHO-word and participant-entry projections use the same private
single-item admission. Reference-derived multi-item inputs must be refused;
PHO groups cannot stand in for a word. Empty input produces diagnostics at the
requested insertion point. A direct single-relation tier control establishes
that GRA needs no synthetic PUNCT relation, so the wrapper adds no such
scaffold and filters no diagnostic for it. PHO likewise needs no appended dot. Passing caller coordinates
directly into these tier wrappers avoids losing zero-width positions through
an intermediate 0..0 span.
Complete dependent-tier fragments also require full-input coverage. A private
parse-owned phase associates the selected tier with its wrapper, admits only
one tier and checks that no non-whitespace caller bytes lie outside its span.
Both generic and content-only adapters consume that phase. Reference slices
containing adjacent tiers or a tier followed by speech must be refused with
rebased diagnostics; normal tier and continuation inputs remain covered by the
complete reference fragment contract. Wrapper-supplied tier labels are outside
the caller range and are accounted for by this coverage check.
The reference body workflow also checks the concrete constructors and allocating output helpers of all nine ordinary text-tier types. Structured construction preserves the complete body and its spans. Plain-text construction is exercised only after matching a single typed text segment, never by flattening media or continuations; it preserves semantics and output without claiming source spans. Each concrete type requires its own reference witness, including the seven types supplied by the shared text-tier macro.
The canonical serialization-sink workflow includes content-only %pho,
%mod and %gra writers alongside %mor. Parsed reference/spec tiers supply
the values; the test compares allocating, streaming and full-prefix output,
then refuses each actual content-write boundary and requires error propagation.
JSON replay must preserve content, while %gra parser-completeness provenance
returns as Unknown. Recovered spec tiers are serialization inputs here, not
certificates of valid CHAT. Reference roundtrip tests separately own correctness
of the canonical spelling.
Body-only output for bullet-capable dependent tiers uses the existing media and
multiline references. The contract requires text, timing, picture and continuation
witnesses, preserves typed segments through JSON, and compares allocating and
streaming output while checking every sink-refusal boundary. These are ordinary
tier payloads; %com receives no special editing or validation semantics.
Coverage source omissions
An LLVM export describes its linked, instrumented code, not automatically every source file or feature configuration. Compare its paths with the tracked source inventory before claiming complete scope. A missing file can contain re-exports, declarations, test-only helpers, or macro inputs rather than uncovered runtime statements; file counts are not executable coverage percentages.
Derive-generated methods need independent evidence too. A compiler-expansion
and executable-symbol audit found the SemanticEq, SemanticDiff and
SpanShift implementations for both ActTier and UtteranceContent in the
runner but not in its LLVM function records. Instrumented collection helpers
do not substitute for the missing element-method records. This is a measured
attribution limit for those six methods, not a claim about every compiler or
every derive. Keep generated behavior separate from handwritten source totals:
use derive-owner shape tests and reference-backed model workflows, and retain
the broader expansion-completeness obligation. Do not modify runtime code just
to manufacture coverage mappings or silently adjust the denominator.
SemanticEq derives both equality and structured diff implementations. Its
owner tests compare unit, single/multiple tuple and named enum variants against
an independently specified semantic partition, including skipped metadata.
Named/tuple structs check exact diff paths; skipped source spans remain useful
diagnostic context without becoming semantic differences. The feature-gated
talkbank-derive UI suite separately checks accepted inputs and compile-time
rejections. These are generator contracts, not canonical CHAT coverage.
SpanShift owner contracts cover unit, tuple and named enum variants, nested
optional spans and insertion/deletion reversibility. A skipped payload does not
need to implement SpanShift and remains untouched. The compile-pass fixture
denies unused variables: skipped tuple fields must emit wildcard patterns, not
unused bindings. The generator retains an optional binding per field and emits
shift calls only for retained bindings, sharing one skip decision between the
pattern and recursive calls.
The coverage_source_ranges example accepts a JSON array of repository-relative
Rust paths on stdin. It parses each file with syn and binds byte ranges to the
exact source SHA-256. Alongside explicit test-only ranges, non_test_items
distinguishes written function bodies, unexpanded macro tokens, non-doc attribute
inputs, and external module declarations. Trait signatures without bodies are
not counted as written bodies. Unknown cfg expressions remain candidates.
This is a syntax inventory, not macro expansion or cfg evaluation. A helper in
an external test module can appear as a candidate until its parent’s test-only
declaration is traced; names such as tests.rs do not prove that relationship.
Macro and attribute inputs must not be silently reclassified as non-executable.
The tool’s Rust/wire boundary tests are measurement evidence, not canonical CHAT
coverage gains or a substitute for spec/reference-corpus execution.
The adjudication corpus contract derives pending requests from actual parsed
reference speakers, then supplies authored operator choices. Prompt exhaustion,
out-of-order sessions, decision-kind mismatches, and missing rename roles must preserve every
uncommitted request and the already accepted override prefix. Retrying must
resolve only the remaining requests. This exercises the public in-memory
workflow, not filesystem transactionality or inferred speaker-role truth.
Each resolved override is then admitted through to_mapping_spec and applied
to the corresponding parsed document. The authored identity mapping must
preserve the original CHAT bytes, connecting queue recovery to its real consumer.
The separate CLI wire test requires missing-role refusal to exit with status 2,
preserve the pending file byte-for-byte, and create no override file. It is not
counted as canonical CHAT coverage.
The same reference-backed workflow admits authored operator mapping strings, with and without token whitespace, and compares them with the override record’s typed mapping. Both routes must preserve the source bytes for an identity mapping. Drop commands must retain every other speaker’s turn semantically after parsing the output. Empty assignments, stray commas, missing equals signs, missing role separators, and repeated source assignments are typed refusals. Both identical repeats and conflicting rename/drop commands are rejected; the mapping parser only inserts through a vacant map entry, never replacing an earlier assignment. This is not a claim that mapping syntax admission validates speaker identities or full output validity.
Build artifact hygiene and runner choice
The E344 paired specimens change only an intervening speaker code. They pin
the nearest-same-speaker boundary under opt-in strict quotation validation,
while a policy-boundary contract proves default validation emits no E344 and
the strict diagnostic points to the attribution turn. CHECK accepts both
specimens; its acceptance is recorded separately from Chatter’s optional rule.
The same policy/location contract covers first-turn self-completion (E351) and
consumed interruption reuse (E352). Paired E352 specimens verify that a turn
can consume an earlier interruption and then issue a new one with its own
terminator, while an extra completion cannot reuse either consumed interruption.
The validator owns one per-speaker history: map occupancy establishes a prior
turn, and non-copyable interruption tokens are issued only from typed +/.
terminators. There is no second last-seen map or unused utterance index stack.
E354 adds an actual trailing-off control and ordinary/CA terminator deletions.
The contract confirms typed None for both deletions and checks strict-policy
diagnostics on the following completion. CA waives E305, not the requirement
for explicit trailing-off evidence under E354. These specimens reach the
missing-terminator path naturally; that path must not be removed as impossible.
The E545 empty birth-date control similarly reaches a real allowed-empty
boundary. Its contract requires a retained typed Header::Birth with empty
date text and no diagnostics, rather than inferring success from a missing
error code alone. CHECK accepts the control; no fabricated date or dropped
header is allowed to stand in for unknown information.
CA delimiter traversal uses the shared ContentStructure walk and typed
WordRef::words ordering rather than separate main-tier/bracketed variant
lists. The delimiter policy covers replacement targets as well.
E230’s paired replacement specimens retain this policy despite CHECK’s opposite
results. Source analysis traces CHECK’s duplicate scan of replacement text:
once inside its bracket token and once after re-entry into the target words.
The authoritative CHECK assessment
defines adjudication policy and completion; this chapter records test mechanics.
Underline traversal remains
separate because it requires leaf marker payloads that this view does not carry.
Word validation also uses one structural dispatcher at every depth, establishing
each item’s annotation-selected language scope before validating its payload.
Replaced words retain their own validation rather than being flattened. Group
and retrace payloads provide enclosed content infallibly; nested-quotation
search similarly receives known enclosed content, not a possibly leaf node.
The E542/E546 specimens also exercise typed unsupported header values. Header
dispatch selects Unsupported once and passes its payload to the diagnostic
reporter; reporters do not reclassify an already-selected enum. Valid and absent
values remain distinct from unsupported text, which is preserved for roundtrip.
E546’s absent-field/comma-only pair proves that distinction through the parsed
Option<SesValue>, not just diagnostic presence. CHECK source review explains
its different token-list policy: commas and whitespace are skipped, so a field
containing only separators has no token to reject. This is distinct from the
duplicate-traversal execution defect in the CA example above.
E725’s complete two-word control and final-source-word deletion cover the
opposite cardinality direction from its original example. The deletion reports
E725 and the independent main-to-model count error E733, not E737: an absent
source word cannot be compared with reconstructed text. This is a reachable
absence guard, not a candidate for removal based on clean-input coverage.
E256’s shortening mutations place each curly single quote inside a larger CST
error region. The generated CST and stage observations pin recovery’s specific
E256 report; only the legal shortening control is byte-exact. Current CHECK’s
runtime results and isBadQuotes source path agree on the character-validity
policy without serving as Chatter’s parsing architecture.
E220’s adjacent-fault specimens additionally keep digit validation independent
of quote recovery. CHECK’s three-byte quote advance plus its loop increment
skips the following byte: it misses the digit in hello’3, but catches it in
hello’x3. Chatter reports E256 and E220 in both. This source-derived mutation
family records an execution defect rather than treating CHECK silence as a
new legal language context.
The E761 subtype pair changes one character in the universal dependency head
while retaining a multi-part subtype. Its contract inspects the typed relation
and checks the diagnostic’s triple, separated head, severity and tier span.
CHECK accepts both; that observation does not redefine the stricter UD rule.
The canonical incremental_corpus family compares producer-owned revisions
against cold parsing across the finite reference and diagnostic populations,
including complete deletion and restoration. Both models (including spans) and
ordered diagnostics must match. A separate raw-tree misuse witness retains
the range-refusal boundary: an unedited stale tree is not a revision proof.
Three small reference files additionally supply a finite typing/backspacing
sweep: clitics/compounds, Unicode IPA and multiline continuation. Every UTF-8
character boundary is visited in both directions through the revision owner.
Each partial buffer must match a cold parse in model and ordered diagnostics,
retain its exact source, and avoid internal-tool errors. Partial buffers are
not asserted valid, nor are observed diagnostic codes promoted into goldens.
The complete references remain the controls; no production-corpus scan is used.
An interior-edit deck uses the MOR/GRA, PHO-grouping and SIN-grouping references.
Parsed dependent-tier content spans select each Unicode scalar for deletion,
followed immediately by restoration. The rest of the document stays present,
exercising a different incremental-recovery shape from prefix typing. Each
revision must match the cold model and ordered diagnostics; restoration must
recover the complete original model without residual diagnostics. The deck
includes diagnosed and undiagnosed edits, but does not assert that every
undiagnosed edit is semantically valid CHAT or invent new diagnostic goldens.
The same comparison owner exercises identity-header edits selected from parsed
@Participants, @Languages and @ID spans. Multilingual speaker metadata and
Portuguese participant names supply separators, identifier boundaries and
multibyte text; deleting header prefix/newline characters keeps following speech
in place. Every edit is restored before the next deletion. These contracts
test editor recovery consistency, not a separate parser or a new CHAT policy.
Main-tier spans are exercised through that same owner using the nested-group,
retrace/replacement and timing-bullet references. Interior mutations keep later
speech and dependent tiers present, including when a deleted prefix or newline
changes their grammatical attachment. Model/spans and ordered diagnostics must
match a cold parse, and restoration must remove all transient recovery.
The LSP consumes the same revision owner; its focused tests cover editor
changes and source-bound diagnostic presentation. Its manual phase benchmark
measures edit, parse, and lowering together through that same transition.
Retired external corpus
TalkBank/testchat belongs to the retired Java Chatter workflow. Do not use it
as the current Rust Chatter conformance target, update it for new rules, or
infer validity from its good, bad or check-good directories. Preserve it
as historical evidence, not an executable authority.
Current expectations belong in spec/constructs/, spec/errors/ and
corpus/reference/, with reviewed claims and owner-generated fixtures. If an
old example exposes a missing case, adjudicate it against current rules and
promote the justified case into those canonical sources; do not copy its old
good/bad label as the expected verdict. This policy does not claim the current
finite corpus is already complete or that the external repository is archived.
The recursive validation-event test uses owned root/nested copies of a canonical error spec and checks both paths and diagnostics. It does not depend on a developer’s external corpus and never silently skips.
Corpus-backed contracts
E541’s clock-boundary family pairs valid two-/three-component start times with single-component overflows and a bare-seconds shape; five of its values are ones CHECK rejects. Time-start assessment uses the shared clock-range predicate as well as shape admission and issues a borrowed refusal capability; the E541 renderer cannot accept unassessed text. Invalid parsed values remain available for byte-exact roundtrip, rather than being discarded or relabeled as syntactic recovery.
E540 duration assessment likewise issues a borrowed refusal before diagnostic rendering. Its canonical duration examples exercise unsupported values, shape refusals, clock bounds and legal controls through header-only validation and JSON roundtrip. The numeric-boundary deck checks both malformed semicolon endpoints and surplus clock fields: neither may be discarded to recover a valid prefix, and serialization must preserve the original spelling and findings. A separate public-model boundary check clears only the duration in each parsed spec model using its constructor, then roundtrips JSON and retains the existing optional-empty policy. That API/wire check does not claim an empty raw CHAT header parses cleanly or certify its serialization as a valid transcript.
The E537/E538/E539 vocabulary families enumerate supported @Number,
@Recording Quality and @Transcription tokens, alongside near-miss numeric,
case and punctuation variants. Controls are diagnostic-free; malformed values
are preserved byte-exactly and refused during validation. CHECK rejects eight
of the nine mutations with code 11, but accepts eye-dialect, which Chatter’s
existing exact-token rule rejects. Exhausting these vocabularies improves their
parse/serialization coverage; it does not exhaust header recovery states.
E546’s SES vocabulary control exercises every declared ethnicity and all four
socioeconomic codes through real @ID headers. Component substitutions and
deletions retain the comma, separating unknown vocabulary from missing parts.
Both unknown components are rejected by CHECK (144) and Chatter (E546);
the incomplete pairs are an existing stricter Chatter boundary. A separate
space-delimited control records accepted normalization to comma, not byte-exact
roundtrip. This family adds measured source coverage in SES classification
and serialization without changing the implementation or denominator.
E246’s lengthening-category matrix pairs two and three colons in ordinary
words, fillers and phonetic @u fillers. All six parse without diagnostics
and roundtrip byte-exactly. CHECK (21-Sep-2026) rejects only the three-colon
ordinary filler with code 48. Preserve this shape-specific difference rather
than imposing CHECK’s category-dependent colon limit on Chatter’s existing
repeated-lengthening rule. These examples improve the finite behavioral
inventory but added no line, region or branch coverage in the measured run.
E209’s CA-shortening pairs contrast a single (ab) with adjacent (a)(b),
both directly and inside a replacement annotation. All four parse and roundtrip
byte-exactly. Only the single-shortening shapes normalize to CA omissions;
the double-shortening variants reach E209, including replacement-specific
spoken-content validation. CHECK accepts all four, so these are coverage
witnesses for an existing behavior difference, not a CHECK-parity gain.
E203’s embedded-marker mutations reuse its legal shortening controls. Inserting two markers inside a shortening, or inserting one inside a word that already has a valid outer suffix, reaches the model’s repeated-marker recovery check. Both malformed sources produce parse E316 and validation E203; CHECK rejects both too. An undeclared outer suffix and an embedded marker are different producer states: coverage of one does not justify deleting recovery for the other.
E243’s word-control family inserts U+0007, U+007F and U+0085 into a printable word control. Bell is rejected during parsing; delete and next-line controls are retained and rejected during validation as well. U+0085 witnesses the word-whitespace validation path directly from CHAT source, so that path must not be dismissed as requiring synthetic model construction. CHECK’s acceptance of the latter two variants is recorded separately from Chatter’s validity rule.
E342’s empty-scoped-content pairs remove the sole word from a retraced or
explained group while preserving subsequent speech. They pin E342 recovery
and the empty retrace’s additional E378. Source emptiness must not be confused
with an empty model collection: these cases do not demonstrate the serializer’s
empty BracketedContent path, which remains a separate coverage residual.
The E305 bullet-retention specimens separately assert timing preservation:
internal plus terminal bullets, consecutive bullets without a terminator, and
an internal bullet immediately before a terminator. The regression checks each
model position and byte-exact serialization even for the invalid specimen.
Lowering installs the grammar-owned terminal slot first; fallback extraction
transfers at most one owned bullet only when no terminator or terminal bullet
already establishes the boundary. A non-bullet tail is returned unchanged.
CHECK’s separate error 73 for empty inter-bullet scopes is not adjudicated by
these E305 claims; preserving evidence is not certification of every rule.
Additional paired controls isolate multi versus multiple options (both
unsupported E534 in Chatter), explicit 0 between timing scopes, and a leading
bullet. The CHECK witness for a bracket-first code does not establish parity
for these bullet-specific paths; keep each observation tied to its shape.
E770 covers the leading-bullet shape. Its canonical contract checks exact diagnostic counts and original bullet spans, nested retrace traversal, explicit zero/event/pause controls, linker non-material, and recovery uncertainty. The state machine distinguishes awaiting material, established material and unknown preceding material after main-tier recovery. It does not reset after each bullet or implement the separate inter-bullet option policy. The invalid leading-bullet retrace also roundtrips byte-exactly: the container’s first/rest split owns separators, and a bullet leaf adds no leading space of its own. The recovered stray-bracket example retains its parse refusal and is excluded from byte-exact roundtrip claims.
The E305 timed-terminator specs pair the media-bullets reference with two source-bound single-token deletions. They retain dependent-tier timing, pictures and continuation text. Both mutations parse cleanly but fail the missing-terminator validation rule; they also exercise terminal-bullet extraction at the parse-to-model boundary. A runtime-tool contract checks that the promoted fixtures are byte-identical to the admitted seed and its generated deletions. This provenance check is separate from the canonical fixture runner’s claims.
The repetition-recovery E375 specs intentionally combine retired [x N]
notation with missing terminators, with and without spoken material. They
complement single-fault cases; their claims do not imply a particular CST
recovery location or a new coverage gain. Keep observed recovery and intended
validity separate when promoting mutation candidates.
E330’s free-text recovery specs retain a shared %eng/%com/%xnote control
and bare-bullet-opener mutations in each dispatch family. Their typed
clean/recovered test cases distinguish valid roundtrip from refusal evidence:
recovered tiers must report E316/E330 at the damaged line and preserve healthy
siblings. %com retains an empty tier (also E756); %xnote is omitted after
its parse refusal. These witnesses do not justify removing other recovery arms.
Disk validation obtains its named context from StoredTranscript, not the
argument’s Unicode spelling. StoredNameResolver reuses directory snapshots
within a run and invalidates them when the directory timestamp changes;
resolution failures are explicit I/O failures, never anonymous validation.
The snapshot is retained through map-entry ownership, with no fallible second
lookup. Canonical media-control bytes also exercise directory-entry admission,
snapshot refresh after a rename, missing-name refusals, and symlink identity.
The symlink test checks the link’s own basename, not its target’s name. Linux
additionally exercises non-UTF-8 filesystem names; normalization-sensitive
filesystems must refuse nonexistent aliases rather than inventing a match.
CLI filesystem tests compare NFC and NFD aliases with directory traversal on
normalization-insensitive filesystems. They also verify that fixing content
does not rename a file or normalize a media URL.
The W109 catalog uses the generated, source-bound MediaHeaderNode fields to
admit one clean filename token. catalog_fix consumes a borrowed ParsedSource
retained by parse_chat_file_with_source, rather than an independently supplied
string. The CLI retains this owner from its original parse; catalog planning
does not create a parser or reparse unchanged input. Missing source ownership
refuses planning, while changed output still undergoes verification. Diagnostics
must still come from that input: the capability binds CST and source, not an
arbitrary diagnostic to its producer. Its private header-edit capability does not
license arbitrary header edits or recovered headers. Canonical corpus tests
compare the repaired bytes with the authored NFC control, then revalidate
under the same transcript identity: a remaining file-only warning is expected
when the stored basename still needs normalization.
The E241 partial-recovery example tests fixing a clean first utterance while preserving malformed morphology in a separately owned second utterance. The CLI contract checks both parse-health states before expecting a repair. An unstructured following main-tier fragment may instead taint the enclosing domain; its refusal control requires unchanged bytes. Apparent line boundaries do not authorize a narrower recovery domain or a guessed repair.
E604’s authored controls place an ordinary dependent tier before a continued
GRA tier, with an independent expected-output control. The catalog locates
the unique complete GRA node inside the typed utterance; it does not assume
physical adjacency or treat the intervening tier specially. Multiple GRA nodes
refuse selection. The CLI verifies explicit semantic opt-in and byte preservation.
Missing participant/role/language facts never produce guessed proposals; the
canonical diagnostic sweep retains refusal witnesses for reachable cases.
The E306 empty-content control contains a grammar-recovered missing node;
its diagnostic does not authorize deletion of the recovered main tier.
The separator-only E306 specs supply a cleanly parsed semantic-invalidity
counterpart. Deletion remains an explicit semantic choice and is proposed only
when the complete source-bound utterance is exactly its main tier. An attached
dependent tier refuses the proposal, preserving its content and ownership;
%com is an ordinary dependent tier here, not a special diagnostic channel.
The repair contract checks the exact main-tier range, successful admission for
the isolated turn, and diagnostic-free parsing after that turn is removed.
W109’s named canonical specs cover media-only, transcript-only and both-side Unicode normalization, canonical controls, identical-decomposed warnings, and a real filename mismatch after normalization. The diagnostic contract consumes the manifest’s authored transcript identity, not the generated storage filename, and checks side-specific advice and unchanged CHAT serialization. Explicit comparison outcomes rule out a warning with no noncanonical side.
The E307 speaker boundary matrix separates length and ASCII constraints, then combines their mutations. Its diagnostic contract checks that each reported source site retains both applicable findings, including recovered prefixes. A multibyte seventh character must not be mistaken for two characters. CHECK observations are recorded separately: agreement about non-ASCII rejection does not establish agreement about the existing seven-character length limit.
The E243 punctuation specs pair valid CHAT controls with deliberate mutations:
Unicode ellipses and standalone slash words, including nested and replacement
positions. Their shared boundary contract checks exact diagnostic multiplicity,
source spans and unchanged serialization. The slash controls retain repetition
annotations and free-text %com: slashes, so a broad slash ban cannot satisfy
the contract. CHECK observations establish rejection, not diagnostic-count
identity: Chatter reports each invalid word once.
The canonical parse, fragment, incremental, validation and roundtrip families
exercise source-preserving document/header dispatch and participant/media lowering.
The scalar/text family and three single-value option headers use generated
associated payload projections too. Existing valid and malformed header specs
exercise those entry points across full-file, fragment and incremental APIs;
the shared reader retains missing/error/absent and range-refusal handling.
Participant/value headers use the same source reader with a generated
speaker-kind constraint and named admitted fields. Canonical fragments retain
their recovery diagnostics. The reader still refuses on the speaker before
attempting the value; semantic date/language checks remain separate from source
admission.
Comment bodies also retain generated association through admission, with
non-present content still retained as an Unknown header. Their bullet-text
adapter is an explicitly separate boundary; source binding does not license
discarding its recovery segments. There is no unbound HeaderSite constructor.
Media filename/type/status and whitespace-before-comma specimens retain their
diagnostics and typed payloads through full-file, fragment and incremental APIs;
source-bound payload reads do not replace recovery or filename admission.
The media-field deletion specimens in E342 and E535 distinguish source
association from lexical admission: a present, zero-width generic media value
can reach an Unsupported model variant without a parse diagnostic. An omitted
optional status and an empty status introduced by a comma are different states.
Keep both their authored controls and mutations; neither a typed CST identity
nor absence of MISSING recovery proves a nonempty, supported value. CHECK’s
token-skipping policy is documented separately from Chatter’s validity policy
in the E535/E536 reference pages.
The E303 multiword-header control/deletion pair also guards the structural
MISSING-tab route. The generated header_sep slot proves which token is
missing, permitting E303 instead of generic E342 without guessing from ERROR
text. The backstop retains other missing-token diagnostics and does not treat
the recovered, separator-inserting serialization as valid authored CHAT.
The companion @Tape Location pair instead produces a whole ERROR header:
the same missing-separator policy needs both shapes, not a fixture assumption
that every malformed header receives an identical CST.
Underline validation consumes ContentStructure throughout. Its leaf view
retains opening/closing identity and optional source span, while group/retrace
variants carry their contents without an optional-container check. Existing
E356/E357 specimens cover word-internal, nested standalone and replacement-target
markers; their canonical contract additionally checks the exact authored marker
span in each unmatched diagnostic. These are pairing/location policy tests, not
tests that duplicate the structural invariant enforced by the enum variants.
Prefix-marker diagnostics similarly require an admitted illegal position, rather than accepting a legal position and relying on a preceding boolean check. E762/E763 specifications continue to own the position/language policy; no fixture should be invented to reach a diagnostic description for a legal position that the admitted type cannot represent.
The E763 language-header deletion deck guards against treating unresolved
language as “no language allows this”, which would emit E763 alongside the
genuine missing-header E504 even with full local coverage. The admitted
language-refusal value requires at least one candidate and no permissive
candidate. The finite deck preserves both sides: missing evidence suppresses
only the language-specific diagnosis, while an explicit @s:eng word marker
still establishes that diagnosis without a file language header. Coverage
closure does not replace these policy counterexamples.
The ambiguous-language E763 pair additionally distinguishes eng&heb (one
permitting alternative) from eng&fra (neither permits the marker). Both parse
and roundtrip cleanly; only the second emits E763. Owned governing marks
dispatch their actual marked payload directly to the borrowed resolver, leaving
the utterance-default variant on its separate no-marker path rather than
reclassifying an already matched value.
Overlap anchors retain the main-tier span captured during extraction. Orphan validation consumes that origin directly rather than reconstructing it with an index into a second utterance view. The seven E347 specimens pin the exact diagnostic tier for top/bottom/index-mismatched orphans and retain matched, unindexed and one-to-many controls. The producer-owned span is a snapshot of model provenance, not certification that arbitrary model spans are valid ranges in a separately supplied source string.
Existing participant recovery specs remain the witnesses for missing speaker placeholders and doubled-comma ERROR groups; source association does not justify removing them. Generator reconstruction tests separately verify source-bound choices, groups, extras and ERROR-root extraction against real parsed trees.
Conformance inventory admission distinguishes generated positional carriers
from generic runtime wrappers by their declared shape, not a Children name
suffix. A carrier or choice may declare the tree lifetime followed by the
defaulted range-phase and kind-proof axes, and is inspected as X<'tree>, the
reading the dispatch produces; any other generics are refused, never skipped.
The boundary regression also retains collision-suffixed carrier names. This
generator-input check is not a new CHAT specimen or a canonical coverage gain.
The authored edge-cases/bracketed-content-combinations.cha reference combines
existing constructs inside groups, including annotated and retraced quotations.
It is not a new production attestation. It exercises both serialization and
generic transform consumers; cross-parser semantic comparison additionally
guards marker attachment and ordering. Parser admission checks separately
ensure quotation markers do not consume word replacements or tier-level codes.
The same fixture includes an internal timing bullet and a nested replacement.
The same refusing-sink contract also runs over every canonical error-spec fixture, including its recovered model. These outputs are not treated as valid CHAT or as byte-exact reconstructions of invalid source. The contract checks only prefix preservation and immediate error propagation at every actual write boundary, with separate witnesses for clean and recovered parsing. Existing spec claims and roundtrip observations remain the semantic authorities.
The canonical validation-proof workflow exercises accepted and rejected models from both reference and error-spec fixtures. Accepted proofs serialize to the same CHAT and JSON as their payloads; neither wire format carries validation authority. JSON-decoded utterances retain unknown parser provenance and cannot be admitted merely because the original model was accepted. Consuming a proof before editing preserves the content, while consuming a rejection makes that content available for repair. Rejection presentation retains all diagnostics and distinguishes incomplete parsing from ordinary model-validation failure. These are model-boundary contracts, not permission to ignore source parsing diagnostics in a file pipeline.
Accepted-proof CHAT output and rejected-model diagnostic presentation also pass through the one-way refusing sink. Each observed write boundary is refused once; the formatter must stop immediately and preserve the accepted prefix. Canonical specs supply separate witnesses for accepted models, ordinary model rejections and incomplete parsing. This tests fallible reporting without turning any rejected model into a valid transcript.
Same-speaker overlap validation consumes the extractor’s closed top/bottom
region kind, not a pair of arbitrary marker kinds with an ignored end marker.
Endpoint completeness still comes from the extracted region’s is_well_paired
check: this type restriction does not authorize dropping malformed overlap
recovery. Existing E704 specs and CA reference workflows retain the behavioral
obligation. Removing an impossible input branch is recorded as a denominator
change, not as new fixture coverage or new CHECK parity.
reference_serialization_propagates_every_writer_refusal parses each reference
file and records its normal serialization writes. It then refuses each observed
write boundary once, requiring error propagation and preservation of the
already-accepted output prefix. A refused sink cannot resume writing. This is
reference-backed output-boundary coverage, not an invalid-CHAT golden corpus or
evidence of operating-system I/O behavior; spelling and semantic preservation
remain the responsibility of the independent roundtrip checks.
The same parsed references also exercise the public Line writer for both
header and utterance variants. Its output must equal the owned payload’s
serialization and propagate every observed refusal. This checks enum dispatch
that whole-document writing may bypass, without another corpus parse pass.
The CA overlaps reference pins the public utterance-item overlap queries to
authored item positions: opening and closing, top and bottom, indexed and plain,
plus lexical negative controls. Attached closing marks in b⌉ and h⌋ remain
inside their words and do not make the enclosing word a standalone overlap
item. The contract checks those retained word-content markers separately;
classification must respect the model’s structural level.
The missing-@End spec supplies a real zero-width EOF diagnostic with and
without a final newline. Display enrichment must retain its exact source span
and the source index’s EOF line/column, rather than treating EOF as out of bounds
and moving the caret to a previous byte. The corpus-wide coordinate contract
also admits EOF when independently counting source newlines and byte columns.
Deleting that spec document through the incremental revision API supplies the
adjacent empty-editor-buffer case. Parsing and validation must still report
diagnostics; enrichment with an empty source index must preserve their complete
serialized payload rather than inventing coordinates or display context.
Before enrichment, the spec-wide diagnostic contract also admits every source
label through checked source-location construction and verifies UTF-8 boundaries.
Display-relative labels are checked separately against the resulting context;
passing that display check alone would not establish valid producer coordinates.
Compound word content delegates output to its typed marker’s writer rather than duplicating the delimiter. Reference word-wire and output-refusal contracts check its spelling, preserved spans and propagation of sink failure.
The reference group traversal also checks the public bracket-content writer
directly, apart from each enclosing group’s spacing and decoration. Existing
grouped actions and annotations supply its inputs; every observed write can
refuse without losing the accepted prefix. The word-wire traversal attaches
caller-owned alignment IDs to parsed words and requires JSON to change only
the word_id field. Source spans, raw/cleaned spelling and CHAT output remain
unchanged; JSON reconstruction retains the ID without inventing source spans.
The spec writer traversal applies the direct group-content sink contract to
both parse-clean and recovered models and requires witnesses for each. It uses
only parser-produced groups, including nested groups. Successful serialization
does not certify recovery as valid or make its output an expected repaired file.
The grouped-content reference distinguishes bare 0 from 0 [=! points]
through typed action variants. The nested E342 scope pair deletes only the
inner annotation from its legal control. That deletion yields a whole-tier
ERROR and E316, not a retained inner group with a MISSING slot; its subsumption
claim prevents reconstructing narrower structure merely to emit E342.
The 1082 reference’s cm|cm and punct|‡ morphology items exercise the
public punctuation-counting classification alongside lexical negative controls.
They remain real MOR/GRA chunks: the contract counts every main word,
post-clitic and terminator and compares that total with the authored GRA tier.
This counting classification is not permission to discard punctuation during
alignment. These witnesses do not establish coverage of the helper’s older
beg and end spellings.
The scoped-annotation display contract traverses both canonical populations
through ContentStructure::walk, including nested and replaced annotations.
It requires diagnostic Display to agree with CHAT serialization and propagates
every observed writer refusal through the same irreversible sink state. This
exercises a separate formatting boundary without inventing AST values or
mistaking a recovered model’s printable annotations for valid source.
The reference non-word-display contract exercises all thirteen non-word leaf
adapters exposed by the canonical content traversal: other-speaker spoken events,
long-feature beginnings/endings, nonvocal beginnings/endings/simple markers,
separators, sound events, pauses, actions, overlap points, freecodes and internal
timing bullets. Word and replacement display retain their separate contracts;
underline markers do not offer a standalone Display adapter.
The canonical content traversal supplies parsed values, including nested content,
without alignment-domain filtering. Every adapter requires a reference witness;
its standalone Display must agree with WriteChat, and both must propagate
each observed sink refusal without writing after rejection. The existing
irreversible writer state owns that refusal check. Reference roundtrips remain
the independent spelling contract; adapter agreement alone is not a golden
spelling oracle or evidence that every supported content type was exercised.
The authored word/pos-hint-vocabulary.cha reference covers transcriber $POS
hints. A parsed-reference test pins the public CLAN-to-UD helper’s conservative
mapping, including proper-noun refinement and unknown-tag refusal. These are
API compatibility expectations, not automatic linguistic gold annotations or
evidence that every downstream tagger uses this helper. The direct API tests
retain boundary inputs, such as an empty tag, that do not require CHAT fixtures.
Language-metadata queries use the language-switching reference, E504’s
missing-header spec, and E249’s bilingual control and context mutations.
They pin counts, switching decisions and unresolved-word
counts after the explicit Uncomputed-to-Computed transition, without changing
CHAT text or language declarations. The single-word ambiguous reference
control distinguishes ambiguity itself from switching between separate words.
The E249 variants remove the declaration, remove its secondary member, or
replace the secondary precode with an undeclared language. Ordinary words may
still inherit a precode while bare shortcuts remain explicitly unresolved;
metadata computation must not invent an alternate language or repair headers.
The public validation context and word-language helpers share one ordered-language
classification: a borrowed alternate declaration, missing alternate, tertiary,
or undeclared. Primary/secondary switching therefore has one policy owner rather
than two implementations. Compatibility accessors may return None, but the
shared classification does not conflate why an alternate is unavailable.
Reference-backed builder contracts preserve speaker sets and declaration order
while checking copy-on-write isolation: configuring a clone cannot alter the
empty original, and an explicit mode override cannot alter its parent. Tier
overrides retain shared file metadata and affect only their local context.
Alternate-language and tertiary queries are checked against parsed declaration
order, including the absence of any alternate in an empty context.
Spec diagnostics also exercise fragment-coordinate projection through single,
vector and inline-batch sink delivery. Primary and secondary document spans
must use the same clipped projection; snippet text and snippet-relative spans
remain unchanged. These checks retain the diagnostic’s code, message and other
payload rather than treating a rendered string as the golden authority.
Forward/inverse document-rebasing tests derive cached coordinates from actual
spec sources, require those caches to be cleared on translation, and preserve
the documented dummy-span policy for unlocated diagnostics.
For confirmed non-dummy source ranges with retained context, the suite also
constructs a fresh excerpt and checks snippet-relative span, source line offset,
unchanged found/expected evidence and JSON roundtrip. The public excerpt helper
returns a typed error for invalid source slices, then uses FragmentSource to
admit snippet coordinate capacity. Reference-derived Unicode controls distinguish
valid whole-scalar and zero-width ranges from reversed, out-of-source and
split-scalar ranges. No empty excerpt is fabricated on refusal. These controls
do not allocate oversized source buffers or certify reconstructed contexts.
The JSON word boundary has a separate suffix contract: serialize a parsed
reference word, append a dangling or repeated @ to only its wire spelling,
then deserialize and validate. A dangling marker requires E202; a repeated
marker requires E203 without an additional E202. The unchanged spelling is the
control. These intentionally inconsistent imports carry no parser provenance;
they justify retaining model guards, not claiming that CHAT parsing produces
those states or treating arbitrary constructed ASTs as CHAT goldens.
The legal E203 shortening/form-marker seed adds explicit payload controls:
unchanged suffix, a $ part-of-speech tail, invalid trailing text, a repeated
marker and a missing payload. These assertions concern E202/E203 only; excluding
a part-of-speech tail from the form-marker rule does not certify consistency of
the remaining imported fields.
The validator classifies raw marker suffixes into absent, missing, repeated and
single states. Only the single state proceeds to payload checks, so repeated or
missing markers cannot fall through to a second classification. A single suffix
contains no further @; downstream checks do not repeat that impossible case.
This is validator-input admission, not a new claim of parser provenance.
Raw spelling and typed suffix metadata are not yet one admitted value: public
deserialization and recovery mutation can make them disagree, while CHAT output
uses the typed fields. The existing prefix comparison must not be deleted based
on clean-parser invariants. A mismatching import’s current diagnostic behavior
is not an approved consistency policy; boundary hardening must distinguish
retained recovery spelling from an admitted typed word before reconciling them.
The same parsed lexical seed supplies JSON shortening controls: balanced text,
an unmatched closing parenthesis, and a closing-then-opening sequence. The last
must retain both errors instead of allowing equal totals to hide invalid order.
The validator’s nonnegative nesting depth is bounded by traversed source bytes;
checked closing transitions leave zero depth on refusal without underflow.
An independent lexical-content import places U+0015 inside the seed’s serialized
text element as well as its raw spelling. Decoding must derive cleaned text from
that element, retain absent timing metadata, and report the illegal lexical
bullet. The unchanged seed has no illegal-character diagnostic. This is a wire
boundary control, not a CHAT timing-bullet specimen or parser-reachability claim.
The E525 model-boundary control starts from an imported spec-derived model and
varies optional recovery metadata. Its unsupported-header producer retains a
reason but supplies no correction; validation owns the E525 report. Removing
the imported reason must not invent parser evidence. Explicit caller advice
following the spec’s @Comment policy must survive validation; absent advice
uses the manual reference. These metadata variants are not new CHAT goldens.
The CA-omission import contract starts with E212’s legal normalized omission,
then deletes its content or substitutes/appends typed shortening and compound
components from existing E212/E232 specimens. Both construction policies must
admit the untouched control as Constructed, never parser-backed Clean, and
reject each malformed draft with E212 without silently normalizing it. JSON
replay preserves the edited structure and supplies no parse provenance; failed
construction leaves provenance Unknown. These are editable-model workflow
checks, not claims that the parser can produce the malformed omission shapes.
Retrace import controls instead leave the serialized model unchanged: E370,
E377 and E378 specimens are parsed, encoded and decoded. Located diagnostics
must retain source labels; imported models must report the same violations,
messages and advice without inventing labels at byte zero. Missing source
coordinates are not permission to suppress the semantic violation. These
checks exercise both literal and constructed diagnostic-message paths.
The legal E315 underline specimens also supply incomplete editor buffers:
exact source prefixes end after either marker’s lead byte or its complete pair.
A lone lead at EOF requires one E315 at that byte; the complete pair does not.
The parser may report other errors for the unfinished document. These lexical
boundary assertions neither suppress that recovery nor declare the prefix valid.
CA quotation controls exercise omitted ordinary terminators under explicit
strict-linker selection. The quotation-follows chain remains valid, as does a
quotation-precedes chain with its explicit ending. Removing only that ending
reports E346 at the orphan chain’s opening quoted turn, not at every following
quoted continuation. Default validation keeps this optional rule disabled.
Reference semantic-diff contracts exercise both vector-backed document lines
and small-vector-backed word content. Both storage adapters share one borrowed
slice comparison: shared elements precede a tail difference, and equal lengths
produce no tail. Bounded reports preserve the same difference prefix and restore
the caller’s path and source context; storage choice does not own a second policy.
Reference-backed judgment prompt tests continue the sampling/context workflow
through transport-neutral rendering without a network call. English and Chinese
samples preserve speaker order, utterance numbering and Unicode; absent context
stays explicitly unknown. Authored sidecar ages exercise both sides of the
18-, 36- and 72-month prompt-hint boundaries. These approximate hints and labels
are input controls, not demographic, consent or speaker-identity findings.
The advisory-response workflow uses speaker codes sampled from the reference
model and explicitly authored responses. Even confidence 1.0 returns a pending
human-review proposal with model/endpoint/prompt provenance, never an applied
mapping or invented lexical score. Missing adult roles and a merge claim with
no adult are refused; all-drop advice invents no role. Shared adult-role
proposals retain deterministic numbered codes and labels. These controls do
not certify the advice as true or validate a live external model.
Transform lenient-parsing contracts use existing malformed morphology and
grammar-tier specs plus a primary-speech failure and clean reference control.
Suppressing generated-tier parse diagnostics must retain the recovered model
and its E600 alignment refusal; it cannot certify validity. Primary-speech
diagnostics remain visible, and clean strict/lenient results agree. These
controls include similarly named %morx and %grax E315 specimens:
their diagnostics must remain visible. Suppression is owned by a private typed
Mor/Gra view of the same parse, not by text-prefix matching. Only wholly
contained, located diagnostics may be suppressed; unknown or cross-tier spans
are retained. Continuation ownership comes from the parsed tier’s full span,
not a second line scanner.
E347’s authored overlap controls start with matching indexed speakers and deliberately remove a marker pair or substitute an index. Unindexed orphan controls pin the rule’s exclusion, and a third speaker pins one-to-many matching. The spec claims own diagnostic presence/absence; the corpus boundary test additionally checks exact diagnostic multiplicity and utterance spans. These are reviewed, finite mutations with retained valid controls, not goldens inferred from whatever diagnostics the current implementation emits.
E756’s namespace controls distinguish retained empty unsupported tiers from
intentional %x tiers. They assert the existing E605-over-E756 suppression
policy, Unicode-whitespace emptiness, optional payload presence, tier labels
and diagnostic spans. Unsupported-tier trimming may normalize whitespace;
the contract does not falsely promise byte-exact spelling for that case.
The options-based validation helper is exercised over parsed reference and
error-spec models at every CheckLevel, with and without strict
linkers. Its diagnostics and derived alignment state must match the
explicit rule-selected validator. ParseValidateOptions::validation_policy
admits either no validation phase or an explicit typed policy, shared by the
model helper and transform pipeline. Strict linkers alone remain parse-only;
alignment without validation is not a value the options can hold. An admitted policy is a request, not proof that
the document is valid; the separate accepted-document API owns that proof.
E243’s Unicode boundary specs replace one scalar inside an otherwise unchanged word. They retain ordinary high-BMP and supplementary-plane controls, all 66 noncharacters, every private-use range endpoint and the endpoints of CLAN’s exemption ranges. The diagnostic contract checks retained word text, source spans and one error per rejected word; a deduplicated error-code set alone would miss lost diagnostics. The recorded U+10000 disagreement with CHECK does not override the independent Chatter claim or turn one private-use parity fixture into universal parity.
Semantic-report contracts compare both reference documents and individual parsed reference words. Word-level comparisons exercise nested content without earlier document metadata deciding the first mismatch. They compare semantic equality in both directions, preserve bounded report prefixes and source/path context, and check zero, exact, spare and maximum capacity. The report budget uses explicit Available, AtCapacity and Truncated states: filling the last slot does not prove truncation until another difference is found. The original display limit is derived from stored entries plus remaining capacity. These are reporting-API contracts, not generated CHAT-validity goldens or a new production-corpus differential-testing process.
Participant-map reports additionally compare unchanged conversation, overlap, and speaker-metadata reference transcripts. They pin the first declared speaker key, map length differences, both comparison directions, and report truncation inside keys and values. Container traversal consumes paired borrowed entries from zipped iterators, so an indexed lookup cannot silently abort an otherwise complete comparison. Sequence tail differences remain reported after value differences unless the report has actually been truncated. An authored speaker-removal workflow also parses the transformed reference CHAT and pins the unmatched sequence tail in both comparison directions. Exact and smaller report budgets preserve the full report’s prefix; a filled budget alone must not invent truncation. This is a report of an explicit operator transform, not a speaker-identity inference.
The user-defined-tier reference contains authored labels around a legacy reader
length boundary and a longer descriptive %xcommunicativefunction label.
Validation and roundtrip tests must preserve their full labels and payloads;
acceptance by these tests is not itself a runtime CLAN CHECK observation.
Terminator classification uses canonical spellings of terminators parsed from the reference corpus, comparing recovered typed meaning and requiring absent source provenance for free-string conversion. Parsed content separators supply the refusal cases. This complements file roundtripping without maintaining a second hand-written inventory of terminator spellings in the test.
Both Cargo workspaces set split-debuginfo = "off" for development and test
profiles. Line tables stay in the linked artifacts, so diagnostics and stack
traces retain source locations without macOS’s default unpacked layout
leaving one .rcgu.o file per codegen unit in target/debug/deps.
The setting is based on a measured failure analysis, not a cosmetic
preference. Under the default layout the root workspace had 55,141 entries in
target/debug/deps; the generator-heavy specification workspace had 842,704
entries, occupied 40 GB, and took 29.3 seconds merely to enumerate with
os.scandir. The exact spec
test executable itself started, listed its tests and exited in 0.00 seconds,
while a warm cargo test --manifest-path spec/Cargo.toml --workspace --quiet
took 46.7 seconds. The file layout, rather than the test harness executable,
was the first bottleneck to remove.
To reproduce the diagnosis without running tests:
python3 - <<'PY'
import os
import time
for path in ("target/debug/deps", "spec/target/debug/deps"):
started = time.perf_counter()
entries = sum(1 for _ in os.scandir(path))
elapsed = time.perf_counter() - started
print(path, entries, f"{elapsed:.3f}s")
PY
du -sh target spec/target
A target directory built without this setting holds unpacked artifacts;
remove them once with
cargo clean and cargo clean --manifest-path spec/Cargo.toml. Both commands
delete derived build output only. A warm run should then be measured with
/usr/bin/time -p just test-spec rather than inferred from the per-test times
printed by libtest.
With the setting in effect, a clean spec/target holds no .rcgu.o files and
the deps directory has a few hundred entries rather than hundreds of
thousands. Its three generator commands and six runtime
commands are declared with test = false, because their behavior is already
covered by library and integration tests and their binary sources contain no
tests. This avoids compiling and launching nine empty harnesses.
The project uses plain cargo test. Whole-workspace nextest is not used: its
eager test enumeration launches dozens of new binaries at once and repeatedly
wedges macOS syspolicyd. The cache migration race is closed at its source, so
no test needs process isolation. A future runner change
needs measurements on a clean and a warm target and must demonstrate that it
does not recreate that first-execution burst. Full Disk Access is unrelated to
repository build artifacts, and Developer Tools permission is not a remedy for
an oversized Cargo target directory.
On a small suite (the generators library’s 51 tests, one binary, four workers, warm) nextest is slower than Cargo, so it gives no reason to change the default runner; that says nothing about the full workspace or newly compiled binaries. Reproduce the comparison by alternating:
/usr/bin/time -p cargo test --manifest-path spec/Cargo.toml -p generators --lib --locked
/usr/bin/time -p cargo nextest run --manifest-path spec/Cargo.toml -p generators --lib --locked --test-threads 4 --status-level fail --final-status-level fail
The nextest macOS guide separately describes XProtect startup overhead and Developer Tools permission. That mechanism matters when launching even trivial tests is slow; it does not explain time spent enumerating hundreds of thousands of build artifacts.
Exercise the owned behavior
Property tests must call the production operation whose contract they claim
to verify. A property test that copies a hashing algorithm instead of calling
the real cache-key function can pass with the real cache completely broken, and
the absence of sampled hash collisions is not a correctness property. Cache
coverage lives in the cache_tests integration module, which exercises the
real CachePool with temporary files, including independent paths, parser
identity, alignment mode, overwrites, and clearing.
Regeneration must preserve unchanged outputs
The generators stage command output, publish only changed bytes and prune only
obsolete files in exclusively owned directories. An unchanged just regen
must leave generated Rust, C and fixture modification times alone, so Cargo
does not rebuild merely because a generator ran: a no-op regeneration
preserves bytes and nanosecond modification times of every tracked file and
compiles nothing, and the following just test compiles nothing either. Any
timing taken this way is a warm measurement, not a clean-build timing.
To reproduce the preservation check, snapshot tracked files before and after
just regen without editing or staging files between the snapshots:
python3 - <<'PY'
import hashlib
from pathlib import Path
import subprocess
paths = [Path(p) for p in subprocess.check_output(
["git", "ls-files", "-z"]).decode().split("\0") if p and Path(p).is_file()]
def snapshot():
return {p: (hashlib.sha256(p.read_bytes()).digest(), p.stat().st_mtime_ns)
for p in paths}
before = snapshot()
subprocess.run(["just", "regen"], check=True)
after = snapshot()
changed = [str(p) for p in paths if before[p] != after[p]]
assert not changed, changed
print(f"Preserved contents and modification times of {len(paths)} files")
PY
/usr/bin/time -p just test
One integration binary per crate
Each crate has a SINGLE integration test binary (tests/integration/), so
tests are selected by NAME FILTER, never by target name:
cargo test -p talkbank-parser-tests --tests <filter> # correct
cargo test -p talkbank-parser-tests --test <name> # fails: no such target
--test <name> names a compilation target, and per-file targets do not exist. It does not fall back to filtering: it errors with
available test targets: integration, parser_suite. Every command on this page
was checked by running it.
Test generation pipeline
Specs are the source of truth. Grammar corpus tests, Rust parser tests, the validation fixture corpus and the local error pages are all generated from specs and are never hand-edited.
flowchart LR
subgraph sources["Source of Truth"]
constructs["spec/constructs/"]
errors["spec/errors/"]
templates["spec/tools/templates/\n(Tera wrappers)"]
end
subgraph generators["spec/tools generators\n(run only what changed)"]
gen_ts["just spec-gen: corpus tests"]
gen_rust["just spec-gen: construct test bodies"]
gen_validation["just spec-gen: validation fixtures"]
gen_docs["docs/errors/ (spec-gen artifact)"]
end
subgraph outputs["Generated Outputs (DO NOT EDIT)"]
ts_tests["grammar/test/corpus/generated/"]
rust_tests["parser-tests generated tests"]
val_corpus["validation fixture corpus\n(.cha + manifest.json)"]
error_docs["docs/errors/"]
end
constructs & errors --> gen_ts
templates --> gen_ts
constructs --> gen_rust
errors --> gen_validation
errors --> gen_docs
gen_ts --> ts_tests
gen_rust --> rust_tests
gen_validation --> val_corpus
gen_docs --> error_docs
To add a grammar or error test, add a spec under spec/constructs/ or
spec/errors/ and regenerate. Spec Workflow owns those
commands and writes each one out; they are not repeated here.
Never-regress gates
These guard behaviour a successor cannot easily re-derive. Any commit touching the grammar, parser, model, validation, serialization or alignment runs the matching gates and keeps them green.
A red gate is a bug until proven otherwise, never a test expectation to quietly update. That cuts both ways: a diagnostic that looks BETTER after a change earns the same scrutiny as one that looks worse.
| Gate | Command | What it protects |
|---|---|---|
| Experimental backend comparison | cargo test -p talkbank-parser-re2c --test integration equivalence_reference_corpus | Compares re2c and tree-sitter reference models using SemanticEq. re2c is experimental and incomplete, not a specification oracle. Investigate disagreements against independent specs; canonical-suite success does not prove cross-backend equivalence. Completing re2c is lower priority than CHECK parity and representative spec/reference coverage. |
| Reference corpus parses | cargo test -p talkbank-parser-tests --tests reference_corpus_parses | Every reference file parses cleanly with the tree-sitter parser. Checks parser acceptance, not cross-parser equivalence. |
| Reference transform workflows | cargo test -p talkbank-parser-tests --test integration transform_corpus:: | Public normalization admits a loss-checked Rewrite, preserves typed semantics and is idempotent. Compact and pretty CHAT-to-JSON output deserialize into semantically equivalent typed models. These wire tests do not imply optional model validation or JSON-schema admission. |
| Roundtrip idempotency, and reference coverage | cargo test -p talkbank-parser-tests --tests roundtrip_reference_corpus | parse, serialize, re-parse yields a semantically identical AST (SemanticEq) for EVERY reference file. One test carries both guarantees: it iterates the whole corpus (coverage) and checks semantic equality on each (idempotency). |
| Generated spec tests | cargo test -p talkbank-parser-tests --tests generated_tests | Every construct spec still parses cleanly. (Error specs do not feed this: string-based error tests would be strictly weaker than the fixture corpus plus the observation snapshot.) |
| Validation error corpus | cargo test -p talkbank-parser-tests --tests validation_error_corpus | Every ERROR-spec example (both stages) still satisfies its CLAIM against its generated .cha fixture, absences included. |
| The gate registry | cargo test -p talkbank-parser-tests --tests gates | Runs every gate registered in gate::ALL. Ask the registry what that is rather than a list here: cargo run -p talkbank-parser-tests --bin audit_gate_probes names each gate, runs every probe against it, and prints the rules no probe reaches. |
The transform corpus tests also send canonical error-spec inputs through
required validation under structural and alignment policies. They compare
typed ValidChatFile admission, parse/validation refusal and streamed
diagnostics with the compatibility API. This checks a public boundary contract
under anonymous/default-rule settings; the validation-corpus runner separately
owns each authored claim, transcript name and opt-in rule selection. Acceptance
or rejection under one policy is not a substitute for those claims.
The media-name workflow additionally passes the E531 specs’ authored transcript identity and rule selection through required validation, checking their claims and the identity retained by both admitted and rejected products. Matching, case-only, mismatching and anonymous controls distinguish transcript identity from the generated fixture’s storage filename.
Media-timing workflows reuse E544, E552 and E752 examples to check typed
untimed/linked admission and missing-media refusal. Internal main-tier bullets
and recorded %wor bullets count as timing; a %wor tier without bullets does
not. Reconciliation also checks untimed preservation, serialization semantics
and idempotence using canonical spec inputs.
Timed E535/E536 controls and single-field mutations additionally check that
unsupported media types and statuses produce distinct typed refusals retaining
the authored value, rather than a linked-media capability.
E501’s timed duplicate-header mutation also checks refusal with the exact
declaration count. A unique declaration is retained by exclusive borrow during
reconciliation rather than located again after a separate count.
The fix-catalog contract separately requires an earlier byte-identical,
source-bound header and refuses any conflicting declaration of the same kind,
including a conflict after an identical pair. This proof allows a proposal,
not a header-write capability: existing admission restrictions remain intact.
No line-prefix scan or independently supplied header text establishes identity.
Continued language-header controls require the proposed range to include the
complete continuation and conservatively refuse equivalent values with different
line layouts. Canonical output may normalize that layout; source edits must not.
Mixed-error E501/E258 spec pairs retain the actual duplicate-header/comma
finding alongside E316 from an unmatched closing bracket. Their controls retain
E316 without the duplicate finding. The catalog contract requires a proposal
for the clean duplicate controls and refuses one for the recovered carriers;
diagnosing an error is not proof that its source is safe to rewrite. These
checks use observed diagnostics and the original source-bound parse, not
synthetic diagnostic locations or a second catalog parser.
The same matrix includes E259’s unlicensed comma and distinguishes two other
recovery boundaries. E305 is absent on a tainted main tier, even after deletion
of its terminator: E316 records the incomplete evidence. E244 remains reportable
for a complete stress-bearing word before an unrelated stray bracket. Its local
proposal survives, but edit admission refuses writing into the tainted turn.
The test uses explicit absent/proposed/refused verdicts rather than equating
any document recovery with suppression of every local proposal. These controls
do not prove reachability of a stress diagnostic inside a recovered word itself.
E305’s existing untimed and timed terminator-deletion specs also exercise LF
and CRLF transport. Each user-selected alternative preserves the complete line
ending, passes source-edit admission, and resolves the missing-terminator
finding without structural recovery. Its insertion position comes from the
typed utterance ending for main tiers and terminal newline for MOR, never
subtraction from the diagnostic’s last byte. The reduced postcode reference
pair additionally requires insertion before both final postcodes; inserting
after them leaves E305 and introduces structural recovery. The canonical
catalog contract exercises all three choices on that pair under both transports.
The separate morphology deletion pair proves parser-stage E305 is reachable
without CST recovery. Rejected morphology nevertheless taints the utterance:
the catalog can propose an insertion, but the admission contract refuses each
alternative under LF and CRLF. A clean CST is not proof of valid typed content
or permission to write a proposal.
E259 initial-comma specs exercise one and several separator spaces under LF
and CRLF. Semantic deletion must reproduce the legal spoken control exactly.
The proposal binds a complete comma token and clean tier body; initial position
comes from their structural ranges, and widening consumes the generated
whitespace node. Interior and nested-group cases instead retain their separator.
No preceding-tab heuristic or one-byte whitespace assumption establishes this
distinction. These remain user-reviewed semantic proposals, not automatic fixes.
E244’s later-run pair places an isolated primary stress before a duplicate
primary-stress run. The repair must traverse typed stress tokens rather than
stop at the first matching character, and reproduce the paired control exactly.
A checked primary-token witness licenses deletion only of an immediately
adjacent primary token. Distinct token spans compose across every run, retaining
the first token of each; there is no first-run-only accumulator. This resolves
E244 without suppressing the separate E247 finding about distinct primary
stress positions. Secondary and mixed-stress repair policy is outside this contract.
The stress-run spec deck also covers three primary markers, runs followed by
separated primary/secondary stress, two separate duplicate runs in one word,
and refusal of mixed/secondary-only pairs.
Each word emits one E244: reporting once per adjacent pair would propose the
same edit twice for a triple run, making the real edit batch fail on overlap. The regression submits all actual diagnostic proposals together and
checks exact paired controls or unchanged source, not isolated successful edits.
E258’s three-comma specs exercise two distinct edits in one admission batch,
both at top level and within an annotated group. The changed source must equal
the paired single-comma control, with no diagnostics. E258 and E259 share exact
source-bound comma-token admission and require a clean tier body; a comma-shaped
substring elsewhere is not token evidence. No new raw-text scan is involved.
E750’s group-edge pairs cover complete space runs and nested groups. The catalog
binds the exact whitespace node and the nearest annotated group’s typed content
field, requiring a content edge rather than guessing from neighboring delimiter
bytes. All edge edits compose to the exact clean control through normal repair
admission, preserving interior separators and annotations. Non-space whitespace
remains refused; unrelated structural recovery within the group also refuses
a proposal rather than relying on delimiter-shaped neighboring bytes.
E241’s marker-boundary specs sample shortened and miscased forms across the
three marker categories. Their admitted batch must equal the canonical control
while preserving identical marker-like text in %com. Omission and shortening
notation remain diagnosed but refuse whole-word replacement; sound-material
controls do not acquire marker-spelling semantics. The catalog binds an exact,
clean standalone word before asking the existing model spelling classifier,
rather than interpreting an independently supplied diagnostic substring.
Reference language-retagging tests establish a collision-free temporary code,
then rename each declared language and reverse the operation. They check model
semantics and per-notation counts, and require unsupported-span refusal to leave
the model unchanged. Both success and refusal must be witnessed by the corpus.
Identity retagging must preserve the language-list owner’s declarations and
return unchanged transform statistics, even on files with span notation.
This inverse-property test complements, rather than replaces, authored examples
of exact forward retagging output.
The authored word-features/retag-source.cha and retag-expected.cha pair
checks exact forward serialization, declaration deduplication, utterance scopes,
nested and replacement word markers, and unchanged ordinary words. Both files
must pass required validation. Direct transformed-model equality is inappropriate
here because words retain original raw_text provenance; semantic comparison
is made after reparsing the serialized result at the wire boundary.
The async_corpus workflow enables the model’s async feature in the test
dependency graph. Cleanly parsed canonical spec documents are admitted through
both synchronous and async default-policy validation with their authored names.
Proof/refusal models, policies and diagnostics must agree, even when the async
streaming sink discards diagnostics. Both acceptance and refusal must occur.
This is a transport/admission contract, not an oracle for optional-rule claims;
the canonical validation-corpus runner continues to own those claims.
The same workflow sends each manifest’s rule selection through
validate_with_rules_async and the real unbounded channel sink, requiring its
complete diagnostic stream to match synchronous rule-selected validation.
Successful task completion is not confused with document validity.
An external-consumer failure control uses the canonical duplicate-language
header (E501 example 3) and a test-only diagnostic sink that panics. Both async
APIs must return AsyncValidationError::Join carrying that panic, not successful
completion, a document-validity refusal, or a ValidChatFile proof. This tests
the task boundary; it introduces neither a new CHAT rule nor a production panic.
The legacy output-check contract uses canonical E243 and E362 control/error
pairs to demonstrate a different boundary: validate_output can succeed on
invalid CHAT because it checks only selected command invariants. A standalone
slash and backwards turn ordering are not covered by those limited checks.
Consuming full validation must refuse those documents with their authored
diagnostics, while issuing ValidChatFile proofs for the corresponding valid
controls. Both proof and refusal retain the parsed document unchanged. An
Ok(()) from a compatibility check must never substitute for that proof.
Dependent-tier regeneration uses parsed reference entries for every supported
family (%mor, %gra, %wor, user-defined). Identity replacement must preserve
order and separator provenance; remove-and-append must retain other entries and
create a clean separator. E758’s non-CA %mor fixture separately proves payload
replacement retains the original spacing evidence and resulting diagnostic.
A separate contract uses semantically distinct parsed reference donors with a
matching typed regeneration key (including the user-defined label). The exact
expected tier sequence changes only that payload, and every family must have
a distinct donor. This tests the regeneration helper’s boundary, not whole-file
alignment validity after combining tiers from different reference utterances.
Whole-utterance language-switch rewrites compare E255’s per-word violations
against independently authored legal precode controls, including exact output,
change counts and post-rewrite admission. The reference corpus also checks that
one rewrite reaches a stable serialized form. This exercises the shared typed
switch decision; it is not authorization to run parser-based corpus cleaning.
The E255 deck also includes grouped, replacement and filler words in one
authored before/after pair. Its governing-span control must yield no
UnspannedSwitchTarget and remain unchanged, pinning the capability’s safety
boundary rather than merely checking the absence of E255.
E504’s mixed-language control and missing-language-header mutation both remain
unchanged by this rewrite. The transform may extend an existing declaration;
it must not fabricate a missing header or report an append it did not perform.
The fragment corpus also exercises public age-token parsing from E517’s typed
@ID fields. Lexical token acceptance is deliberately distinct from complete
CHAT date-pattern validity: the public validator must still emit E517 for the
authored violations, including values the token API accepts.
The age deck includes legal full, omitted-day and year-only forms plus missing
or non-digit component mutations. In particular, legal year-only CHAT is not
misclassified as an E517 violation merely because this token API rejects it.
Reference-derived fragment boundaries cover the inverse of successful header
parsing: two adjacent headers cannot become one header, and a speech tier cannot
be admitted as a header. Their rejection diagnostics must remain within caller
bytes and rebase with the requested offset. The ID-only adapter is a typed
selection: a valid non-ID header returns no ID value without inventing a syntax
diagnostic. These cases use parsed source spans, not fabricated CST nodes.
The strict Result APIs independently refuse adjacent headers as one header,
headers as main tiers or words, and whole main tiers as words. Each refusal must
retain nonempty diagnostic evidence; streaming-fragment checks alone do not
establish this contract for the strict entry points.
The validation corpus exercises header-only and alignment-only public entry points over reference/spec models, including actual parser recovery. Header-only diagnostics form the initial phase of full validation; anonymous trait validation preserves the complete stream. JSON serialization/deserialization preserves semantic content but deliberately loses parser provenance. Alignment-only validation must warn for each relevant pair in that unknown state, rather than certify alignment or report speculative count errors. This is a real wire-state transition, not a test that manually assigns a parse-health flag.
The checked-construction contract rebuilds canonical main/dependent tiers and preceding headers through the public utterance constructor, retaining authored separators. Structure-only and alignment-inclusive policies must agree with their parser-backed counterparts on acceptance and JSON output. Each policy requires positive witnesses for acceptance, invalidity and recorded recovery refusal. Constructed models cannot authorize source-byte edits; single-tier or dependent-wide taint withdraws construction admission. Appending a dependent tier also clears admission and derived alignment metadata. These are supported model-construction workflows over canonical content, not fabricated parser faults.
Reference decoration-editing contracts relocate parsed linkers and postcodes through their public mutable iterators, then restore them. Both families require real source-span witnesses, exact coordinate shifts, unchanged CHAT/JSON and semantic content, and exact restoration. The postcode decoder also checks its span against the actual CST token boundary; successful decoding must not discard that provenance. Its input is source-bound; an independent equal-text parse cannot admit the node, and source-field failures remain internal failures.
Coordinated morphological/grammatical replacements use parsed reference tiers.
Reversed, empty and out-of-range replacement requests must refuse without
mutating either tier. Admitted lexical-block replacements preserve donor heads,
and outside dependents follow the caller’s HostRedirects: a block replaced by
itself item by item leaves every outside head unchanged, and a donor block with
no item correspondence sends them where the stated per-item targets say. They
are not assumed to be whole-tier identity operations. L2 fixtures written as
CHAT and parsed (host_redirects_corpus.rs: a host utterance and one donor
utterance per span) pin the corrected heads for two @s sentences, one
spliced span at a time, a dependent placed through a unique head chunk with a
span root anchored after a growing range, and the refusals (unequal item
counts by item, a wrong target count, an out-of-block target or a missing
counterpart per item, an ambiguous head chunk only when a host relation
depends on it, and a span root inside the replaced range, past the host,
under a host chunk that depends on the span, or at the utterance’s root
beside the host’s own), each leaving both tiers unchanged. SplicedBlock’s
unit tests refuse every block that is not a one-rooted tree, and
AttachmentRelation’s refuse every spelling of a root label. The model’s
private admitted host range exclusively borrows both tiers before mutation, and
single-item replacement shares the same admission and rewrite path.
Two admitted reference donor blocks with different chunk counts exercise both
growth and shrinkage, preserving index validity and the declared head mapping.
Single-item cases cover admission of a block holding the utterance’s root,
and the refusal of unrebased donor heads and wrong relation counts where the
block is built. Donor admission
retains the actual parsed tier pair with its checked lexical-block extent.
Word-timing sequence tests follow the full capability chain: count binding,
lexical corroboration, then complete positive timing assessment. Same-count
lexical refusals must expose every mismatching slot in order, with its projected
main-tier text and recorded %wor text under the selected membership policy.
The canonical workflow checks that complete explanation against the two typed
tiers; count agreement alone never admits timing. E544’s timing
deck supplies gaps, touching intervals, overlaps, backward starts, a partially
timed/non-positive sequence and an empty projection. All sequence states and
adjacency classes require corpus witnesses. Hulls must equal the minimum onset
and maximum offset of their admitted slots; every refusal retains all missing
or non-positive timing issues. Neither linkage evidence nor a complete timing
hull certifies acoustic accuracy, and lexical ownership stays on the main tier.
Sanitizer corpus tests run every reference document through deterministic redaction, fresh parsing and a second redaction for byte-idempotence. Speaker codes, main-tier timing and grammatical relations must survive. These wire contracts guard against delimiter collisions in inline placeholder output; they do not certify complete privacy coverage or authorize disclosure of sanitized data. The same population witnesses all nine additional free-text header payloads: each must become the redaction marker without changing header kind or its speaker reference, while unrelated preserved headers remain semantically equal.
Builder corpus tests project reference headers and main tiers into the public
transcript-description schema, preserving all representable ID demographics.
They compare the built main tiers and participant join with serialized output,
and require missing-language refusal. A description that emits @Options: CA must
not interpret parentheticals as shortenings, so the builder carries that CA
parsing context. Its admitted context owns nonempty languages and contextual
fragment parsing. Header and utterance construction obtain their inputs from
the description borrowed by that capability, rather than accepting a second
description. Participant names and first-language headers are also preserved
where represented by the input schema. These are explicit projections, not a claim that the builder
can reconstruct every source header or dependent tier, or that a returned
mutable model carries full validation evidence.
Recorded reference bullets also supply the builder’s separate start/end fields.
The complete pair must preserve main-tier semantics; deliberately removing
either endpoint or removing the timed text must refuse, not silently erase
timing. A private admitted text input separates empty, untimed and completely
timed cases before rendering. These mutations concern the description API,
not newly authored CHAT invalidity claims.
Reference media declarations also exercise exact type preservation and the
documented absent-type audio default. Replacing the type with the authored
unsupported E535 value must refuse. Media is admitted once into the bound
construction context, making header rendering infallible instead of allowing
it to silently change an unknown declared type to audio.
Morphological co-construction tests admit parsed reference %mor/%gra pairs
through the canonical alignment owner, then reconstruct their actual items,
relations and paired terminator. Semantic identity and the constructor’s shared
span policy are checked separately. Clitic witnesses distinguish chunks from
items. Deliberately omitting a relation or double-including the terminal relation
must produce the documented typed count mismatch. These are API error variants
derived from corpus data, not newly adjudicated golden CHAT or a certificate of
dependency-tree validity. Direct constructor boundary tests remain necessary.
E720 separately supplies authored CHAT claims: a reference-derived clitic
control and mutations removing or appending one terminal relation. These run
through ordinary parsing and validation, checking both count-mismatch directions
without changing the control’s morphology. Their observations are reviewed
separately from the claims.
Mismatch rendering consumes the typed MorChunk variants directly; there is
no second kind classifier with an uncalled main-chunk fallback.
The clitic reference also pins authored projection results: it~be a cookie
has five chunks including punctuation, but only three word items. Both clitic
chunks retain the same borrowed host; item starts and dependency heads use
their distinct typed index spaces, including ROOT. Out-of-range item/chunk
requests refuse without inventing a host, and error-spec tiers witness missing
relation slots. These projections do not certify an invalid dependency graph.
Content-only %mor serialization is checked against full-tier framing and
the same one-way refusing sink, with real empty and post-clitic tier witnesses.
Recovered-tier output remains boundary evidence, not a valid-CHAT golden.
Replacement serialization has one payload owner: ReplacedWord emits its
original word and separator, delegates [: ...] to its typed Replacement,
then emits trailing scoped annotations. Canonical reference and error-spec
replacements exercise both standalone payload output and enclosing Display,
including multiword and annotated cases. The one-way refusing sink checks every
observed write boundary; successful output must agree with CHAT serialization.
Recovered annotation output is not treated as proof of source validity.
E711’s post-clitic feature matrix retains flat and keyed valid controls, then deletes the main-word value, the clitic value, and both. All five remain syntactically parseable. The corpus contract checks zero/one/two diagnostics, retained feature keys and host structure, public feature-constructor roundtrip, JSON preservation and unchanged CHAT spelling. CHECK accepted these variants in the recorded observation; its silence does not weaken Chatter’s existing nonempty-feature rule. Claims and code-set observations remain separate from the diagnostic-multiplicity assertion.
E220’s bare-numeral matrix distinguishes digit-bearing words from numerals: tone and homonym digits keep their resolved-language exemption, but a mixed language candidate cannot license a word consisting only of ASCII digits. Omission and unresolved-language policies remain separate. CHECK observations and the manual’s number-spelling rule support this boundary. The written Mandarin reference supplies authored spelling controls for the numeral spec; the generation API must match them without rewriting either source document. Its skipped-group case guards against duplicate zero emission: one group-prefix state carries first/adjacent/skipped context to the sole zero-emission path. This is a generation contract, not automatic repair or pronunciation inference.
The Spanish number-spelling reference checks standalone cardinals around the
hundred boundary: cien for 100, ciento when further cardinal content follows,
and the corresponding form inside 1101. The expected forms are authored from
the RAE numeral table,
not copied from transform output. Controls also cover thousands, including
100000 -> cien mil, which must not multiply the complete phrase diez mil.
The Spanish composer admits standalone groups and thousands below one million,
selecting hundred forms from numeric structure. Thousands multipliers requiring
apocopation are refused (for example 21, 31 and 101); one thousand is mil.
Larger noun-based scales are unsupported. Refused input is preserved, including
currency tokens and ranges containing an unsupported numeral; it is not thereby
certified as valid CHAT. These controls do not certify feminine agreement,
prenominal apocope, other languages’ decomposition or non-English currency names.
Other table-backed languages use exact entries only, never generic arithmetic composition of complete phrases. French and German reference controls retain their authored spellings. Unsupported numerals preserve the whole input token, including currency, digit-leading compounds and number groups; a supported part cannot license a partial rewrite. Existing language-specific composers remain separate. Preservation is not certification of CHAT validity.
The English number-spelling reference extends that contract to irregular and
compound ordinals, short-scale cardinals, decade shorthand and century decades.
Authored written controls pair with the E220 numeric-form specimen.
Thousands with remainders retain the ordinal conjunction convention but omit
prose commas: generation emits spoken words, not a formatted prose number.
The ordinal composer admits only 0-9999 through a private checked type.
Unsupported suffix-bearing inputs are preserved exactly, not given a guessed
th suffix. E220 specs exercise this refusal independently of CHAT validity;
leading-zero preservation is a raw-string API test because CHAT gives an
initial zero its own omission semantics.
Decade composition likewise consumes a private admitted shorthand/full-year
sum, so unsupported magnitudes or nonmultiples of ten cannot reach inflection.
Supported endpoint controls and deliberate unsupported suffix variants enforce
the public number-generation policy.
Year-form lexical tests remain separate: a valid year such as 2007 need not
be admitted as a decade. No guessed decade phrase becomes a golden control.
The contract admits the written document through validation, requires E220 for
every numeric token, compares generation output with the authored word sequence,
and verifies that neither source was edited. These selected pronunciations are
explicit policy examples, not an inference about an unseen recording.
Digit-leading compounds preserve their alphabetic tails after expansion;
all-numeric dash sequences expand each group separately. Already-written
reference words, including alphabetic hyphen compounds, must remain unchanged.
The cardinal cases include multiplied scales, skipped groups and the u64
maximum. They guard against generic concatenation of complete phrases (2000 must not
become “two one thousand”). English admits a nonzero decimal scale and its unit
from the existing lexical table before composing the multiplier. Other-language
generic decomposition remains a separate policy-review target; this English
contract does not certify its linguistic correctness.
Generic table decomposition admits numeric keys as NonZeroU64 before
iteration, so a zero divisor is not an iterable entry. Arithmetic admission
controls retain zero-value lookup and refusal for tables containing only zero,
malformed or overflowing keys; they are separate internal-boundary evidence,
not corpus coverage or certification of multilingual pronunciation.
The splice corpus contract proposes source-span identity edits in reverse order, then verifies sorted mappings, distinct edit provenance and exact byte identity. Canonical main-tier serialization supplies non-identity replacements whose mapped regions and reparsed semantics must agree; original line framing stays intact. Empty edit sets, unrecorded tails/truncation, duplicate targets and mid-UTF-8 insertion points exercise protocol boundaries. None of these checks certifies arbitrary semantic edits, and the tests never write corpus files. Source-derived cross-utterance replacements must refuse even when their start is clean; replacements spanning tiers of the same utterance remain admissible. An internal utterance-scoped edit binds the insertion point or both replacement endpoints before health admission. This does not certify arbitrary mutable model spans or bind the external source string to that model. The catalog contract additionally observes real parse/validation diagnostics from canonical error-spec inputs, keeping each recovery model and diagnostic set with its exact source. Only deterministic mechanical proposals enter edit admission; a partially refused proposal is not applied piecemeal. Admitted proposals must preserve bytes outside their edits and reduce the triggering diagnostic count after reparse/validation. Semantic and ambiguous proposals remain review-only. These checks do not establish whole-file validity, prove the semantic correctness of every proposed repair, or replace authored spec claims and their rule-selection-aware runner.
The timed-gem reference additionally supplies complementary speaker projections with speech strictly before and after a named, fully timed gem. Source-bound exterior admission must reconstruct the original speech/gem order, preserve payload and timing, and return both placement receipts. An equal-content clone cannot substitute for the bound reference. Existing untimed gems must refuse the timed-exterior capability. These are structural timing contracts, not evidence of acoustic accuracy or permission to omit speech. The E526-E530 authored controls and mutations also enter this boundary: unpaired, mismatched, duplicate and lazy markers cannot issue timed placement evidence. Legal untimed, nested or unlabelled examples still cannot provide that capability; refusing the operation does not change their CHAT validity claims.
Structural merge corpus tests bind unchanged reference documents to their
original donor coordinates and refuse extra, out-of-bounds or reversed parent
mappings. Gem boundary specs also exercise repeated ends after a scope has
closed, with and without a different scope still open, and unlabelled versions
of unmatched/nested begins. Bare @G has a legal reference control; inserting
it inside an explicit scope violates E530. A colon without its required label
is a separate malformed-input claim, not the legal bare form.
Sequential reuse of a closed label and multiple unclosed begins exercise the
active-scope transition. Validation retains a nonempty set of actual begin
locations per active label; consuming the last begin removes the entry. Closed
labels cannot survive as zero-count pseudo-scopes, and begin multiplicity is
derived from retained locations rather than a separately mutable counter.
Fully validated, nonempty reference documents without section-placement
ambiguity exercise retain-all assembly, total reference/donor fates, unchanged
speech and dependent tiers, mandatory reporting, and output wire semantics.
Empty retain sets and overlapping unretained speakers must refuse. This does
not establish cross-source acoustic correspondence or authority to omit speech;
section placement and distinct-donor integration remain separate contracts.
The same documents supply typed header-only projections for donor insertion,
with both no stripping and the default donor-tier stripping policy. Every
inserted origin and stripping receipt must agree with the source; main tiers
and unstripped dependent tiers retain their semantics and order. An empty donor
speech selection still retains its participant declarations: collisions with
live non-retained reference speakers must refuse, even when roles match.
Complementary speaker projections reconstruct each eligible multi-speaker
reference’s original utterance sequence. Cross-source adjacency constraints
come from that same original sequence, not fabricated timing. Exact origins,
speech, dependent tiers and recorded bullets survive; contradictory order and
an equal-content but differently owned reference cannot obtain admission.
Rediarization corpus tests use absent timelines and single-track timelines derived from the reference documents’ recorded bullets, with both existing and new anonymous track labels. They check exact attribution/flag accounting, preservation of everything except speaker attribution, reconciled participant and ID sets, and full model wire equivalence. A returned model must not carry an empty participant map despite populated headers, so rediarization travels the canonical participant-join reporting transition before returning the model; the tests do not certify acoustic truth or invent source timestamps. The first corpus-derived contested row also tests every truncated JSON output capacity using standard bounded byte buffers. Each refusal must propagate an I/O error and preserve exactly the accepted output prefix; exact-capacity output must succeed. This is an output-boundary contract, not an extra CHAT construct. Contested ownership tests duplicate source-backed turns for one track and retain simultaneous turns for another. Same-track duplication cannot inflate held time; cross-track overlap remains in both shares. Threshold and turn-order changes affect neither attribution nor payload. Header-only references must preserve declarations without inventing an empty participant list. E524’s timed legal control and one-field birth-reference mutation exercise the content wrapper’s admission boundary. Identity attribution retains the valid birth reference; replacing its sole speaker preserves the birth header, reports the resulting orphan through the required join sink, and refuses serialized output. Invalid input is refused before attribution rather than silently repaired.
Alignment metadata tests recompute derived state over both parser-produced
models (including real recovery) and their JSON-decoded counterparts. All eight
structural alignment families require unknown-provenance witnesses with no
trusted pairs and one warning each; %wor cannot gain a timing binding from
unknown provenance. Recalculation preserves content and parse health, produces
stable metadata, and replaces rather than accumulates diagnostics.
The same dependency-policy table drives a recovery transition matrix over
clean parser-produced specimens. Each tier is tainted through the public
provenance API before recomputation: affected alignments must withdraw their
pairs and replace cached diagnostics with one recovery warning, while unrelated
alignments remain identical. Repeating computation must be stable. Every
structural family requires a nonzero withdrawal witness; %wor timing bindings
must likewise disappear when their participating tiers lose trust. These are
trust-transition contracts, not additional Phon syntax or validity rules.
The wire contract also includes already-computed alignment metadata: decoded
cached pairs do not restore parse provenance, and recomputation replaces them
with warnings. Diagnostic contexts that serialize an empty expectation list by
omission must decode without requiring that field.
The legacy wire payload remains inspectable; its presence is not validation
evidence, and consumers must use the provenance-aware computation boundary.
Semantic-report corpus tests compare adjacent reference models and parsed models with their JSON-decoded equivalents. All five difference kinds require witnesses. Bounded reports retain exact prefixes of uncapped reports, including source locations, and traversal restores its caller’s path and span context. A zero-capacity report may be empty but truncated; emptiness alone is not an equality verdict. Wire-only provenance loss must not appear as semantic edits. The same pass compares original top-level content items within each observed enum variant, preventing an earlier header mismatch from short-circuiting every payload equality check. Word, pause, replacement and annotated-group differences require nonzero reference witnesses; bounded payload reports preserve the full report’s prefix and restore source/path context. This is not exhaustive coverage of every derived field, nested variant, or procedural-macro expansion. Error-spec models with parser diagnostics use the same payload/report contract. They must compare equal to themselves even with a zero report budget, and same-variant pairs require actual unequal recovered payloads. This inspection does not validate, reparse, or require wire admission of recovered models; recovery findings remain separate from semantic-comparison results.
Reference JSON roundtrips cover both the schema-skipping and default schema-checked pipelines, in compact and pretty forms, with identical output. The underline marker wire type is shared by decoding and schema generation; in-memory source metadata must not make the schema reject serialized markers. E531’s authored name controls and mutations also run through both schema policies: skipping JSON Schema never skips requested CHAT validation. E356/E357 controls and deliberate word-internal/grouped marker mutations also run through JSON roundtrips. Parsed markers retain optional source locations; decoded markers explicitly have none, rather than a fabricated zero span. Semantic equality ignores that provenance transition, while validation must still satisfy each authored underline claim using the available enclosing span. This includes standalone markers inside groups and markers inside replacement text. The marker inspection exhaustively handles content variants and visits both original and replacement words; ignoring the replacement wrapper would leave the wire-location contract untested for its editorial text. The balanced controls keep standalone markers away from angle-bracket edges: CHECK strips the controls before its edge-spacing check. CHECK accepts the isolated opening/closing-marker deletions too; Chatter retains its paired-marker rule rather than treating that silence as a validity guarantee.
The diagnostic corpus runs spec-derived parser/validation errors through source-location contracts for nested commas as well: E259 controls and event/omission variants require the exact comma span, one diagnostic for a violation, and unchanged CHAT. A later word in the group cannot license an earlier comma; editorial replacement text cannot turn an omission into speech. The validator consumes the shared in-order content traversal, with explicit awaiting-content/licensed states instead of a separate recursive look-ahead.
The diagnostic corpus also runs spec-derived parser/validation errors through source-indexed rendering. It preserves diagnostic codes, severity, messages and help while checking byte-based line/column coordinates independently and requiring primary/secondary highlights to be valid slices of display text. Empty contexts retain zero-width positions, not fabricated one-byte spans. The same raw diagnostics run through shared plain and ANSI rendering, requiring the same enhanced evidence in each result and leaving raw evidence untouched. Display-relative diagnostics are not fed back into the one-shot source enhancement API; repeated enhancement is not an idempotence contract. Standalone rendering consumes the enhanced error’s embedded context. The shared-source helper is exercised with both raw and enhanced spec diagnostics: sharing the source buffer must not change rendering, and embedded display context must take precedence over the full-file fallback. Each rendered form retains its diagnostic code. These are presentation-boundary contracts, not proof that raw and enhanced diagnostics are distinct types; that API distinction remains an open hardening task.
Selected diagnostic contracts also check more than code presence. E532’s canonical-role control, paired misspellings and lowercase substitutions require two diagnostics per invalid role (one per header occurrence), exact corrective advice, and unchanged serialized spelling. This distinguishes role admission from suggestion text: a heuristic hint must not become an accepted alias or a silent model rewrite. The invalid-case variant owns its expected advice, so a canonical control cannot accidentally inherit another case’s correction table. The suggestion matcher uses only the shortest equivalent substring predicates; longer substrings already implied by them do not need separate runtime branches.
E518/E545 date contracts distinguish fixed-width ASCII digits from general
integer syntax. Paired recording/birth controls and signed or nonnumeric
components require the header-specific diagnostic and unchanged serialization.
The model-owned private digit-admission type is shared by date construction,
JSON decoding, and header validation. It proves width and alphabet before day
or year is interpreted numerically; the corpus still owns the 01-31 policy and
wire-format behavior. These tests do not certify full calendar validation.
The same canonical fixtures require invalid spellings to remain Unsupported
through parsing and JSON roundtrip, not merely receive a validation diagnostic.
Timed-pause boundary specs distinguish admitted CHAT spelling from a bounded
numeric projection. Minutes-to-seconds multiplication and addition are checked
before entering PauseTimedDuration::Parsed; overflow retains the spelling in
Unsupported, with no fabricated wrapped duration. Canonical boundary cases
exercise parsing, validation, JSON decoding, numeric projection and unchanged
CHAT serialization. These are robustness witnesses, not production-frequency
claims. Submillisecond text remains intact even when its numeric projection
truncates to milliseconds under the existing policy.
The Parsed variant owns a ParsedPauseDuration whose fields are private.
Use PauseTimedDuration::new for admission and the payload’s seconds(),
millis() and as_str() accessors for inspection; callers cannot pair a
spelling with independently supplied numeric components.
The timed-pause reference also tests the external JSON boundary: seconds
contains the authored string, not a numeric projection. Removing that field or
replacing it with a JSON number must fail admission; neither operation may
invent a spelling or silently introduce a default duration.
The separate public-API boundary deck checks integer duration strings (including
the maximum supported seconds), numeric conversion overflow and malformed
fractions, including Unicode. Constructor and JSON admission must agree on
Parsed versus Unsupported, preserve the exact string, and never fabricate
a numeric duration. The model’s integer-string API is broader than CHAT’s
decimal-point pause grammar; these cases are not canonical CHAT coverage.
The validation-runner corpus sends canonical reference and spec paths through the worker pool with caching disabled. Per-file event states require exact diagnostics before completion, reject duplicate or missing events, and reconcile the terminal statistics against the complete input. It tests actual storage filename identity, not the manifest-authored names used for spec claims. Recursive reference discovery additionally requests roundtrips: admitted files must report successful serialization roundtrips before completion, and terminal roundtrip counts must agree with those events. Invalid files skip that phase. The cache workflow uses an isolated in-memory SQLite cache with the runner’s actual parser/rule identity. Only a cold reference-file run populates it; the warm read-only run must preserve verdicts, diagnostics and roundtrip totals while reporting the learned hits. No cache verdicts are preloaded as answers. Validation-only entries first demonstrate that no roundtrip result exists; requesting roundtrips then performs and stores the missing work, followed by a warm roundtrip reuse pass. Validation success alone cannot certify roundtrips.
The file-pipeline corpus checks disk parsing against the typed reference models and compares spec-file admission/refusal evidence with the named in-memory entry point, including strict-linker and alignment policies. Both sides use the actual storage filename; authored manifest names remain the responsibility of the separate spec-claim runner. A directory must produce an I/O refusal, never an empty CHAT model.
Strict quotation/completion validators receive a UtterancePosition issued
by the complete FileUtterances view. Its private current/before/after fields
bind one real utterance to its own neighbourhood; callers cannot supply an
out-of-range or cross-file index. First-position absence of a predecessor is
still a real diagnostic case, not an impossible-index fallback. Canonical
strict-linker specs retain policy and diagnostic coverage across this boundary.
External overlap-analysis indices still use the separate fallible lookup.
Reserved bullet-rule specs are not evidence that stricter timing policy is implemented or required. Their real bullet-bearing controls also run through default validation: cross-speaker overlap, gaps, untimed turns, and exact-500-ms self-overlap retain the adopted policy. CHECK’s optional continuity flags are observed separately. A paired E316 delimiter-deletion case distinguishes real media bullets from bare timestamp text after an utterance terminator.
E744’s phone-interval matrix covers both sides of the 1 ms media-boundary
tolerance, absent media timing, and the unsigned timestamp limit. Maximum
integer cases are robustness boundaries, not claims about real recording
durations. A borrowed PhoneExtent couples the first and latest observed
intervals; its presence certifies observation only, never valid ordering.
Bounds use saturating differences rather than overflowing tolerance addition.
The canonical contract keeps E742 interval-order evidence independent of E744
and proves validation preserves even invalid timestamps byte-for-byte.
Grouped E714/E718 controls and single-target deletions exercise the diagnostic
side of alignment: phonological/sign groups stay atomic in their own domains,
pauses appear in phonological positions, and action markers and their annotations
appear in sign positions. The contract pins displayed positions, descriptions,
the missing-target marker, and unchanged CHAT. It does not introduce another
counter or restate count/extraction equality: PositionalDomain, AtomicUnit
and the shared traversal remain the single policy owners. These authored pairs
specialize existing reference-corpus shapes rather than inventing model trees.
The E704 untranscribed-timing matrix pins the adopted CHECK133 rule: a timed
xxx, yyy or www turn constrains the same speaker’s following speech just
as a lexical turn does. Transcription availability is not timing eligibility.
The canonical contract checks that lexical classification differs while all
four cases report exactly one E704 at the following bullet, preserve their
source, and accept the middle turn’s exact-500-ms overlap. Retained real CHECK
observations agree with these four rejection claims.
The public collecting cross-utterance API also runs the authored E341 quotation and E347 indexed-overlap controls and mutations with strict linkers off and on. Disabling strict quotation checks must not suppress indexed overlap errors; changing an overlap index must retain both orphan diagnostics. This is a phase-specific policy contract, not full-file validity admission or a new claim about CHECK’s quotation policy.
The E704 marker pair also distinguishes an ordinary same-speaker continuation from one carrying a bottom overlap pair. A later other-speaker response does not turn the ordinary continuation into self-overlap, nor excuse an adjacent same-speaker top/bottom pair. Both default and strict-linker configurations exercise these authored controls; no new CHECK observation is implied.
E220’s mixed and ambiguous language cases pair a digit-permitting candidate
with a candidate substitution that removes that permission. Both + and &
forms exercise the resolved candidate set: any permitting language suffices;
otherwise the diagnostic must retain the corresponding language interpretation.
These are parsed spec fixtures, not hand-constructed language resolutions.
The same E220 specimens and E504 header controls exercise downstream candidate
ordering, display, universal selection, and serialized language-resolution
identity. Universal selection is distinct from E220’s permissive policy:
every candidate must qualify. The explicit Unresolved variant survives the
wire roundtrip; a vacuously true query over its empty candidate set does not
establish a language or authorize a language-specific operation.
Scoped E220 controls also run through NLP extraction in Mor, Pho, and Sin domains. The enclosing language governs unmarked words and the Mor comma; the word’s own marker still wins. Extracted words resolve through their opaque governing mark, preserving the source position captured during traversal. CHECK rejects the Chinese-span digit control, so that observation is retained as a scoped discrepancy rather than reported as parity.
Existing replacement and retrace reference files also have authored extraction
sequences for each domain. Mor selects replacement text and excludes retraced
or [e]-marked material; Pho/Sin retain the eligible spoken originals. Expected
sequences are not computed with the same selection helper as the implementation.
E370’s CA repetition control and following-speech deletion exercise retracing without a terminator. CA’s terminator exemption does not waive the requirement for substantive speech after a repetition marker. The pair tests those rules independently, through the canonical spec runner rather than a fabricated tier.
Replacement-category spec pairs exercise omitted, untranscribed, fragment, nonword, and filler filtering. Invalid examples retain their validation diagnostics even when the extraction API can produce a sequence; extraction does not certify validity or repair the source. Ordinary annotated-word exclusion belongs to the shared scoped walker, while replacements carry their own annotations into the replacement-specific selection path.
E769 pairs a comma control with a semicolon substitution, including nested and retraced variants. The semicolon remains a typed separator for lossless legacy parsing, but modern CHAT validation rejects it at its own span. The extraction contract also verifies that this non-tag punctuation is not an NLP word; that does not make the input valid. Current CHECK evidence supports the rejection.
E243 ellipsis specs retain valid trailing-off and nested/replacement controls, then insert U+2026 into word text. The diagnostic contract checks each rejected word’s full source span, exact multiplicity, and byte-preserving serialization. CHECK accepts both controls and rejects both mutations; this grounds another specific CHECK 48 shape, not every branch of that broad diagnostic.
Compound-part specs delete lexical material while retaining stress markers before, between, or after compound joins. A typed progress state tracks whether the current part has spoken material and whether a join has been crossed; each join resets that evidence. E232/E233 cannot be bypassed by placing prosody in an otherwise empty part. The corpus contract verifies exact codes, source spans, and unchanged serialization. CHECK catches the leading case but accepts the empty middle/final cases; its silence does not establish validity.
Prosodic measurement exhaustively matches the closed stress-marker enum. Primary and secondary stress are the only representable variants; there is no third, uncounted fallback. Existing E244/E247/E250 controls and mutations remain policy tests, while extending the marker enum requires handling it at compile time rather than relying on another boolean-predicate test.
Roundtrip text-report tests use canonical transcript lines to check the
five-difference limit, actual truncation, missing trailing lines, and identical
text. Missing lines remain optional values until rendering; present lines are
quoted, so literal <missing> text cannot masquerade as absence. These are
report-format boundaries, not a claim that valid reference CHAT fails roundtrip.
File and test counts deliberately appear nowhere on this page. They change
weekly; ask the tree (rg --files -g '*.cha' corpus/reference | wc -l) rather
than trusting a number in prose.
The gate registry
A repository-wide gate computes findings and must FAIL when there are any.
Written freehand that is two steps, and the second step is easy to omit: a
check inside main() that CI never invokes, a #[test] that prints its
findings and asserts nothing, a --check-only mode that reports “Found N
invalid words” and returns Ok(()), a coverage percentage compared to
nothing. Every one of those type-checks, because () and Ok(()) are
perfectly good return types for “I printed something”.
So a gate implements the Gate trait in
crates/talkbank-parser-tests/src/gate.rs, whose only output is a verdict:
there is no method that yields findings without one, so “compute the list and
forget to act on it” is not expressible. Registration in ALL is the whole
mechanism, and a second gate checks the registry against the impl Gate for
declarations in the sources, in both directions, so a gate that is written and
not listed is a failure rather than a silence.
Two checks are not yet gates and are named in that module: verify_error_coverage.rs still prints a coverage percentage
and compares it to nothing, and validate_golden_words.rs keeps a path whose
only caller is its own main. A [[bin]] in that crate sets test = false,
which is target selection, so such a binary is excluded from --tests as well
as never being run by CI. If you are citing a check as a gate, run it, then
break it on purpose and watch it fail, before believing the citation.
The ratchets among them, and how you lower one
Three gates hold reviewed baselines: fabricated_ast (a per-crate
CEILING on new_unchecked and Span::DUMMY), error_code_demonstration (an
UNDEMONSTRATED list of identities absent from the canonical file snapshot), and content_catch_alls (an
UNPROTECTED list). Each baseline is a const in its own module, so lowering
one is an edit in the commit that earned it, reviewed like any other line.
There is no --write; instead each gate names exactly what to edit. The two list ratchets print
the entries that are accounted for and must go; fabricated_ast, whose
baseline holds numbers, prints its replacement row verbatim, so banking a drop
is a paste rather than a retyped number.
These counts are investigation tools, not CHAT policy. Backend/API-only diagnostics must not force tree-sitter to reconstruct their diagnostic identities. The inventory retains those identities explicitly while their specs continue to enforce structural rejection and legal controls. A residual entry does not claim the code has been reached or that coverage is complete.
Likewise, fabricated_ast::BOUNDARY_SPANS names reviewed diagnostic-admission,
splice-refusal and parsed-decoration tests. Their unknown-span sentinel counts
are checked exactly before the general count is computed; additions, removals,
missing files and unchecked-constructor substitutions refuse the inventory.
They are boundary evidence, not manufactured CHAT coverage. The general
per-crate ceiling remains strict. This lexical inventory does not prove test
semantics; retain the functional assertions and review same-count substitutions.
They need a Rust build:
cargo test -p talkbank-parser-tests --tests gates # every gate, verdicts only
cargo run -p talkbank-parser-tests --bin audit_gate_probes # + can each fail?
The layers
flowchart TD
unit["Unit + integration tests\n(cargo test)"]
specgen["Spec-generated construct tests\n+ the claim-judging fixture corpus"]
grammar["Grammar corpus\n(tree-sitter test)"]
ref["Reference corpus\n(corpus/reference/)"]
gates["Registered gates + CI"]
unit --> specgen --> grammar --> ref --> gates
Unit and integration. just test (cargo test --workspace --tests).
Doctests are separate and are NOT run by cargo test; run
cargo test --doc --workspace when you change public API examples.
Grammar corpus. cd grammar && tree-sitter test, the right gate for
grammar structure changes. It does NOT detect a stale parser.c; see
Grammar Workflow.
Reference corpus. corpus/reference/, organised by surface
(annotation/, audio/, ca/, content/, core/, edge-cases/,
languages/, tiers/, word-features/). It must stay at 100%, but it is a
SYNTHESIZED regression signal, not a validity authority. When a change rejects
a reference file, adjudicate the FILE against spec/, the grammar and real
corpus data, and fix the data or move it to spec/errors/. Weakening the
parser to keep a reference file green is the one response that is always wrong.
The corpus is not “the ultimate arbiter of correctness”; that reasoning would
entrench a bad fixture.
Coverage scope and completeness
The grammar-node inventory (corpus_node_coverage) measures concrete named
node presence, with explicit exclusions; it does not establish construct
combinations, semantic validity or Rust code coverage. Its result refuses
success when any input contains ERROR or MISSING nodes, even if every required
kind appeared. A paired instrument test preserves both the clean complete
control and the complete-but-recovered refusal. Keep this distinction when
using the inventory to select new reference cases.
An excluded node appearing is a policy-review trigger, not automatic progress.
The absent-node list includes unselected lexical alternatives, generic validation
fallbacks, recovery syntax and unsupported declarations. For example, clean CST
recognition of @Thumbnail still leads to E525 refusal during model lowering;
its witness belongs in the error specs. Verify support and validation before
promoting a specimen or removing an exclusion. Do not change denominators simply
to make the node inventory read 100%.
core/headers-ses-vocabulary.cha promotes the authored E546 legal control into
the reference corpus: standalone ethnicity values accompany combined
ethnicity/SES fields. Its explicit contract requires clean parsing, model
validation, all eight ordered SES payloads and exact CHAT output. The grammar
inventory therefore counts ethnicity_value instead of excluding it. Invalid
component mutations remain owned by the E546 specs, not the valid corpus.
When adding a file, verify that file-glob test discovery has rebuilt; a cached
integration binary can retain its previous case list. An explicit fixture
dependency and a nonzero named test witness prevent mistaking that old list for
verification of the new specimen.
Fragment-coordinate limits belong to the API-boundary track. Small UTF-8 inputs
at boundary origins exercise checked byte extents without allocating enormous
CHAT files. The public-boundary tests require exactly one rejection diagnostic,
no invented source location, correct multi-step offset rebasing, and unchanged
snippet-relative context and unknown spans. These are coordinate/API contracts,
not new CHAT specimens or canonical coverage credit.
The same track checks diagnostic presentation with a downstream producer’s
out-of-range offsets. An invalid start clamps to inclusive EOF, never the byte
before EOF (which may split a multibyte scalar); an invalid end clamps to EOF
without changing a valid start. ASCII, two-/three-/four-byte final scalars and
newline controls preserve sliceable source/context spans and the original
internal-failure identity. These injected locations are not parser recovery
specimens or evidence that any CHAT construct generates such offsets.
The public dependent-tier parser is checked separately from the range helper:
MOR/GRA, PHO and SIN slices from reference CHAT preserve their payload and exact
tier span across i32::MAX and at the last representable u32 byte extent.
Moving the same slice one byte beyond that extent must refuse admission with
one location-free diagnostic, without entering ordinary parser recovery.
The same track checks that mixed input findings and internal failures retain
their order and payload, reject completion regardless of severity or validation
profile, and propagate truncated diagnostic-output failures.
Replacement admission is nonempty in Rust, JSON and both parser backends.
ReplacementWords::new/TryFrom<Vec<Word>> return a typed error for empty input;
Replacement::new consumes that admission proof. Element editing cannot resize
the list. To filter or rebuild it, consume into_vec and re-admit the result;
an empty result is a decision for the caller, never a fabricated replacement.
Canonical replacement workflows still check valid spelling, ordering and writer
refusal. Separate API-boundary tests check empty constructor/JSON refusal.
Re2c tokenizes the replacement opening separately and uses its existing word productions, with a required first word and an optional remainder. Conversion neither reparses text nor falls back to a plain word. The reference compound and multiword alternatives check full-file source spans; the E208 specimen and unclosed-bracket controls require parse rejection. E208 has no model emit site; it is not preserved through an impossible empty model value. The primary parser’s E376/E342 recovery remains. Re2c spacing diagnostics identify its opening token rather than duplicating tree-sitter recovery locations.
A coverage percentage answers which instrumented code ran in one selected configuration. It does not establish that every supported source file, feature, platform or generic instantiation was present. Keep the two acceptance questions separate; missing executable code is not covered code.
| Evidence | What it establishes | What it does not establish |
|---|---|---|
| Canonical spec/reference workflows | Observed parser, model/validation and transform behavior for authored controls and errors | Completeness of supported CHAT or unlinked code |
| Source census and macro attribution | Written bodies and explicit ownership of macro invocation sites | Every derive expansion or generic instantiation |
| Feature-specific public workflow tests | Behavior in the selected configuration | Other feature combinations or platforms |
| Fault-injection boundary tests | Tool-failure and refusal contracts | Additional malformed-CHAT specimens |
| Build/package/platform checks | The particular build, installation or platform boundary checked | Canonical semantic coverage or native UI acceptance |
The canonical harness enables model async and channels together and uses
the transform’s default validation-runner feature. Test runner-disabled
consumers separately: workspace feature unification can conceal a missing
feature guard. Do not blend these checks or internal unit-test profiles into
canonical coverage totals. Generated code must remain separately identified.
Report three distinct evidence tracks: canonical spec/reference workflows,
API/internal-failure boundaries, and their combined coverage. The initial
boundary population reuses closed_newtype_consumer_view and
public_error_types; these are API contracts, not additional CHAT specimens.
Its catalog-boundary module deliberately supplies independent diagnostics:
typed-node fixes must refuse out-of-source locations, and unsupported repair
requests cannot infer missing participant/language facts from a code or span.
Codes attached to ordinary words, canonical media names and URL headers must
not manufacture a repair. Real marker/comma spec diagnostics are positive
controls against an always-refuse implementation. These synthetic requests
belong only to the boundary track; they are not parser-emitted CHAT findings
and do not establish that every diagnostic is bound to its original parse.
The JSON-export boundary module supplies a deliberately failing downstream
serializer to each compact/pretty, schema-checked/unchecked API. Each must
invoke it once, return no output, and retain the underlying serialization
error instead of claiming schema invalidity. Generic intermediate arrays are
permitted by unchecked export but refused by the CHAT schema; malformed JSON
retains its decoding error and position. These are generic wire/API contracts,
not extra CHAT fixtures or evidence of embedded-schema load failure.
The failure-evidence consumer contract carries an explicitly injected producer
fault through diagnostic admission, ownership transfer and the public pipeline
error. Its user-facing text must state that validity was not determined; the
original diagnostic and source location survive recovery of the error payload.
Keep existing canonical tests in their original population, including public
output-refusal and cache/lifecycle workflows already measured there.
Measure the combined track with both populations in the same pinned runner, source revision, feature configuration and instrumentation. LLVM then counts shared lines, regions and branch outcomes once. Never add covered counts or percentages from separate runs. All three tracks retain identical source exclusions and distinct comparison identities; report each separately and keep source/feature completeness as an independent obligation. A configured track or a passing uninstrumented test is not a coverage measurement.
For macro and generic APIs, identify the actual type-specific public workflow; an executed shared source line does not prove that every expansion ran. A source classifier’s unexpanded attribute/token inventory is a list of obligations, not a compiler expansion or a reachability proof. Preserve internal-failure and recovery handling until producer invariants justify narrowing it; do not invent CHAT fixtures solely to force tool faults or remove code from the denominator. An unused generic coverage record need not identify an untested concrete type. Rustc can emit dummy records with placeholder type arguments even when downstream code instantiates the same library function. Check concrete function counts, region counts and binary symbols against the pinned compiler’s unused-function mapping implementation before inventing a fixture. This does not justify dropping records or changing coverage denominators; retain the raw result and document the bounded attribution.
The tier-construction reference contract exercises macro-generated lexical wrappers through borrowed/owned construction, text views, formatting and JSON. Nonvocal constructors preserve parsed labels across begin, end and simple forms; these open wrappers do not confer validation or source-ownership proof. Reuse retained per-type function/region mappings when checking expansion execution, keeping unexecuted generic instances distinct from executed shared lines. The header-text contract independently requires reference witnesses for all 23 selected wrappers: PID, situation, gem label, tape location, location, room layout, birthplace, transcriber, warning, activities, background, page, videos, thumbnail, font, window, color words, participant name/role and ID corpus/group/education/custom fields. Each preserves its parsed payload through borrowed/owned constructors, From, AsRef/Deref, Display, WriteChat and JSON reconstruction, while rejecting non-string JSON. Interned roles are tested for text/value equality, not an unsupported universal pointer-identity promise. Rebuilding a header field is not semantic validation; structured headers and checked filename/date fields retain separate admission contracts. These per-type witnesses close the selected adapter contract, not universal macro or feature completeness.
Running specific tests
cargo test -p talkbank-model # one crate
cargo test -p talkbank-parser-tests --tests mor # by name filter
cargo test -p talkbank-model -- --nocapture # show stdout from passing tests
--nocapture goes after --; it is an argument to the test harness, not to
cargo.
What to run when
| What you changed | Run |
|---|---|
Grammar (grammar.js) | the whole Grammar Workflow, including the typed-traversal regeneration |
| Parser (CST to model) | cargo test -p talkbank-parser, plus parser equivalence and roundtrip |
| Model (types, validation, alignment) | cargo test -p talkbank-model, plus roundtrip |
| CLI | cargo test -p chatter |
| LSP | cargo test -p talkbank-lsp |
| Spec files | regenerate per Spec Workflow, then just test-spec and the gate registry |
| Either registry (symbols, form markers) | just test-spec, which includes the drift gates |
| Anything, before pushing | just gate, or just push which runs it |
Mutation testing
cargo-mutants finds code that can be changed without any test failing, which
is the real coverage question. It is not part of CI; run it periodically after
significant changes.
cargo install cargo-mutants
cargo mutants -p talkbank-model --file 'src/validation/**' --timeout 180
cat mutants.out/missed.txt # mutations no test caught
Scope it, and read the result as a work list rather than a score. The
validation tree is the highest-value target: chatter validate is the
authority on CHAT validity, so a mutant that survives there is a rule that can
be silently disabled. Running -p talkbank-parser unscoped spends most of its budget on src/generated_traversal.rs,
over half that crate and generated, where a survivor indicts the generator
rather than this repository. To see the size of a target before committing an
evening to it, use cargo mutants --list --file '<glob>'.
Each job runs a full workspace build peaking around 8 GB, and the failure mode
is an out-of-memory kill during overlapping linker phases rather than steady
state, so measure peak memory at a small --jobs before raising it. A fixed
--jobs 1 is not a property of the tool; the right value depends on the machine.
Configuration is mutants.toml at the repo root.
Adding tests, and when not to
Before writing a test, ask whether a TYPE could make the bad value unrepresentable instead. A test guarding an invariant is a standing admission that nothing enforces it; changing the type deletes the test, covers callers the test never enumerated, and fails at the point of the mistake rather than in CI. Reducing the test count this way is an explicit pre-1.0 goal.
What legitimately survives that question: wire formats, roundtrips between a formatter and a parser that are two separate functions, measurements, policy choices with real alternatives, and behaviour a signature cannot describe. A surviving test says which of those it is, in its own docstring.
When a test is the right answer:
- Model behaviour: the crate’s
tests/directory or a#[cfg(test)]module. - Grammar shape or validation contract: add or update a SPEC and regenerate. A parser bug fixed without a spec will regress.
- A repository-wide invariant: implement
Gateand register it, rather than writing a binary that prints findings.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Coding Standards
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
Rust Conventions
- Edition: 2024
- Formatting:
cargo fmtbefore every commit - Linting: clippy is release-time, not per push and not per edit. Stated in full under Rust Conventions below; do not restate it here, because this file already carried the policy twice and the two copies claimed different things.
Error Handling
- No panics for recoverable conditions, use
thiserror/miettefor error types - Library code uses the
ErrorSinktrait for error reporting, notResult - Use
ParseOutcome<T>in parser code (parsed or rejected)
Logging
- Library crates use
tracing(neverprintln!oreprintln!) - CLI binaries write to stdout (results) and stderr (diagnostics)
- Use appropriate log levels:
error!,warn!,info!,debug!,trace!
Naming
- Follow standard Rust conventions (snake_case for functions, CamelCase for types)
- Conventional Commits for commit messages:
<type>[scope]: <description>- Types:
feat,fix,refactor,test,docs,chore
- Types:
Dependencies
Preferred crates:
clap: CLI argument parsingserde: serializationmiette: user-facing diagnosticsinsta: snapshot testingtracing: structured loggingrayon/crossbeam, concurrencysmallvec: small-buffer optimization
Code Organization
- Keep crate boundaries clean, lower crates should not depend on higher ones
- The model crate should not depend on any parser
- Parsing code should not depend on serialization/transform code
- All CHAT parsing and serialization goes through the AST, never ad-hoc string manipulation
- Treat 10 or more named struct fields as an audit trigger. Wide boundary or
report records can be acceptable, but wide runtime state bags need explicit
review. See
architecture/chat-model/wide-structs.md.
Testing
- Prefer spec-driven tests over hand-written tests for parser behavior
- Use
cargo testfor unit tests (except doctests) - Snapshot tests with
instafor complex output comparisons
Generated Files
Never hand-edit generated artifacts:
parser.c: generated fromgrammar.jsgrammar/test/corpus/: generated from specscrates/talkbank-parser-tests/tests/integration/generated/: generated from specscrates/talkbank-model/src/generated/symbol_sets.rs: generated from symbol registry
Always regenerate from source inputs.
Full Rust Standards Charter (canonical)
Edition and Tooling
- Rust 2024 edition.
cargo fmtbefore committing. Usecargo fmt(not standalonerustfmt) for workspace-consistent formatting.- Prefer
cargo testfor faster parallel-per-test execution. Usecargo test --docfor doctests (they are not part of the normal run those). - Clippy is release-time, in
just release-lintand inrelease-lint.ymlon a tag, never per push and never per edit: each pass is its own cargo unit that recompiles the workspace. It is a single pass (--workspace --all-targets) carrying-D warnings, so ANY finding is red, not only a violation of the panic family the workspace[lints]table denies; test code relaxes that family via in-source attributes. A finding you intend to keep gets a scoped#[allow]naming the reason. See the clippy policy section above.
Error Handling
- No panics for recoverable conditions. Use typed errors
(
thiserror); usemiettefor rich diagnostics where appropriate. - No silent swallowing. Every unexpected condition must be
handled with explicit error reporting, no
.ok(),.unwrap_or_default(), or silent fallbacks that hide bugs.
Output and Logging
- Library crates:
tracingmacros (tracing::info!,tracing::warn!, etc.), neverprintln!/eprintln!. - CLI binaries: results on stdout through
outln!/out!(thechattercrate’sstdout.rs), neverprintln!/print!, which panic when a consumer such asheadcloses the pipe; the shared writer ends the command with exit status 1 instead.eprintln!for diagnostics;tracingfor debug logging. - Test code:
println!is acceptable (cargo captures it).
Lazy Initialization
LazyLock<Regex>(fromstd::sync) for constant regex patterns. Never callRegex::new()inside functions or loops.OnceLockfor per-instance memoization of runtime-determined values.- Prefer
constwhen possible (even better than lazy). - All lazy init via
std::sync, no external crate dependencies needed.
Type Design
- No boolean blindness. Enums over bools for anything beyond
simple on/off. This is a hard rule.
- Banned: 2+ bool parameters on a function, 2+ related bool
fields on a struct, opposite bool pairs (
foo/no_foo), bool return where meaning is unclear without reading docs. #[derive(Default, clap::ValueEnum)]enum with named variants. For clap CLI args, use#[arg(value_enum)]instead of--flag/--no-flagpairs.- OK as bool:
verbose,force,quiet,dry_run, singleinclude_*/skip_*flags, anything where the parameter name fully communicates whattruemeans.
- Banned: 2+ bool parameters on a function, 2+ related bool
fields on a struct, opposite bool pairs (
BTreeMapfor deterministic JSON in tests and snapshot tests (notHashMap). Ensures consistent, reviewable diffs.- Prefer explicit enums over ambiguous
Optionwhen there are multiple meaningful states.
Newtypes Over Primitives
- No primitive obsession. Domain values must have domain types. Function signatures should be self-documenting through type names, not parameter names.
- Use newtype structs (e.g.,
struct TimestampMs(u64),struct SpeakerId(String)) or theinterned_newtype!/string_newtype!macros fromtalkbank-model. Newtypes should implementDisplay,From/Intofor the underlying type, and deriveClone,Debug,PartialEq,Eqas appropriate. - Scope: Applies to public API boundaries, struct fields, and function signatures. Local variables inside a function body may use bare primitives when the context is unambiguous.
- Parsing boundaries: Parse raw strings into newtypes at the boundary (file I/O, CLI args, IPC). Interior code should never handle raw strings for typed values.
- No ad-hoc format parsing. Use real parsers (JSON:
serde_json, etc.) not regex or string splitting for structured formats. Regex is appropriate only for flat text pattern matching (search, normalization, validation of simple formats).
Integer Discipline
- Distinguish meaning. Not all
usizevalues are interchangeable. Separate:- Index: position into a collection (
UtteranceIndex,GraIndex) - Count: accumulated quantity (
WordCount,UtteranceCount) - Limit: upper bound for iteration or reporting
(
UtteranceLimit,WordLimit) - Threshold: minimum value for inclusion
(
FrequencyThreshold) - ID: opaque identifier (
NodeId,SpeakerIndex)
- Index: position into a collection (
- Non-negative quantities use unsigned types; newtypes enforce domain semantics.
- No bare numeric literals except
0,1, and simple loop bounds. All other numbers must be named constants. Assess whether each constant should be configurable.
Closed-Set Strings and Constants
- Closed sets must be enums. If a string value comes from a
known finite set (tier labels, command names, output formats),
represent it as an
enumwith aFromStrparser andDisplayserializer. UseOther(String)escape hatch only when the set is genuinely extensible. - All remaining string literals must be defined constants. No
scattered
"mor"or"cod"strings, useTierKind::Mororconst DEFAULT_TIER: &str = "cod". - Config defaults: Use
constvalues or enum variants inDefaultimpls, not"string".to_owned()(avoids runtime allocation, makes the default visible at the type level).
File Path Discipline
- File paths use
PathBuf/&Path, neverString. Convert to strings only at display/serialization boundaries via.display()or.to_string_lossy(). - Distinguish base filename (e.g.,
MediaFilenamenewtype, no extension) from full filesystem path (PathBuf). - Use
.display()for user-facing output;.to_string_lossy()only for cache keys or hashing.
Configurability
- Hardcoded thresholds and limits belong in config struct fields with documented defaults.
- If a default is useful to change per-invocation → CLI flag.
- If a default is useful to change per-user → future
defaults.tomlfile (not yet implemented). - Config structs must be constructible in tests without filesystem or network access.
Rustdoc as Primary Documentation
- Types are the primary documentation layer. A reader of crates.io rustdocs should understand the domain by reading type definitions alone.
- Every
pubtype and function must have a doc comment explaining role, ownership, invariants, and CHAT manual references where applicable. - Newtypes must document valid values, units, and meaningful operations.
- Enum variants must document when each variant applies.
File Size Limits
- Recommended: ≤400 lines per file.
- Hard limit: ≤800 lines per file (must be split).
Testability
- No global mutable state. All command state flows through
explicit
Statetypes (theAnalysisCommandtrait pattern). Enforce this going forward. - Config structs must be constructible in tests without filesystem, network, or environment setup.
- Stateful resources (caches, pools, registries) must accept injected dependencies for test control.
Refactoring Triggers
Stop and refactor when you see:
x: i32, y: i32for domain data → use domain structsstart_ms: u64, end_ms: u64→ useTimestampMsnewtype orTimeSpanstructfn foo(lang: &str, speaker: &str, path: &str)→ useLanguageCode,SpeakerId, typed path- Multiple booleans for state → use enum with variants
fn foo(a: bool, b: bool)or--flag/--no-flagpairs → use enum withclap::ValueEnumfn parse() -> Option<T>where failure reason matters → useResult<T, ParseError>match s { "win" => ... }on raw strings → parse toenumat boundary"mor"or"cod"string literals → useTierKind::MororTierKind::Codlimit: usizeormax_X: usize→ use domain-specific newtype (UtteranceLimit,WordLimit)- Bare
0.5or60in logic → named constant or config field - Regex or
split()/find()on XML, JSON, or other structured formats → use a proper parser
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Coding Standards and Engineering Practices
Status: Current Last updated: 2026-05-21 08:38 EDT
Objective
Set enforceable, language-specific standards that reduce ambiguity and improve long-term maintainability.
Global Standards
- Prefer explicit domain types over ad-hoc strings.
- Keep parsing, validation, and rendering logic separated.
- Eliminate magic numbers/strings/paths via named constants and config.
- Treat generated code as immutable artifacts.
- Require tests for every bugfix and behavior change.
Rust Standards
- Enforce formatter and clippy in CI.
- Minimize
#[allow(clippy::...)]; each allowance needs rationale. - Prefer small focused modules with clear ownership.
- Public APIs require doc comments with examples and error behavior.
- In parser code, disallow
ErrorSink+Option<T>signatures for fallible parse operations.- Use explicit outcome enums or
Resultwith structured diagnostics. - Guardrail script:
scripts/check-errorsink-option-signatures.sh.
- Use explicit outcome enums or
- For model enums that encode validation state, require
ValidationTaggedderive.- Explicit annotation:
#[validation_tag(error|warning|clean)]. - Naming-convention fallback (per
crates/talkbank-derive/src/validation_tagged.rs:118-123): variants ending inError→Error; variants ending inWarningORUnsupported, plus a variant named exactlyUnsupported, →Warning; otherwise →Clean.
- Explicit annotation:
Grammar Standards
- Grammar rules must map to documented token/category semantics.
- No duplicated symbol sets in free-form literals.
- Every non-obvious precedence/conflict decision must include rationale.
Spec and Generator Standards
- Spec files must follow strict metadata template.
- Generators must be deterministic and pure with respect to inputs.
- No hardcoded user-specific paths in docs or generated outputs.
Magic Value Policy
Disallowed
- Inline path literals tied to local machines.
- Unnamed numeric constants encoding protocol behavior.
- Repeated header/tier string literals across modules.
Required
- Central constants/modules:
- path defaults,
- tier/header prefixes,
- token categories,
- formatting policies.
Review and PR Standards
- PR template must include:
- subsystem touched,
- contract impact,
- generated artifact impact,
- tests added/updated,
- docs updated.
- Require at least one reviewer with subsystem ownership for core modules.
Internal Decision Records
Adopt short ADR format in the book’s architecture section:
- context,
- decision,
- alternatives considered,
- consequences,
- rollback path.
Acceptance Criteria
- Coding standards are documented once and enforced automatically.
- Magic values are systematically reduced and tracked.
- Every behavior change includes tests and doc impact assessment.
- Architecture decisions are recorded and discoverable.
This page last changed: 2026-06-21 (commit 1952fb27). The whole book last changed: 2026-10-07 (commit 5e895791).
Toward Chatter 1.0
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
This is the current readiness record. Version 1.0 means a documented compatibility contract and reproducible evidence, not merely a release-number change. No date is promised here.
Where the numbers live
just spec-status reports the current spec counts (error specs by status,
examples verified, deferred or failing) and the CHECK assessment summary. These
are spec and example counts, not a count of independent validation rules. To
inspect the deferred specs, run
cargo run --manifest-path spec/Cargo.toml --bin spec_status -- --deferred:
prioritize rules reachable through supported CHAT input, preserve legal and
invalid examples together, and use the observation snapshot to adjudicate
backend differences.
The generated CHECK assessment owns the current scope, adjudication counts, completion status and reopening criteria. This readiness record does not maintain a second CHECK status.
Release conditions
- Define the stable CLI, exit-code, JSON/schema and Rust API surfaces. State which surfaces remain experimental and test the chosen compatibility contract.
- Adjudicate every deferred spec for the stable surface. Implement missing behavior or document why the rule is deprecated, unreachable or deferred; do not activate a spec merely by changing its status.
- Ground CHECK obligations against an identified executable and fixture set. Keep deliberate divergences justified in the manifest. Every observed disagreement needs a spec or validator adjudication before data repair.
- Review typestate transitions at parsing, validation, repair and writing. Output-producing APIs must consume the evidence their operation requires; a curated label, parsed AST or recovered node is not proof of validity.
- Keep generated artifacts current and reader-facing documentation consistent with actual commands, validation behavior and release mechanics.
- Pass the established release gate and platform jobs on the exact release commit. Verify packaged artifacts and deployment using the release runbook.
Evidence boundaries already in place
These are properties of the code that the release conditions build on; each is a typed or gated boundary rather than a convention.
- CHECK mapping audit. It reads the compiled spec registry, renders only mapping evidence (unmapped or nonempty curated states) and has a report-currency integration gate. Explicit CLAN grounding cannot silently succeed without its wrapper.
- Source-bound diagnostic indexes.
SourceIndex<'source>owns the line boundaries and immutably borrows their source; its constructor is the only way to pair them, so a line map from another source cannot be supplied. Indexed callers useenhance_errors_with_index. One-off lookups scan the source prefix; repeated batches use one explicit index (O(source bytes) construction, O(log lines) lookup). - Prosodic validation.
Wordpasses through a privateProsodicWordmeasurement before its prosodic checks. It borrows the word immutably and owns the stress counts and the first and last spoken positions; only its constructor can assemble them, so they cannot describe a different word or outlive a mutation. - JSON conversion.
JsonSchemaPolicyselects serialization after the shared named CHAT parse, so skipping schema validation cannot discard transcript identity or bypass E531. - Deserialized model values need validation. CHAT text cannot construct an
empty text, phonetic or shortening segment, so E251 is implemented at the
model validation boundary (serde-constructed values), covered by
test_e251_deserialized_word_content. A wrapper type’s name is not evidence that deserialized content is valid. - Schema. The generated JSON schema is native Draft 2020-12, with
$refsiblings kept as the dialect allows; consumers must support that dialect. The generatedBracketedItemdefinition is exercised bygenerated_ref_siblings_enforce_tag_and_payload.just schema-genis an explicit operation; the normal suite checks currency without writing it. - Generated outputs preserve unchanged files. Generators publish through
scripts/generate_if_changed.py(staging, atomic replace, identical bytes keep their mtime, a failed generator leaves the existing file). The spec artifact writer deletes only explicitly retired names in shared directories, andGeneratedDirproves exclusive generated ownership before pruning.just node-types-checkandjust grammar-generate-checkverify these. See Testing. - Publication set. The publication check derives every non-first-wave
workspace package from Cargo metadata and requires
publish = falsefor it;--metadata-onlyruns it without package assembly. It is evidence of publication policy and metadata only; full package and registry verification remain part of the release review. - re2c backend. It borrows the caller’s source, owns only reconstructed
subtoken text (
Cow<str>) and has no productionBox::leak.just verify-vendored-lexerconfirms the committed lexer against the pinned generator.SinTier::from_tokensreturnsResult; JSON remains an unvalidated boundary, so utterance validation descends through%sinitems to report empty tokens.
Reproducing the CHECK evidence
From the Chatter checkout:
cargo test -p talkbank-parser-tests --test integration check_mapping_audit:: --offline
cargo test -p talkbank-parser-tests --test integration check_validity_parity:: --offline
CHATTER_CLAN_RUN=/path/to/clan-run.sh \
cargo test -p talkbank-parser-tests --test integration clan_check_grounding \
--offline -- --ignored --nocapture
The wrapper resolves its default CHECK executable from the CLAN checkout;
CLAN_BIN_DIR selects another binary directory. Retain the file-mode PTY
wrapper. Record the CHECK executable, wrapper and parity manifest SHA-256 and
the CLAN checkout HEAD with any result: the executable hash identifies what
actually ran, while the checkout HEAD identifies only the source inspected.
These results cover the committed fixture set, not all possible CHAT input.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CI and Release
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
Pre-Merge Verification
Run the shared local gate from Developer Verification Checks:
just gate
This runs the checks used by per-push CI, including doctests, both Rust
workspaces, generated-artifact currency, and the book. Wait for GitHub Actions
on the exact pushed commit before announcing it as ready. Release-only checks
run separately through just release-lint. That recipe first checks app versions
and the changelog section/link, before formatting and compiler checks, so a
missing release entry fails before compilation.
Generated artifact drift
After changing grammar, spec, or a registry, run just regen, then just test.
The regeneration recipe builds derived artifacts in dependency order; currency
tests detect stale output. Never hand-edit generated artifacts.
See Spec Workflow and spec/AGENTS.md for the current
source-of-truth guidance.
Release Process
TalkBank/chatter is the public release source of truth. release.yml
(cargo-dist) builds the CLI artifacts and publishes the GitHub Release;
release-desktop.yml, called from it, creates the unpublished draft and adds
the desktop installers first. Signing differs by platform as described below.
The full publication order is in docs/strategy/coordinated-release.md.
Cutting a release: the two-command procedure
The version literal lives in many places (the workspace version, every
internal path-dep pin, the desktop package.json, the CHANGELOG section),
and the tag is the release trigger, so both steps are mechanized and
fail-closed. Hand-editing version fields or tagging with raw git tag is
how releases break (v0.1.1 shipped a desktop version mismatch; v0.5.0
tagged a bump commit before its CI reported and the desktop build died on
drift CI would have caught). The procedure:
just release-bump X.Y.Zrewrites the canonical[workspace.package] version, everypath = "crates/…"pin, andpackage.jsonand both root-version fields inpackage-lock.json, then refreshes both Rust lockfiles (root +spec/). The app-version check exercises this command against temporary manifests and lockfiles before checking the checkout, including independent drift in each lockfile version and preservation of dependency versions.- Write the
## [X.Y.Z]CHANGELOG section (the one deliberately manual step; every gate enforces its presence). - Format, run
just release-lintandjust gate, then squash the commits since the last push into one release commit whose message is the CHANGELOG section. Verify the gate on the squashed tree and, with maintainer authorization, push and wait for CI on that commit. The content stamp survives a squash that leaves the checked bytes unchanged. Push rarely; preserve already-published commits instead of rewriting history at release time. just release-tag X.Y.Ztags and pushesvX.Y.Z, refusing on a dirty tree, an unpushed HEAD, any version-copy drift, a missing CHANGELOG section, or CI/Cross-platform not yet green on the exact tagged commit.
After the tag: release-tag-dispatch.yml dispatches release.yml for that
tag, which builds, runs the desktop publication job, publishes and then adds
the app banner; verify the release page carries the CLI archives, the LSP
standalone artifacts, the desktop installers and the CHANGELOG notes before
announcing.
Workflows that actually exist in this repo
| Workflow | Purpose | Notes |
|---|---|---|
.github/workflows/ci.yml | Main build/test/book CI | Primary shared signal on pushes and PRs |
.github/workflows/cross-platform.yml | Cross-platform build coverage | Supplements the main CI workflow |
.github/workflows/crates-io-foundation.yml | First-wave crates.io readiness | Checks foundation-crate metadata, package surfaces, hold-backs, and publish order |
.github/workflows/release-tag-dispatch.yml | Tag trigger | On a version tag push, dispatches release.yml for that tag |
.github/workflows/release.yml | cargo-dist release automation | Generated by dist (dist generate; never hand-edit). Builds CLI artifacts, calls the desktop publish job, then publishes the draft in its announce step |
.github/workflows/release-desktop.yml | Desktop installer release automation | dist publish job: creates the draft with dist’s announcement title and CHANGELOG body, builds and uploads the installers and updater bundles, and verifies the candidate; workflow_dispatch runs build-only |
.github/workflows/release-app-banner.yml | Release page banner | dist post-announce job: prepends the desktop app banner to the published release notes |
.github/workflows/release-lint.yml | Release-time lint | just release-lint: clippy over both workspaces plus the feature-off build. Runs on a version tag and on workflow_dispatch, never per push |
.github/workflows/clippy-rolling.yml | New-stable clippy lints | Weekly scheduled clippy, so a new stable’s lints surface within days |
Current release stance
Cross-platform workspace verification uses Cargo’s normal Windows dynamic C
runtime through a job-local TAURI_CONFIG override. Tauri 2.7’s static-runtime
build shim otherwise shadows msvcrt.lib for unrelated workspace doctests.
Doctests remain enabled. This override does not apply to desktop release
packaging, which retains Tauri’s static-runtime default; passing workspace tests
does not replace verifying the packaged Windows artifact.
release.ymlis about workspace artifact packaging via cargo-dist, not about crates.io publication.- The first-wave crates.io path is documented separately in
Crates.io Publication and is checked by
just crates-io-foundation-checkplus.github/workflows/crates-io-foundation.yml.
Desktop release workflow: how the release jobs compose
dist-workspace.toml sets create-release = false, github-release = "announce", publish-jobs = ["./release-desktop"] and post-announce-jobs = ["./release-app-banner"]. After the CLI artifacts build, release.yml calls
release-desktop.yml, which creates the draft release with dist’s announcement
title and body, builds the Tauri installers and updater bundles, uploads them
and verifies the candidate. Only when that job succeeds does the announce step
upload the CLI archives, checksums and installer scripts and publish the draft;
the banner job then edits the published notes. Nothing polls, and a failed
desktop build leaves an unpublished draft. Two platform notes baked into the
workflow:
- macOS: Tauri signs, notarizes, and staples the
.app, but NOT the.dmgit wraps around it. The workflow therefore submits the.dmgitself to the notary service and staples it, then verifiescodesign,spctl, andstapler validateon both artifacts. The signing identity is supplied via environment, never hardcoded intauri.conf.json. - Windows / Linux: artifacts are currently unsigned by decision; see
docs/strategy/distribution-and-signing.md(“Decisions, 2026-06-12”) and the SmartScreen guidance in the install docs.
Release secrets (Actions secrets on this repository)
Required by the macOS jobs of release-desktop.yml (and by cargo-dist
macOS codesigning if macos-sign is enabled, which uses the separate
CODESIGN_* names documented in the strategy doc):
| Secret | Content |
|---|---|
APPLE_CERTIFICATE | base64-encoded Developer ID Application .p12 |
APPLE_CERTIFICATE_PASSWORD | password for the .p12 |
APPLE_SIGNING_IDENTITY | full identity string, Developer ID Application: <Name> (<TEAMID>) |
APPLE_API_KEY | App Store Connect API key ID (notarization) |
APPLE_API_ISSUER | App Store Connect issuer ID |
APPLE_API_KEY_CONTENT | contents of the AuthKey_*.p8 file |
Rotation: replacing the certificate or notary key means updating these secrets and nothing else; no workflow edits are needed. A maintainer must re-create all of them on any new repository (secrets do not transfer).
The development loop
The loop exists because a single parser fix can cost a day to the process around it rather than to the fix.
- Inner loop:
just test. Write the failing test or the type change first, then make it green.clippyandfmtare run before a release, not per edit. - After any change under
grammar/,spec/or a registry:just regen, thenjust test. Every derived artifact has a currency test, and they are far cheaper to satisfy together than one gate run at a time. - Before committing: review the final diff once.
- Before pushing:
just gate, once. It mirrors per-push CI exactly, so CI is a confirmation and never a discovery. The pre-push hook refuses a push without the stamp; the stamp hashes tree content, so a gate run on uncommitted changes stays valid once the same bytes are committed. Clippy and the feature-off build arejust release-lint, run before a release. - Releasing: format,
just release-lint, gate, squash every commit since the last push into one release commit carrying the changelog section, gate once more, push, wait for CI, thenjust release-tag. Push rarely and preserve already-pushed history.
Nothing on this path needs data that is not in the repository.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Crates.io Publication
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
Scope
The 1.0 publication contract includes foundation libraries and registry-installable CLI/LSP binaries. Publication is a deliberate maintainer action, not a tag-triggered release path. The foundation-named check covers the complete dependency closure below; enabling publication is not proof of readiness.
The publication order is:
talkbank-buildtree-sitter-talkbanktalkbank-derivetalkbank-modeltalkbank-cachetalkbank-parsertalkbank-parser-re2ctalkbank-transformsend2clantalkbank-llmtalkbank-lspchatter
talkbank-build is build-only support for the model and parser source
fingerprints and must be published before those consumers.
talkbank-parser-re2c is included because
talkbank-transform has a runtime dependency on it. Holding it back would
make talkbank-transform unpublishable. Inclusion in the dependency closure
does not promise parser equivalence: re2c remains experimental.
The CLI additionally requires send2clan and talkbank-llm at runtime. They
must be published before it; optional Cargo dependencies would still require
registry resolution and are not a way to hide unpublished packages.
Every workspace package outside this publication set must be explicitly marked
publish = false. The check derives this complement from Cargo metadata,
so a newly added crate cannot silently escape the publication decision.
Internal test, vocabulary, desktop and task-runner packages remain held back;
the script prints the complete current set. CLI/LSP binary releases and desktop
installers remain required alongside registry installation.
Before declaring registry installation supported, verify the actual published
candidate with cargo install chatter --locked and
cargo install talkbank-lsp --locked, including CLI validation and LSP protocol
smoke tests. These are acceptance targets, not a claim that current registry
versions are available. MSRV, supported platforms and post-1.0 compatibility
policy still require explicit decisions and candidate-bound verification.
What the repo automates
The existing foundation-named entry points cover the full publication set:
| Surface | Purpose |
|---|---|
just crates-io-foundation-check | Local preflight for crates.io readiness |
bash scripts/release/check-foundation-publication-readiness.sh --metadata-only | Fast manifest, dependency and hold-back review without packaging or registry access |
.github/workflows/crates-io-foundation.yml | CI enforcement for metadata, package surfaces, hold-backs, and publish order |
The readiness check enforces:
- required crates.io metadata (
repository,homepage,keywords,categories,readme) - readme-file existence
- package file enumeration for every selected crate via
cargo package --list - the selected runtime and build dependency graph
publish = falseguards on every workspace crate outside the publication set- real
cargo publish --dry-runchecks for the standalonetalkbank-buildandtree-sitter-talkbankcrates
The metadata-only mode uses locked Cargo metadata and reads README paths. It does not validate assembled package contents or registry resolution and cannot replace the full pre-publication check.
Important limitation: Cargo cannot fully dry-run the bootstrap wave
For the first publication of an interdependent workspace, cargo publish --dry-run is not a complete CI gate for every crate. Cargo rewrites path
dependencies to registry dependencies while preparing the package. That means a
crate such as talkbank-model cannot complete a registry-style dry-run until
its prerequisite talkbank-derive already exists on crates.io.
So the current automation is intentionally honest:
talkbank-buildandtree-sitter-talkbankget real crates.io dry-runs because neither depends on an unpublished workspace crate.- The remaining selected crates are validated by metadata, readme, and
dependency checks before publication. (No MSRV is declared yet; set a
deliberate
rust-versionand re-add an MSRV check when publication is actually pursued.) - As each prerequisite crate lands on crates.io, rerun targeted
cargo publish --dry-run -p <crate>checks for the later crates before publishing them.
This is a real limitation of the initial bootstrap wave, not a missing script. If we later want full registry-resolution rehearsal before publication, that requires a staging registry/local index strategy, not just another shell loop.
Publication procedure
Before publishing anything:
- Verify crates.io name availability for every selected package.
- Run
just crates-io-foundation-check. - Ensure
.github/workflows/crates-io-foundation.ymland the main CI workflow are green on the commit you intend to publish. - Publish in the order in Scope above, waiting for the crates.io index to observe each crate before moving to the next. That list is the single documented order; do not omit build-only or CLI runtime dependencies.
- After each prerequisite becomes visible on crates.io, rerun any newly-unblocked
cargo publish --dry-run -p <crate>checks before the next publish step.
Example command shape:
cargo publish -p tree-sitter-talkbank --locked
Tagging policy
Do not use version tags to drive crates.io publication from this repo.
.github/workflows/release.yml is reserved for cargo-dist GitHub Releases of
dist-enabled artifacts. Crates.io publication remains a deliberate manual
maintainer flow.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Testing and Quality Gates
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
How local verification relates to CI. The local commands themselves live in Developer Verification Checks, which is their single owner; this page says which of them CI repeats and which it does not.
Local pre-merge contract
just gate runs everything CI runs; just push runs it and then pushes. See
dev-checks for what it contains and why each
step catches something the others cannot.
The list is deliberately not reproduced here: a set of commands assembled from
memory forgets the easiest one to forget, cargo test --doc --workspace, which
is exactly the one a green just test gives no signal about.
Never-regress gates
The CHAT core has five gates that must stay green for any change touching the grammar, parser, model, validation, serialization or alignment: parser equivalence, roundtrip idempotency (which carries reference-corpus coverage in the same test), the generated spec tests, the validation error corpus, and the gate registry. Each has a fast targeted command, listed with what it protects under Testing, Never-Regress Gates.
Those commands take --tests <filter>, not --test <name>: each crate has one
integration binary, so a per-file target name errors out.
A red gate is a bug until proven otherwise, never a test expectation to quietly update. That rule has teeth in both directions: a diagnostic that LOOKS better after a change earns the same scrutiny as one that looks worse: a specific, plausible-looking error message can be a symptom of corruption rather than an improvement.
What CI actually runs
.github/workflows/ci.yml is the authoritative shared signal, and it runs
these jobs:
| Job | Checks |
|---|---|
rust | build, test, and the spec/ workspace. NOT clippy: that is release-time |
wasm | the re2c parser still compiles for wasm32 |
book | mdBook build plus a lychee link check |
rust-version-sync | version pins in workflows, and doc date headers |
app-version-sync | the desktop app version tracks the workspace version |
shellcheck | every tracked shell script, default severity |
grammar | the grammar’s own checks |
dependency-audit | dependency advisories |
Separate workflows cover release-time lint (release-lint.yml: clippy over
both workspaces plus the feature-off build, on a tag or on demand),
cross-platform builds (cross-platform.yml), the weekly scheduled clippy
(clippy-rolling.yml), crates.io readiness, and the release and desktop
pipelines.
What CI does NOT cover
Worth knowing, because these are the gaps where a local run is the only signal:
- The vendored re2c lexer. No workflow installs re2c, so nothing verifies
that the committed lexer matches
lexer.re.just verify-vendored-lexeris the only check, and it must be run by hand. - The observation snapshot (
spec/observations/example-diagnostics.json) records, for every spec example, the codes each stage emitted and whether the parsed model serializes back byte-exact. It IS in CI, through its currency test, but the gap is human: a regenerated snapshot with a changed entry passes the test, so every diff in it must be adjudicated in the commit as intended or unintended rather than committed becausejust regenproduced it. - A consumer’s behaviour after regenerating a generated module. A differential over generated TEXT is blind to a change in behaviour precisely when the text is expected to change; only running the consumer’s own suite sees it.
Legacy labels
References to numbered gates such as G0-G14 come from the predecessor
workspace and name nothing here. There is no Makefile in this repository.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Documentation Architecture
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
Principle: Centralized Book + Subsystem Satellites
User-facing and contributor-facing prose lives in mdBook
(book/). The repo-level docs/ directory holds operator-facing
material (release contract, versioning, code-signing, platform
support, validation feature flags). Maintainers can also generate a
local error-reference tree under docs/errors/ while working on
diagnostics, but that output is not the canonical checked-in docs
surface. Subsystem-specific working docs stay in place
only when tightly coupled to files in that directory.
flowchart TD
main["book/ (the unified Chatter mdBook)\nSurfaces: chatter, chat-format, architecture, contributing\nAudiences: users, integrators, contributors"]
spec["spec/docs/\nSpec authoring guides"]
errors["docs/errors/\nOptional local generated error reference"]
api["cargo doc\nRust API docs (auto-generated)"]
main -->|"links to"| spec
main -->|"links to"| errors
main -.->|"complements"| api
Where Documentation Goes
| Content type | Location | Examples |
|---|---|---|
| User guides, CHAT format reference | book/src/chatter/user-guide/, book/src/chat-format/ | CLI usage, validation errors |
| Architecture and design | book/src/architecture/ | Parsing, data model, concurrency, memory |
| Contributor workflows | book/src/contributing/ | Grammar workflow, testing, coding standards |
| Integrator contracts | book/src/chatter/integrating/ | JSON schema, diagnostic contract |
| Technical reference and audits | book/src/ (Technical Reference section) | Parity audits, UTF-8 audit, risk register |
| Spec authoring guides | spec/docs/ | Error spec format, curation workflow |
| Generated error docs | docs/errors/ | Registry artifact, written by just spec-gen and gated by just spec-check; source of truth stays in spec/errors/ |
| Historical/archived docs | project archive | Old audits, superseded proposals |
| AI assistant context | AGENTS.md files (per repo/subdir) | Not documentation for humans |
Rules
- One canonical page per topic. No duplicate coverage across locations.
- No crate-level
docs/directories. Architectural explanations go in the book. Crate API docs come from///doc comments viacargo doc. - Satellites stay only when the audience is editing files in that directory.
Spec authors need
WRITING_ERROR_SPECS.mdnext to their specs. Everyone else reads the book. - Generated docs are build artifacts. Never hand-edit
docs/errors/;just spec-checkreports a hand-written file there asextraand fails. Regenerate withjust spec-gen. - Historical docs go to project archive. Don’t keep old audit logs, investigation notes, or superseded proposals in the public repo.
Publication dates and content review
SUMMARY.md uses an HTML <a href="…">Git history</a> link to its own
history; mdBook would interpret a Markdown link there as a chapter entry.
Last modified is publication metadata, not a certificate of content review.
Book chapters use **Last modified:** 2026-10-02 (commit [2d7e886b](https://github.com/TalkBank/chatter/commit/2d7e886b)); the configured
Git-date preprocessor renders the date and commit from that chapter’s history.
Documents outside the rendered book, including SUMMARY.md, use a header
linked to their own Git history, for example
**Last modified:** [Git history](https://github.com/TalkBank/chatter/commits/main/CONTRIBUTING.md).
Both forms follow ordinary edits and content-preserving squashes without a
manual date sweep. Content review remains part of change review.
The date check admits the metadata header itself. A placeholder mentioned in
the body, a non-book rendering placeholder, or a history link to another file
does not exempt a document. Handwritten dates remain supported and checked
against actual and prospective commit dates; update them from real date
output when editing. Existing known-stale handwritten dates remain in the
ratchet until their pages are reviewed and corrected.
One unified book
There is one mdBook for this repo at book/,
titled “Chatter, TalkBank CHAT Toolchain”, organized by audience-first sections
under book/src/:
| Section | Audience | Content |
|---|---|---|
book/src/chatter/ | chatter CLI users + integrators | CLI reference, library usage, JSON contracts |
book/src/chat-format/ | All users + integrators | CHAT format reference (headers, tiers, symbols) |
book/src/architecture/ | All devs | Cross-surface architecture, parser/grammar/data-model design |
book/src/contributing/ | Contributors | Setup, testing, coding standards, dev checks |
One book.toml and one SUMMARY.md for the whole tree. Cross-section
links resolve as ordinary in-book paths.
Diagram Authoring Rules (canonical)
Architecture and design documentation MUST include Mermaid
diagrams. GitHub renders Mermaid natively; all mdBook builds have
mdbook-mermaid enabled.
When to Create a Diagram
Add a diagram when documenting:
- Data flow pipelines (how data transforms through stages)
- Architecture boundaries (what owns what, who calls whom)
- State machines and lifecycles (valid transitions, terminal states)
- Decision trees (option routing, fallback paths)
- Type relationships (trait hierarchies, enum variants, ownership)
- Protocols (request/response sequences, IPC message flows)
If a page describes a pipeline, boundary, or decision flow in prose without a diagram, the page is incomplete.
Diagram Type Selection
| Situation | Use | Not |
|---|---|---|
| Data flows through stages | flowchart TD or flowchart LR | sequenceDiagram (no named participants) |
| Request/response between components | sequenceDiagram | flowchart (hides back-and-forth) |
| Type hierarchies, trait impls | classDiagram | flowchart (wrong semantics) |
| State transitions, lifecycles | stateDiagram-v2 | flowchart (no state semantics) |
| Decision trees, option routing | flowchart TD with diamond nodes | Text lists (hard to follow branches) |
The Seven Diagram Rules
These rules exist because a successor who has never met the team will read these diagrams to understand the system. Every rule directly addresses a documented failure mode that produces misleading diagrams.
- Name every resource. Every node must have a specific name
AND its type/role. Not
"Cache", use"SQLite cache\n(talkbank-cache crate)". A reader must be able to grep the codebase for the node label and find it. - One concept per diagram. Each diagram tells one coherent story. When in doubt, split.
- No conveyor belts for interactive flows. If two components
exchange messages (request/response, IPC, HTTP), use
sequenceDiagram. Reserveflowchartfor genuinely one-directional data pipelines. - Show real decision points. Decision diamonds must use real
function names, flag names, and condition expressions, not
"check condition". - Include error and fallback paths. Every decision node must
show what happens on failure. Mark optional paths with dashed
lines (
-.->). - Anchor to source locations. Architecture diagram nodes should include the crate, module, or file path in the label or in prose immediately below.
- Never generate diagrams from source code without verification. Read the actual source files for every entity in the diagram; verify every node corresponds to a real module, function, or type; if you cannot verify a connection, omit it, gaps are better than lies.
Formatting Standards
- Node labels:
["Name\n(role or path)"]for multi-line - Decision nodes:
{"condition?\ndetail"}diamond syntax - Edge labels:
-->|"label"| targetfor all non-trivial edges - Colors/styles: Do not use custom colors. Default Mermaid themes ensure consistent rendering across GitHub and mdBook
- Size limit: Keep diagrams under about 30 nodes. If larger, split into focused diagrams.
- Angle bracket escaping: Raw angle brackets in Mermaid labels
(
Arc<str>,Cow<str>,&str) trigger mdBook “unclosed HTML tag” warnings. Escape as<str>inside labels.
Placement
- Place each diagram inline, immediately after the prose paragraph that introduces the concept it illustrates.
- Every diagram must have a prose introduction explaining what it shows and why the reader should care.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
CHAT Processing Playbook for Developers
Status: Current Last updated: 2026-09-28 20:59 EDT
Objective
Provide an implementation playbook for developers building or extending CHAT parsing, validation, transformation, and serialization logic.
Mental Model
Treat CHAT processing as a layered pipeline:
- Ingest bytes and normalize line boundaries.
- Parse syntax into structured model with exact spans.
- Validate semantic rules with structured diagnostics.
- Transform or enrich model without breaking invariants.
- Serialize in canonical form.
Developer Workflow
- Start from a concrete fixture or corpus case.
- Add/adjust parser behavior with contract tests first.
- Add semantic validator rules separately from parser acceptance.
- Confirm roundtrip and equivalence gates.
- Update docs for any visible behavior or policy change.
Tier Dispatch Strategy
Use cheap byte-prefix dispatch before heavy parsing:
@=> header candidate,*=> main tier,%=> dependent tier,- continuation rules and whitespace handled deterministically.
This preserves performance and isolates error contexts earlier.
For downstream batchalign3 consumers, tier dispatch is only the front door.
The important contract is what happens after dispatch: parse-health taint,
recovery vs rejection, and whether a tier is safe to pass into alignment.
Word Parsing Rules of Thumb
- Parse suffix markers in strict order (
@...,@s...,$...) with explicit precedence. - Derive
raw_textfrom typed structure; use the original source and its spans when exact input bytes are required. Keepcleaned_textpolicy-driven and test-locked. See the JSON contract. - Treat CA delimiters and special symbols via centralized symbol sets.
- Never embed ad hoc symbol literals in multiple files.
Error Handling Contract
- Every parser failure should produce structured diagnostics with:
- code,
- severity,
- span,
- context,
- message.
- Avoid silent fallback behavior unless policy explicitly allows it.
- If fallback occurs, emit warning-grade diagnostics where relevant.
- Never fabricate semantic placeholders (empty required text, arbitrary enum default, fake word/chunk) to satisfy type construction.
- Prefer
None/partial outcome + diagnostics over synthetic model values.
Span Discipline
- Offsets are absolute across full file content.
- Nested parser helpers must accept base offset and return shifted spans.
- Add tests for boundary and continuation-line spans.
Performance Policy
- Prefer byte-oriented prechecks for top-level dispatch and simple delimiters.
- Use parser combinators for structural parsing, not for obvious constant-prefix routing.
- Measure parser performance on representative corpus slices before/after major changes.
Common Failure Patterns and Fixes
- Symptom: semantic mismatch only in snapshots.
- Fix: compare parser outputs directly and isolate first structural delta.
- Symptom: generated tests pass, corpus fails.
- Fix: add missing fixture, decide parse-vs-validate placement, lock behavior.
- Symptom: output drift after grammar edit.
- Fix: run full regeneration and equivalent parser contract suite before merge.
Batchalign3 Surface Checks
When a change affects the surface used by batchalign3, confirm:
- full-file parse equivalence still holds for corpus coverage
- alignment-sensitive downstream tiers still gate on parse-health appropriately
Review Checklist for Parser PRs
- New or changed behavior has targeted tests.
- Equivalence suite status is attached.
- Snapshot updates are intentional and explained.
- No hidden magic symbols or magic string literals introduced.
- Docs updated where user-visible behavior changes.
Required Artifacts for Significant Changes
- Design note (architecture decision record in the book).
- Before/after examples.
- Impacted fixtures list.
- Migration implications for integrators.
This page last changed: 2026-09-28 (commit 2cb42a45). The whole book last changed: 2026-10-07 (commit 5e895791).
GitHub Readiness and Open Source Governance
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
Objective
Prepare TalkBank/chatter to operate as a healthy public project with clear legal, security,
contribution, and release processes.
Root Artifacts
| Artifact | Status | Notes |
|---|---|---|
LICENSE-MIT + LICENSE-APACHE | Done | Dual-licensed MIT OR Apache-2.0 (standard Rust convention; both files present at root, no combined LICENSE). Every crate inherits license = "MIT OR Apache-2.0" from [workspace.package]. |
CONTRIBUTING.md | Done | Setup, standards, PR flow, pre-PR checklist |
CODE_OF_CONDUCT.md | TODO (deferred) | Intentionally absent for now: it is held until a durable enforcement contact (an institutional address or successor handle, not an individual) is settled. The plan is to adopt the Contributor Covenant once that contact exists. |
SECURITY.md | Done | Root file present; the issue-template contact link resolves to a real policy |
CODEOWNERS | TODO | Not added yet: repo contents do not currently publish an authoritative GitHub owner/team map for path-level review ownership |
.github/workflows/*.yml | Done | ci.yml (Rust build+test, mdBook) + cross-platform.yml (OS matrix) + release-lint.yml (clippy + feature-off, on a tag) + clippy-rolling.yml + crates-io-foundation.yml + release.yml + release-desktop.yml |
.github/ISSUE_TEMPLATE/* | Done | Bug report + feature request (YAML forms) |
| Pull request template | Done | .github/PULL_REQUEST_TEMPLATE.md mirrors current CONTRIBUTING + PR review requirements |
CI Governance Policy
- Required status checks: the
ci.ymljobs that run on every pull request,Rust build + test,mdBook build, andRust version pins in sync. See Branch Protection for the exact GitHub check names and which other workflow (cross-platform.yml) is deliberately not in the required set. - Branch protection rules: documented in Branch Protection; configure on GitHub once the repo is public.
Release Governance
- Releases: the CLI and desktop app are published as signed GitHub Releases (cargo-dist); the Rust crates are source-available (not yet on crates.io).
- Cargo publication governance: first-wave crates.io foundations are documented
in Crates.io Publication and checked by
.github/workflows/crates-io-foundation.yml. - Binary release governance:
release.ymlis reserved for cargo-dist GitHub Release packaging of dist-enabled artifacts. It is not the crates.io publication workflow. - Tagging rule: do not treat version tags as authorization to publish new surfaces. A surface becomes stable only when its release notes explicitly say so and its public distribution channel is live.
- Release-note rule: every public release note must state the surface’s distribution channel, support boundary, and any closely related surfaces that remain held back.
Community Operations
- Label taxonomy:
bugandenhancementauto-applied by issue templates. Richer taxonomy (drift,spec,grammar,parser,docs,good first issue): TODO (GitHub settings). - Contributor pathway:
CONTRIBUTING.mdcovers setup and PR flow. First-time/advanced contributor pathways: TODO. - Public project roadmap: TODO.
Supply Chain and Security
- Dependency scanning: CI runs
rustsec/audit-checkandcargo-deny(withdeny.toml). Automated update PRs (Dependabot/Renovate): TODO. - Signed release artifacts: TODO.
- Security advisories process: documented in
SECURITY.md.
Acceptance Criteria
- Repo has complete governance artifacts at root.
- CI and branch protections enforce stated policy.
- Contributors can onboard and submit PRs without tribal knowledge.
- Release/support tiers are documented per surface.
- Release process is repeatable and documented.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Rust Compilation Times
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
How the workspace’s dev and test profile settings keep compilation fast, and
what to avoid. The workspace root Cargo.toml comments are the source of truth
for the exact settings; re-run cargo build --timings for current numbers.
Background: How Rust Compilation Works
Rust compilation has two key mechanisms for speed:
-
Incremental compilation: When you change one file and rebuild, the compiler remembers which “codegen units” within each crate were affected and only recompiles those. This is the primary speedup mechanism for local iterative development (edit-compile-test cycles).
-
Crate-level caching: Cargo tracks which crates have changed inputs (source files, dependencies, feature flags). Unchanged crates are skipped entirely. This helps when you edit a leaf crate and don’t need to rebuild unrelated crates.
Additionally, there are external tools:
-
sccache: A shared compilation cache that stores compiled artifacts by content hash. Designed for CI environments where builds start from a clean state. It works by wrapping
rustcand checking a cache before invoking the real compiler. -
Linker choice: The linker runs after all crates are compiled to produce the final binary. Faster linkers (like
lld) can shave seconds off link time for large binaries.
Settings this workspace uses
debug = "line-tables-only"in[profile.dev]and[profile.test]. Backtraces keep file and line information, while the bulky type and variable metadata (full DWARF, large.dSYMbundles and.ofiles that inflate linker input) is skipped. You cannot inspect local variables in a debugger (lldb/gdb); for most development workflows this is the right tradeoff.split-debuginfo = "off"in the same profiles. macOS defaults tounpacked, which leaves one.rcgu.oper codegen unit intarget/debug/depsand makes warm test runs pay for scanning tens of thousands of directory entries.opt-level = 3for build scripts and proc macros ([profile.dev.build-override]).- No workspace-wide third-party optimization.
[profile.dev.package."*"]and[profile.test.package."*"]withopt-level = 1are not set: with the workspace’s third-party dependency surface (axum, async-trait, tokio’s full feature set, and so on) their build-time cost is prohibitive. Where runtime is the bottleneck for a specific test, opt in locally rather than setting it workspace-wide.
Do not let a compiler wrapper disable incremental compilation
A global ~/.cargo/config.toml that sets rustc-wrapper (for example to
sccache) disables Rust incremental compilation entirely, because the wrapper
interposes between Cargo and rustc and breaks the incremental artifact
protocol. sccache also gives near-zero benefit for this workspace: rlib
crates, which most workspace crates produce, cannot be cached by sccache. The
result is that every build after a one-line change is a full rebuild of the
dependency chain; a change to talkbank-model, near the root of the crate
graph, recompiles 11+ downstream crates.
If your global config sets a wrapper, override it for this project only with a
local .cargo/config.toml:
[build]
rustc-wrapper = ""
Other Rust projects on the system are unaffected, and sccache stays available
for CI and other projects. The file is gitignored (the repository’s
.gitignore names /.cargo/config.toml) rather than committed because
the empty-string value trips a cargo-llvm-cov bug that treats "" as a real
wrapper path instead of “no wrapper”; each contributor opts in locally, and CI
does not carry the override.
The lld linker (linker = "lld" in the global config, ld64.lld from
Homebrew’s LLVM on macOS) is fine and slightly faster than Apple’s default
linker for a workspace of this size.
Optional: Cranelift Backend for Maximum Iteration Speed
For the fastest possible “does it compile?” checks during rapid iteration, Rust nightly supports the Cranelift codegen backend:
cargo +nightly -Z codegen-backend=cranelift build
Cranelift generates code ~2x faster than LLVM but produces unoptimized output and is nightly-only. It is useful for compile-check cycles but not for correctness testing or benchmarking.
General Principles for Rust Compile Time
-
Incremental compilation is king for local dev. Anything that disables it (sccache, certain rustc-wrapper tools) is a net negative for iterative development.
-
sccache is for CI, not local dev. It shines when doing clean builds from scratch (CI runners, cross-compilation). For edit-rebuild cycles, incremental compilation is far more valuable.
-
Optimize dependencies, not your own crates, where the dependency surface is small.
[profile.dev.package."*"]withopt-level = 1speeds test execution at little compile cost when dependencies rarely change, but its build-time cost grows with the dependency set. This workspace does not set it (or theprofile.testequivalent); where runtime is the bottleneck for a specific test, opt in locally. -
Debug info has a real cost. Full DWARF debug info inflates binary sizes and link times. Use
line-tables-onlyunless you actively need a debugger. -
Measure before optimizing. Use
cargo build --timingsto generate an HTML report showing per-crate compile times and parallelism. Usesccache --show-statsto verify cache effectiveness. -
Watch for crate graph bottlenecks. Crates that sit at the root of the dependency graph (like
talkbank-model) are the critical path, changes to them trigger the longest rebuild chains. Keep these crates lean and consider splitting them if they grow too large.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Developer Verification Checks
Status: Current Last modified: 2026-10-02 (commit b4ad11fc)
What to run locally, and what each thing costs. The commands are just
recipes; just --list shows them all.
The inner loop
just test # workspace test binaries, complete failure inventory
Narrower is better while iterating. Prefer the smallest thing that can fail:
cargo test -p <crate> --tests <name filter>
cd grammar && tree-sitter test # grammar-only edits
Do not run cargo check before cargo test. cargo test type-checks
everything check would, and the two are DIFFERENT cargo units: check emits
only .rmeta while test emits full .rlib with codegen, so nothing is
reused and alternating them recompiles the whole dependency graph twice. If a crate has no tests, run
cargo test -p <crate> anyway; it compiles and reports zero tests.
Deterministic tests and native-platform feedback
Express admission and lifecycle invariants through typestate, ownership and validated constructors first. Use pure tests for structural and policy logic, and deterministic doubles for clocks, services and scheduling. Keep spec and reference-corpus contracts at their real public boundaries. A double cannot establish a database driver’s worker shutdown or an operating system’s file locking behavior; retain thin, isolated native-platform tests for those facts.
Use explicit barriers and acknowledgments rather than sleeps or scheduler luck. A deadline bounds execution; it is not the synchronization condition. Consume resource capabilities at lifecycle transitions, and await actual cleanup where the external API requires it. A passing retry does not resolve an observed failure.
The shared workspace recipe and native CI use --no-fail-fast: all test
binaries run, failures still fail the command, and the result is one repair
inventory rather than successive first-failure discoveries. This flag does not
retry tests or weaken assertions. CI and the gate share just test-workspace,
which requires prerequisites rather than permitting inner-loop skips.
A local gate verifies its own platform, not
every supported operating system; use authorized incremental push/CI
checkpoints before freezing a release candidate instead of treating release
publication as the first cross-platform test.
The git hooks, and what each refuses
Run just install-hooks once per clone. It points core.hooksPath at the
tracked .githooks/ directory, because .git/hooks does not survive a clone
and an untracked hook is a gate that exists on exactly one machine.
| Hook | Refuses |
|---|---|
pre-commit | invalid publication metadata or stale staged handwritten dates, then chains to the optional local hook |
commit-msg | a type(scope)!: subject that does not touch CHANGELOG.md; and production Rust staged with no test, spec, corpus or fixture beside it |
pre-push | a push with no just gate stamp, or a stamp taken on different bytes |
These hooks have no bypass flag, and pre-push runs no checks of its own: it reads
the stamp just gate writes, because git has already opened its connection to
the remote by the time a pre-push hook runs, so a multi-minute hook is closed
by the SSH idle timeout and fails a push that had passed.
just doc-dates admits each document’s publication metadata. Git-derived
headers follow the document’s own history automatically; see
documentation architecture. Remaining
handwritten dates are checked against actual history and, for pending changes
since the configured upstream, today’s date. The commit hook reads the Git
index, so an unstaged correction cannot conceal a stale staged header.
Detached CI checks actual committed history. Neither Git publication metadata
nor a handwritten date certifies a content review.
The red-evidence gate has one way past it, and it is not a flag. If a change
genuinely admits neither a test nor a type, say so in a Red: trailer on its
own line in the message body, naming what was red:
Red: the compiler, at 14 call sites of Word::new
Red: nothing. A pure deletion; it removes the only caller of X.
That trailer is recorded in the history and names a claim a reader can check,
which a bypass variable is not. In this repo a spec file counts as the
failing test: a construct or parser bug is fixed by writing the spec first,
and just regen turns it into fixtures.
just evidence-gate-test and just breaking-changelog-test prove both gates
fire, in both directions; both run in just gate.
Before pushing
just gate # static checks plus every test CI runs; the pre-push gate
Or just push, which runs gate and then pushes.
just release-lint is separate and is NOT part of this: clippy over both
workspaces plus the feature-off build, run once before a release. Each is its
own cargo unit that recompiles the workspace, and none of them is a thing a
per-push gate needs to know.
gate puts every cheap check ahead of every expensive one, so a workflow typo
or a stale version pin fails in seconds rather than after the test suite.
Do not assemble this by hand from the list below. A green just test is
not a green gate: just test is --tests, and doctests are a separate
compilation it cannot see.
What gate runs, and why each is not covered by the others:
| Step | Catches what nothing else does |
|---|---|
just fmt-check | cargo test does not run rustfmt; CI does |
just grammar-generate-check | a stale parser.c. The traversal staleness guard hashes grammar.json and node-types.json, so a regeneration touching only parser.c passes it correctly; a tree-sitter version bump does exactly that |
just test | the compiled test suite |
cargo test --doc --workspace | doctests, invisible to --tests |
just test-spec | the spec/ workspace, which --workspace does not reach |
just book | the book builds and its links resolve |
just doc-dates | a Last modified header older than the file |
just actionlint, the two sync checks | workflow syntax and version pins |
Clippy is deliberately absent, and so is the feature-off build: both are
just release-lint, which per-push CI does not run either. Nothing in CI
goes red on something the local gate did not run; that equivalence is what
scripts/check_ci_gate_sync.py enforces.
just test-all is the TEST half of the gate (both workspaces, doctests, the
proc-macro UI suite) and is what gate delegates to. Useful on its own when you
want the tests without the lints, the grammar checks and the book.
just fmt-check is not optional. cargo test does not run rustfmt, CI
does, and formatting drift accumulated across 19 files once while every test run
stayed green.
By surface
Parser, model, alignment, serialization, roundtrip (mandatory):
cargo test -p talkbank-parser-tests --tests reference_corpus_parses
cargo test -p talkbank-parser-tests --tests roundtrip_reference_corpus
cargo test -p talkbank-parser-tests --tests gates
Grammar. Follow the full Grammar Workflow;
tree-sitter test does NOT detect a stale parser.c, so regeneration is
mandatory before any parser behaviour can be trusted.
Specs, or either registry:
just spec-status # derived state: statuses, verified/deferred, parity counts
just test-spec # the gates: example codes, manifest, registry drift
The re2c lexer. After changing lexer.re or the generated form-marker code
set it includes, install the exact re2rust version named in
re2c-version.toml, then run:
just verify-vendored-lexer
The recipe fails before regeneration if the installed generator version differs from that source of truth, and then compares the generated bytes. Nothing else checks it: no CI workflow installs re2c, so this is the only check that exists, and it takes under a second.
Docs:
just doc-dates # a `Last modified` header older than the file fails
Dependency updates
Review Rust and JavaScript desktop dependency changes together with both lockfiles. For schema-validation or compiler-helper updates, exercise the reference corpus and JSON contracts:
cargo test --locked -p talkbank-parser-tests --test integration -- transform_corpus::json_contracts reference_corpus_parses
For desktop dependency updates, run npm ci, npm run test:unit, and
npm run build from apps/chatter-desktop, plus the native bridge tests
from the workspace root:
cargo test --locked -p chatter-desktop --test validation_bridge
These are focused compatibility checks, not release acceptance: they do not certify signed installers, updater delivery, or native behavior on every target platform. A manifest-only update does not require grammar regeneration.
Regeneration
Run a generator only when its inputs changed, and never edit its output:
just symbols-gen # spec/symbols/symbol_registry.json
just form-markers-gen # spec/form_markers/form_marker_registry.json
The spec-driven generators (tree-sitter corpus, Rust tests, validation corpus) are in Spec Workflow, with every command written out.
Regeneration is not a substitute for choosing the right regression test.
Failure policy
For CLI subprocess failures, retain the full exit status, stdout and stderr. Check successful completion before interpreting cache counts or other output: an empty stream alone cannot distinguish a product failure from a terminated process. Reproduce the exact failing test before broadening the run.
On Unix, tests that vary the program name should use CommandExt::arg0 on
the original executable. This avoids giving the shared test executable a
second filesystem name during concurrent launches. Windows uses a temporary
same-filesystem hard link because its command API has no arg0 override.
A failing check blocks the change. If a failure is unrelated and pre-existing, verify that by running against a clean checkout, say so, and fix it anyway rather than routing around it: pre-existing defects linger precisely because each person who meets them decides they belong to somebody else.
This page last changed: 2026-10-02 (commit b4ad11fc). The whole book last changed: 2026-10-07 (commit 5e895791).
Branch Protection and Required CI Checks
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
This page defines the required status checks and protection policy for main.
Branch Protection Policy
Enable branch protection for main with:
- Require pull request before merge.
- Require approvals (minimum 1; maintainers may set higher).
- Require conversation resolution before merge.
- Require status checks to pass before merge.
- Restrict force pushes and branch deletions.
Required Status Checks
Configure these CI checks as required. The names are the GitHub check names,
which come from each job’s name: in .github/workflows/ci.yml; that
workflow runs on every pull request to main. This is every job that
workflow defines, so the required set and the workflow do not drift apart:
Rust build + testwasm32 check (model + re2c parser)mdBook buildRust version pins in syncApp version in syncShell scripts (shellcheck, strictest)Grammar (generate staleness, tree-sitter test, queries)Dependency policy (cargo-deny)
A required-check list that silently omits jobs is worse than no
list, because it reads as a deliberate selection rather than an oversight.
When you add a job to ci.yml, add it here in the same commit.
Note that this page states the INTENDED required set; the live setting lives in the repository’s branch-protection configuration on GitHub and is changed there by a maintainer, not by editing this file.
One other workflow is deliberately NOT in the required set:
cross-platform.yml(the Ubuntu + macOS + Windows matrix) runs on push tomain, a daily schedule, and manual dispatch, NOT on pull requests, so it cannot report a status on a PR and must not be required (requiring it would block every merge). It is a post-merge and daily drift gate. Add apull_requesttrigger first if you want it required.
Optional Hardening
- Require branches to be up to date before merging.
- Enable merge queue if PR volume increases.
- Restrict who can dismiss stale reviews.
Operational Rule
If required checks fail:
- Do not bypass protection.
- Fix the issue or revert the breaking change.
- Re-run checks until green.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Reference Corpus
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
The reference corpus (corpus/reference/) is the finite, reusable input suite
that the parser, roundtrip and transformation gates run against. The layout,
the provenance and licensing rules and the validation commands are in
corpus/README.md;
the test harness discovers the current population, so no count is kept here.
This page describes the tools that select, extend and measure the corpus.
What the corpus must cover
Every concrete grammar node type must be exercised by at least one reference
file, so a grammar regression in any construct fails a gate. Constructs that do
not occur in real-world data (rare terminators, uptake_symbol,
scoped_best_guess, the unsupported_* nodes and thumbnail_header) are
covered by small handcrafted files in the corpus rather than left uncovered.
The language files are real conversations morphotagged with %mor/%gra, one
per language, drawn from the corpus data.
Tools
| Tool | Path | Purpose |
|---|---|---|
corpus_node_coverage | spec/tools/src/bin/ | Reports which concrete grammar node types the corpus exercises |
extract_corpus_candidates | spec/runtime-tools/src/bin/ | Scores and ranks candidate files in a corpus data directory for target languages |
perturb_corpus | spec/tools/src/bin/ | Produces error files by mutating a valid .cha file, and mines real data for tree-sitter ERROR nodes |
Selecting language files
extract_corpus_candidates ranks files per target language. Its criteria:
- the file parses cleanly with tree-sitter (no ERROR nodes), which is mandatory;
- short files (a
--max-linesbound, default 200, preferring 15-100 lines); - varied tiers (
%mor,%gra,%pho,%com); - multiple speakers preferred;
Passworddirectories are skipped, for privacy.
Fresh %mor/%gra tiers for a selected file come from batchalign3
morphotag run in place over its language directory.
Producing error files
perturb_corpus --list prints the mutation strategies; each takes a valid
.cha file and applies one controlled mutation:
| Perturbation | Mutation |
|---|---|
delete-participants | Delete the @Participants header |
delete-languages | Delete the @Languages header |
delete-id | Delete all @ID headers |
undeclared-speaker | Change a speaker code to an undeclared XXX |
delete-terminator | Remove the terminator from the first utterance |
extra-mor-word | Add an extra word to the first %mor tier |
fewer-mor-words | Remove a word from the first %mor tier |
delete-begin | Delete the @Begin header |
delete-end | Delete the @End header |
duplicate-participants | Duplicate the @Participants header |
mor-terminator-mismatch | Change the %mor terminator to differ from the main tier |
--mine DIR scans a data directory for tree-sitter ERROR nodes, again
excluding Password directories.
Design notes
- Perturbation beats mining for systematic coverage. Well-curated corpora contain almost no tree-sitter parse errors, and mining is slow on large directories, so controlled mutation is the way to reach a specific error code.
- Parser recovery codes are hard to trigger. Tree-sitter’s error recovery
routes most malformed input through the generic path (E316) rather than the
specific recovery codes (E319-E322, E376), so examples for those codes
usually cannot reach them; their specs are
not_implementedand say why. - Some codes have no emission path (internal or reserved codes such as E001 and E002), so their specs document why no example is possible.
- Adding files is purely additive. New reference files extend the gate without disturbing existing ones.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Desktop App Testing
Status: Current Last updated: 2026-10-07 (commit 5e895791)
This document covers the testing strategy for the Chatter desktop app
(apps/chatter-desktop/). Testing is split into three tiers by speed and scope.
Testing Tiers
┌─────────────────────────────────────────────────────────┐
│ Tier 3: E2E (WebdriverIO + tauri-driver) │
│ Real app, real DOM, real IPC. Slow (~5-10s/test). │
│ Catches: rendering bugs, IPC wiring, platform quirks. │
│ Run: manually before releases, optionally in CI. │
├─────────────────────────────────────────────────────────┤
│ Tier 2: Rust integration tests │
│ Real validation pipeline, real event bridge, no GUI. │
│ Catches: serialization mismatches, event ordering, │
│ stats consistency, single-file handling. │
│ Run: every commit, CI required. │
├─────────────────────────────────────────────────────────┤
│ Tier 1: Unit tests (Rust + TypeScript) │
│ Pure functions and thin runtime seams in isolation. │
│ Catches: protocol drift, reducer bugs, CLAN math. │
│ Run: every commit, CI required. │
└─────────────────────────────────────────────────────────┘
Tier 2 exercises the shared validation runner and event bridge, not the entire
native app. Command-context tests must enter tauri::async_runtime::block_on:
a plain-thread call cannot detect nested-runtime failures at the Tauri boundary.
Use isolated cache directories, never the real user cache. Passing these tests
does not establish that native menus, dialogs, drag/drop, signing or installation
work; those require their respective runtime or release checks.
Release-facing behavioral contracts
The desktop links the same validation engine as the CLI. Keep CHAT semantics in the shared Rust implementation; use canonical specs/reference files for desktop boundary tests rather than inventing a second fixture oracle.
File outcomes and completion
validationState.ts::fileOutcome projects a streamed FileEntry into a
discriminated union: pending, valid, or problem with an explanation. Both the
file tree and the detail panel use it, and fileStatusLabel words it for
the detail panel and the text export. In particular, readError,
internalFailure, roundtripFailed and cached invalid statuses must remain
visible when there is no diagnostic array. Do not infer success from an empty
array.
shouldShowAllFilesValid requires a finished phase whose passed is
true (the runner’s own verdict, the one the CLI’s exit status reads) and no
file with anything to show. A cancelled run is the separate stopped phase,
and a target with no transcript is nothingFound, so neither can reach this
check. An internal failure increments internalFailures, not the valid or
invalid counter: CHAT validity remains undetermined.
finishedRunSummary is shared by the window title,
notification and status bar. Regression tests exercise failure statuses without
diagnostics, warning visibility, cancellation, empty populations and success.
When adding a new wire status, update the exhaustive projection, not a second
component-local classification.
Update lifecycle
updates.ts owns an idle/checking state. A checking state carries the single
in-flight result and feedback mode. Launch, six-hour background and manual menu
requests join that operation; a manual join promotes it to visible feedback.
There can be only one prompt/install sequence for overlapping requests. Completion
or failure returns the capability to idle so a later retry can start.
The update seam tests cover accepted/declined installs, no update, network failure, overlapping requests, a manual join, retry, and native error-dialog failure. The latter must not reject a fire-and-forget menu callback. These tests use transport doubles: they do not verify a real signature, actual installation, or relaunch. Installer and updater artifacts require separate release evidence.
Stored Unicode identity
Reveal-in-file-manager travels through the validation-target capability and the typed command protocol, not a component-local Tauri import. Transport-double tests retain the exact Unicode path and propagate native failures. They do not launch or certify the operating system’s file manager.
The runtime bridge test copies the canonical W109 media specimen into an owned temporary directory, then validates it through stored NFD, NFC alias (where the filesystem supports it), and directory targets from inside Tauri’s runtime. It requires the same both-side W109, no E531 mismatch, and nonempty HTML/plain renderings. Production identity comes from the shared stored-name resolver; the desktop must not independently normalize paths or downgrade failed name resolution to anonymous validation.
Export admission and failure preservation
The async command tests exercise text export with each file status, including failures without diagnostic cards and pending results. Text export deserializes typed file records and the same status enum used by the event bridge; it no longer navigates arbitrary JSON with question-mark fallbacks. Required paths and rendered diagnostic text must be present. An unknown status or malformed record refuses before filesystem writing, preserving an existing report.
Status and diagnostics are separate facts: a roundtrip failure can coexist with diagnostic text, and neither may suppress the other. Tests preserve rendered text verbatim. JSON export retains its existing wire representation. Neither per-file export is a whole-run coverage certificate; cancellation/session metadata is not presently included, as documented in the user guide.
Before calling a desktop candidate ready
Run the focused frontend seam tests, compile the production frontend, and run the Rust bridge tests on the final sources. Inspect the actual app on supported platforms: open a canonical valid file, a diagnostic fixture and a failing target; verify cancellation/revalidation, copy/export, menus and update feedback. Do not call the app ready solely because TypeScript compiled or the headless bridge passed. Use the coordinated release process for signed installers and updater manifests, with versions tied to the exact release commit.
Tier 1 & 2: Unit and integration tests
Candidate evidence matrix
An implemented test is not a passing receipt, and a headless receipt is not native interaction evidence. Record candidate identity, platform, actual result and remaining gaps in the release record; do not copy historical success into a new candidate’s acceptance.
| Workflow | Existing automated boundary | Required native observation |
|---|---|---|
| Valid and invalid files | Reference corpus and canonical spec bridge tests | Select each; verify visible verdict, diagnostic navigation and completion |
| Unreadable or unsupported target | Typed outcome projection and path-contract tests | Surface failure visibly; do not show an all-valid result |
| Directory selection | Nested discovery and event/statistics bridge tests | Select a directory; inspect relative paths and final totals |
| Unicode paths | Stored-identity W109 runtime bridge test | Open an accented path through the native chooser and inspect media warnings |
| Cancel and revalidate | Runner lifecycle and cancelled-state seam tests | Cancel an active run, change input, revalidate; no stale result may claim success |
| Copy and export | Typed export admission and refusal tests | Check clipboard and saved output, including failure without diagnostic cards |
| Menus and updates | Single-flight update seam tests | Exercise menu feedback; test signed installation/update separately |
| Settings preservation | No complete native proof supplied by bridge tests | Verify preferences survive relaunch and upgrade without touching unrelated settings |
Declare platform support and installation/update acceptance explicitly. A configured CI target or available installer is not evidence that every workflow above has been exercised on that platform.
Running
# TypeScript capability/seam tests
cd apps/chatter-desktop && npm run test:unit
# Rust contract/integration tests
cargo test -p chatter-desktop --test validation_bridge
What they cover
| Test | What it verifies |
|---|---|
apps/chatter-desktop/tests/unit/validationRunner.test.cjs | Validation capability uses centralized command names, subscribes before invoke, and disposes listeners exactly once |
apps/chatter-desktop/tests/unit/validationState.test.cjs | Validation reducer computes relative file names and merges diagnostics/status immutably |
reference_corpus_no_hard_errors | every file under corpus/reference/ produces zero Severity::Error (warnings allowed) |
event_lifecycle_has_correct_sequence | Discovering → Started → FileComplete×N → Finished ordering |
frontend_events_serialize_to_expected_json_shape | Every event has type field; camelCase field names match TypeScript types; diagnostics include renderedText |
protocol_contracts_serialize_to_expected_json_shape | Rust command/event constants and request payloads stay aligned with the TypeScript protocol module |
single_file_validation | Single-file path validates exactly the selected file |
finished_stats_match_file_events | Typed valid, invalid, parse-error and internal-failure counts account for the population; FileComplete count matches |
rendered_html_present_for_errors | Every diagnostic carries non-empty miette HTML with box-drawing characters and style= attributes (ANSI colors converted to HTML) |
Adding new tests
Test file: apps/chatter-desktop/src-tauri/tests/validation_bridge.rs
The tests use collect_events() which runs the real validation pipeline and
collects all FrontendEvent values. To test a specific scenario:
#![allow(unused)]
fn main() {
#[test]
fn my_scenario() {
let target = workspace_root().join("path/to/corpus");
let events = collect_events(&target);
let summary = summarize(&events);
// assert on summary fields or individual events
}
}
Miette rendering pipeline
Error rendering is server-side. Each FrontendDiagnostic carries two
renderings:
rendered_html:render_error_with_miette_with_source_colored()produces ANSI-colored text,ansi-to-htmlconverts it to HTML<span style="...">. The frontend displays it in a<pre>block viadangerouslySetInnerHTML. This guarantees identical output to the CLI.rendered_text:render_error_with_miette_with_source()produces plain text (no ANSI codes) for clean clipboard copy-paste.
The rendered_html_present_for_errors integration test verifies that every
error diagnostic includes non-empty HTML containing miette box-drawing
characters and style= attributes from ANSI color conversion.
TypeScript seam tests
The TypeScript unit tests compile a focused subset of apps/chatter-desktop/src/ to a
temporary CommonJS directory, then run Node’s built-in test runner against the
compiled output. This keeps the test toolchain small while still exercising the
runtime seam as real JavaScript.
- Runner script:
apps/chatter-desktop/scripts/run-unit-tests.mjs - Compile config:
apps/chatter-desktop/tsconfig.unit.json - Test files:
apps/chatter-desktop/tests/unit/*.test.cjs
TypeScript ↔ Rust contract
The Rust integration tests verify that serialized JSON matches what the
TypeScript frontend expects. If you change a field name or event structure in
events.rs, the frontend_events_serialize_to_expected_json_shape test will
catch the mismatch before you discover it at runtime.
The key serde attributes:
#[serde(tag = "type", rename_all = "camelCase")]on enums, variant names become camelCase tag values (fileComplete, notFileComplete)#[serde(rename_all = "camelCase")]on individual variants, field names become camelCase (totalFiles, nottotal_files)- Both must be present: the enum-level
rename_allonly affects tag names, not field names within variants
Tier 3: E2E Tests (WebdriverIO)
Prerequisites
cargo install tauri-driver # WebDriver backend for Tauri (Linux/Windows only)
cargo tauri build --debug # Build the app binary
Note: the checked-in configuration drives standalone tauri-driver, which
supports Linux and Windows, not macOS. This is a limitation of that route,
not of every available Tauri automation option; see the evaluation below.
Running
# Terminal 1: start tauri-driver (WebDriver server on :4444)
tauri-driver
# Terminal 2: run the tests
cd apps/chatter-desktop
npm run test:e2e
What they cover
The smoke tests in tests/e2e/smoke.spec.ts verify that the app launches and
renders the expected UI elements:
- Drop zone with Choose File / Choose Folder buttons
- Empty file tree (“No files loaded”)
- Empty error panel (“Select a file to view errors”)
- Status bar showing “Ready”
Limitations
The native file picker is outside this suite’s WebDriver-controlled DOM. The checked-in smoke suite verifies launch UI only: it does not yet automate file selection, a validation run, cancellation, export, or native menus. The Rust integration tests cover the shared validation pipeline and bridge, not those native interactions.
Do not copy a direct window.__TAURI__.core.invoke("validate", { path })
example: this app does not enable the global Tauri API, and validate takes
one request argument (a ValidateRequest: path, roundtrip, parser kind,
strict linkers and jobs), not a bare path. Production calls
use the typed runtime capability and transport. A future automated native
validation test must use that contract, subscribe before starting the run,
await its terminal event with a bounded timeout, and isolate cache/settings.
A fixed sleep is not proof of completion. Do not add a production test-only
command or bypass admission just to drive a test.
Adding E2E tests
Test file: apps/chatter-desktop/tests/e2e/*.spec.ts
Follow the existing selector-based launch checks for layout assertions. New validation-flow coverage must meet the lifecycle and isolation requirements above; passing a DOM assertion after a delay is not a validation receipt.
When to run E2E tests
- Before releases: manual run to verify the built app works end-to-end
- Optionally in CI: requires
tauri-driverand a display server (Xvfb on Linux). Slow, so consider running only on release branches. - Not on every commit: the Rust integration tests are fast and cover more ground
Platform-Specific Considerations
| Platform | WebView engine | E2E support |
|---|---|---|
| macOS | WKWebView | Not supported by our current standalone-driver configuration; service-based alternatives exist |
| Windows | WebView2 (Chromium) | Full support via tauri-driver |
| Linux | WebKitGTK | Full support via tauri-driver; requires Xvfb for headless |
Current macOS coverage: use the Rust integration tests (Tier 2) and explicit native smoke evidence. Do not infer native coverage from browser or transport doubles. The service-based option below has not been adopted or verified here.
Automation follow-up: evaluation, not implementation
The Tauri WebDriver guide
documents @wdio/tauri-service with an embedded WebDriver plugin supporting
macOS as well as Linux and Windows. It also documents renderer-only browser mode
and external-driver alternatives. See the
WebdriverIO Tauri documentation.
These capabilities are upstream claims, not passing Chatter test receipts.
The least disruptive next experiment is rendered-component testing through the
existing DesktopRuntimeProvider injection seam: valid, invalid and internal
failure results; cancel/revalidate; export failure; and overlapping update
requests. Keep typed event sequences and bounded completion assertions. This
would complement the existing Rust bridge tests, not replace real IPC evidence.
Before adding a native automation dependency, review its permissions, lifecycle, dependency cost and release exclusion. An embedded control server must never ship in production artifacts; require an explicit test-build configuration and an automated release-exclusion check. Do not add a second validator, test-only production command, or global IPC escape hatch. Keep settings/cache isolated.
The existing native smoke deck only checks launch UI. Its configuration also
resolves the binary under apps/target, rather than the workspace target, and
does not select the Windows .exe suffix. Correct binary discovery and retain
an actual successful run before treating this deck as a usable release check.
The smoke test’s suggested validate_for_test shortcut is not an approved design.
CSS rendering differs slightly between WebKit (Linux) and Chromium (Windows). Visual regressions are possible, consider screenshot comparison tests if this becomes a problem.
Test Data
All tests use the reference corpus at corpus/reference/. This
corpus is checked into the repo and must always pass validation with
zero hard errors (warnings are allowed). The exact set of files and
the current warning-emitting files are whatever
rg --files corpus/reference -g '*.cha' and the validator
report, do not hard-code those lists here.
Do not create ad-hoc .cha test files. Use existing reference corpus files
or ask the user to provide test data.
CI Integration
Add to the existing CI workflow:
# Rust integration tests (fast, always run)
- name: Desktop integration tests
run: cargo test -p chatter-desktop --test validation_bridge
# E2E tests (slow, release branches only)
- name: Build desktop app
if: startsWith(github.ref, 'refs/heads/release')
run: cargo tauri build --debug
- name: E2E smoke tests
if: startsWith(github.ref, 'refs/heads/release')
run: |
tauri-driver &
sleep 2
cd apps/chatter-desktop && npm run test:e2e
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
Library Usage
Status: Current Last updated: 2026-10-07 (commit 5e895791)
The TalkBank Rust crates can be used as dependencies in your own Rust
projects for parsing, validating, and manipulating CHAT files. This page
shows the most common entry points; the API reference on docs.rs (once
published) is the authoritative source. Until then, treat the rustdoc
comments inside each crate’s src/lib.rs as the source of truth.
Examples on this page are mirrored as a real Cargo test at
crates/talkbank-transform/tests/book_library_usage_examples.rs. The book renders them asrust,ignoreso mdbook doesn’t try to link against the workspace’s many compiled crate variants; the parallel test runs the same code undercargo testand is what catches API drift between this page and the libraries. If you edit either, update both.
Important: some legacy tree-sitter fragment helpers are synthetic rather than semantically honest. They can inject fragment input into boilerplate CHAT text and parse the resulting synthetic file. Prefer full-file parsing for real tree-sitter use, and do not treat legacy fragment helpers as the long-term fragment API. For direct-parser fragment semantics, use direct-parser-native tests instead of treating synthetic wrappers as the oracle.
Adding Dependencies
The TalkBank library crates are source-available from this repository. They are
not yet published on crates.io, so depend on them from the public repo via git
(pinned to a release tag), or via local path dependencies from a
TalkBank/chatter checkout for local development:
[dependencies]
talkbank-model = { path = "../chatter/crates/talkbank-model" }
talkbank-transform = { path = "../chatter/crates/talkbank-transform" }
talkbank-parser = { path = "../chatter/crates/talkbank-parser" }
The published-crate workflow is tracked separately; once it lands these
paths can become version = "X.Y" deps.
Explicit English number generation
talkbank_transform::num_words::expand_number spells supported numeric input
for callers constructing transcripts. It is not a CHAT parser, validator, or
automatic repair of existing speech. The caller remains responsible for the
intended pronunciation; unsupported input is preserved exactly rather than
partially rewritten or assigned a guessed pronunciation.
English ordinal composition supports 0-9999 and retains British-style
conjunctions without prose commas. Larger suffix-bearing ordinals are preserved,
including their original suffix: 10001st must not become 10001th.
English decade expansion requires a multiple of ten in one of these domains:
| Input domain | Examples | Generated words |
|---|---|---|
| Shorthand 0-90 | 0s, 80s, 90s | zeros, eighties, nineties |
| Full-year 1100-2990 | 1100s, 1950s, 2990s | eleven hundreds, nineteen fifties, twenty-nine nineties |
Other suffix forms, such as 21s, 100s, 2001s, and 3000s, remain
unchanged. This deliberately avoids inventing phrases such as “twenty-ones”
or “two thousand and ones”. Neither support nor preservation certifies CHAT
validity: English digit-bearing words still require an authored transcription.
Initial zero has omission semantics in CHAT, so the raw-string 0s and
leading-zero API cases are not interchangeable with parsed CHAT words.
These are bounded generation policies, not claims about every possible English number expression or the speaker’s intent. The testing contract pairs authored written reference controls with invalid numeric specimens and keeps raw-string boundary tests separate.
Parsing and Validating a CHAT File
The simplest entry point is parse_and_validate from
talkbank-transform. It takes the source text and a
ParseValidateOptions, returns a fully constructed ChatFile, or a
PipelineError if parsing or validation failed.
extern crate talkbank_model;
extern crate talkbank_transform;
use talkbank_model::ParseValidateOptions;
use talkbank_transform::parse_and_validate;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let source = std::fs::read_to_string("file.cha")?;
let options = ParseValidateOptions::default().with_validation();
let chat_file = parse_and_validate(&source, options)?;
for utt in chat_file.utterances() {
println!("Speaker: {}", utt.main.speaker);
}
Ok(())
}
parse_and_validate returns a mutable ChatFile and may skip validation
according to its options. Use parse_validated_with_parser for an immutable
ValidChatFile whose policy cannot skip validation.
When original bytes must be returned unchanged, use parse_source_with_parser.
Its opaque ParsedSourceChat permits read-only inspection before admit
validates that same parse with complete rules and tier alignment; this
unchanged-output admission has no caller-selected policy. Success returns
AdmittedSourceChat, coupling
original bytes with their immutable model proof; into_owned keeps the binding
across asynchronous storage. No constructor accepts separate text and model
evidence. into_product relinquishes admission authority for recovery work,
and into_valid_file relinquishes the original-byte binding before editing.
Removing and regenerating tiers requires the separate replacement admission
below; an unchanged source cannot claim a removal exemption.
chat_file.utterances() returns an iterator over &Utterance derived
from the file’s lines (utterances are interleaved with headers and
comments in source order).
For batch workflows where parser construction overhead matters, reuse a
single TreeSitterParser and call parse_and_validate_with_parser:
extern crate talkbank_model;
extern crate talkbank_parser;
extern crate talkbank_transform;
use talkbank_model::ParseValidateOptions;
use talkbank_parser::TreeSitterParser;
use talkbank_transform::parse_and_validate_with_parser;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let chat_files: Vec<std::path::PathBuf> = Vec::new();
let parser = TreeSitterParser::new()?;
let options = ParseValidateOptions::default().with_validation();
for path in &chat_files {
let source = std::fs::read_to_string(path)?;
let chat_file = parse_and_validate_with_parser(&parser, &source, options.clone())?;
let _ = chat_file;
}
Ok(())
}
ParseValidateOptions also exposes with_alignment() (implies
with_validation(), additionally validates cross-tier alignment for
%mor, %gra, %pho, %wor), with_level(CheckLevel) to choose the
level directly (CheckLevel::ParseOnly or
CheckLevel::Validate(AlignmentValidation::Structure | IncludeTierAlignment)),
and
with_strict_linkers() (enables the
opt-in cross-utterance quotation and completion checks; chatter validate --list-checks marks each code it turns on as [Opt-in]).
Admission for dependent-tier regeneration
TreeSitterParser::admit_replacing_tiers is a separate input boundary for a
transform that will actually remove and regenerate dependent tiers.
ReplacementTiers::Morphosyntax removes %mor and %gra together;
ReplacementTiers::WordTiming removes %wor while retaining morphology.
For header-dependent policies, admit_planned_tiers supplies immutable references
to all typed headers to a one-shot callback. Returning None preserves every
tier and requires whole-document validity; returning a replacement selection
physically removes only that domain. Header lowering happens once, before
pending source-bound utterances lower once. The callback sees misplaced headers
too, but cannot excuse header faults. AdmittedReplacement::selection records
the actual optional removal selection; an unfinished producer plan is a tool
failure, never CHAT invalidity or validation success.
The parser parses once. Its generated, source-bound traversal selects concrete tier nodes and omits them before lowering the retained model. The recovery backstop excludes those exact discarded subtrees, not diagnostic codes or overlapping byte spans. Headers, main-tier/retrace errors, retained dependent tiers, unclassified recovery and global forbidden-control checks remain subject to refusal. The retained model must pass normal validation and tier alignment.
The resulting AdmittedReplacement owns both immutable ValidChatFile evidence
for the retained document and the original ParsedSource. Its removal receipt
records each selected tier’s original span and whether it contained syntax
recovery. Absence of syntax recovery does not prove that discarded content was
semantically valid. This API does not certify the original bytes as valid
CHAT and never repairs or marks recovered content clean.
A pass-through, keep-tier or incremental preservation path must validate what
it will retain; it cannot use whole-file tier removal to excuse faults in copied
content. After admission, transformations still need completion and output
validation before writing. into_valid_file explicitly consumes the
source/removal association; ValidChatFile::into_unchecked consumes validity
before editing. No serialized CHAT is reparsed to establish these states.
Adaptive word-timing regeneration
admit_word_timing_plan returns WordTimingAdmission, not unconditional
AdmittedReplacement. Its disposition is preserved, replaced with a completely
valid retained document, or Regenerating(AdmittedTimingRegeneration).
The last branch is possible only when concrete discarded source-bound %wor
tiers contain recorded timing and their removal leaves the linked-media timing
requirement outstanding. The model and original source remain bound to the
removal receipt; no original-byte output proof is issued.
Only the word tiers actually at fault are removed; the others are retained.
Each RemovedTier carries a RemovalCause naming why: OwnLowering (the
tier itself did not lower cleanly), LocatedValidation (validation errors
inside its source span), UnattributedValidation (errors in no word tier, so
every remaining word tier was removed, sharing those diagnostics) or
Selected (a fixed ReplacementTiers request). RemovalCause::diagnostics
returns the diagnostics that cost the tier, so a consumer can report which
codes cost which tier.
PendingTimingChatFile is checked working structure, not valid CHAT.
The payload has no into_valid_file, accepted-document serialization or
validity constructor. It keeps its document and its MediaTimingObligation
together: write the regenerated timing through document_mut, then
discharge checks the payload’s document with the very rule that issued the
obligation (E544), handing that document out once the rule no longer fires
and the payload back while it would. The obligation alone has no discharge,
so the verdict always concerns the document the consumer goes on to admit.
UTR/FA must establish actual timing, and final output must still pass
complete checked construction, including E544. No header is
changed and no emitted diagnostic is filtered. Normal validation still reports
E544; unrelated retained faults, internal failures and preservation/NoAlign
paths still refuse. Missing timing without concrete discarded timing is not
eligible for this adaptive state.
Working with the Model
ChatFile stores participants and language metadata as top-level fields
populated from @Participants / @ID / @Languages headers during
parsing. Utterances live in lines and are iterated via
chat_file.utterances().
extern crate talkbank_model;
extern crate talkbank_transform;
use talkbank_model::DependentTier;
use talkbank_model::ParseValidateOptions;
use talkbank_transform::parse_and_validate;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let source = "\
@UTF8
@Begin
@Languages:\teng
@Participants:\tCHI Target_Child
@ID:\teng|test|CHI|||||Target_Child|||
*CHI:\thello world .
%mor:\tco|hello n|world .
@End
";
let chat_file = parse_and_validate(source, ParseValidateOptions::default().with_validation())?;
// Participant metadata is top-level on the ChatFile.
let _participants = &chat_file.participants;
// Iterate utterances and their dependent tiers.
for utt in chat_file.utterances() {
for tier in &utt.dependent_tiers {
if let DependentTier::Mor(mor_tier) = tier {
for item in mor_tier.items() {
println!("POS: {}, Lemma: {}", item.main.pos, item.main.lemma);
}
}
}
}
Ok(())
}
DependentTier is a closed-set enum (Mor, Gra, Pho, Mod, Sin,
Act, Add, Com, Err, Exp, Gpx, Int, Lan, …); match on the
variants you care about and ignore the rest. MorTier::items() returns
&[Mor]; each Mor has a main MorWord plus optional post-clitics.
Serializing to CHAT
Bring the WriteChat trait into scope and call to_chat_string() for a
fully-rendered CHAT string, or write_chat(&mut writer) to stream into
any std::fmt::Write.
extern crate talkbank_model;
extern crate talkbank_transform;
use std::fmt::Write as _;
use talkbank_model::ParseValidateOptions;
use talkbank_model::WriteChat;
use talkbank_transform::parse_and_validate;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let source = "@UTF8\n@Begin\n@Languages:\teng\n@Participants:\tCHI Target_Child\n@ID:\teng|test|CHI|||||Target_Child|||\n*CHI:\thello .\n@End\n";
let chat_file = parse_and_validate(source, ParseValidateOptions::default().with_validation())?;
// Convenience: render to a fresh String.
let chat_text = chat_file.to_chat_string();
assert!(chat_text.starts_with("@UTF8"));
// Streaming: write into any std::fmt::Write sink.
let mut output = String::new();
chat_file.write_chat(&mut output)?;
Ok(())
}
Serializing to JSON
Prefer the schema-validated helpers in talkbank_transform::json:
to_json_pretty_validated checks the output against the JSON schema and
catches drift between the data model and the schema. The unvalidated
variants are a faster bypass when you’ve already validated upstream.
extern crate talkbank_model;
extern crate talkbank_transform;
use talkbank_model::ParseValidateOptions;
use talkbank_transform::json::to_json_pretty_validated;
use talkbank_transform::parse_and_validate;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let source = "@UTF8\n@Begin\n@Languages:\teng\n@Participants:\tCHI Target_Child\n@ID:\teng|test|CHI|||||Target_Child|||\n*CHI:\thi .\n@End\n";
let chat_file = parse_and_validate(source, ParseValidateOptions::default().with_validation())?;
let json = to_json_pretty_validated(&chat_file)?;
assert!(json.contains("\"speaker\""));
Ok(())
}
The schema for ChatFile lives at schema/chat-file.schema.json and is
regenerated from the Rust types via just schema-gen. For arbitrary
serde values (not just ChatFile), to_json_unvalidated /
to_json_pretty_unvalidated work the same way without the schema step.
To go straight from CHAT text to JSON, talkbank_transform::chat_to_json
parses, validates and serializes in one call, taking a JsonLayout
(Pretty or Compact) rather than a bare bool:
use talkbank_model::ParseValidateOptions;
use talkbank_transform::{JsonLayout, chat_to_json};
let json = chat_to_json(source, ParseValidateOptions::default().with_validation(), JsonLayout::Compact)?;
Custom Error Handling
Lower-level parser entry points stream diagnostics through the
ErrorSink trait. Implement it to collect, count, filter, or forward
errors as they arrive, useful when you need finer-grained control than
the Result<ChatFile, PipelineError> shape parse_and_validate returns.
extern crate talkbank_model;
use talkbank_model::ErrorSink;
use talkbank_model::ParseError;
struct MyErrorHandler;
impl ErrorSink for MyErrorHandler {
fn report(&self, error: ParseError) {
// Custom handling: log, filter, count, etc.
eprintln!("[{}] {}", error.code, error.message);
}
}
ErrorSink is Send + Sync, and a blanket &T: ErrorSink impl means
borrowed references are sinks too, no Arc wrapper required. The
built-in ErrorCollector (gathers into a Vec), ParseTracker (counts
by severity), and NullErrorSink (discards) cover most common needs;
implement ErrorSink directly for everything else.
Crate Selection Guide
| Need | Crate |
|---|---|
Data model types, error types, WriteChat, ErrorSink | talkbank-model |
| Tree-sitter CHAT parsing (low-level) | talkbank-parser |
| Full pipeline (parse + validate + JSON, schema validation) | talkbank-transform |
talkbank-model is the foundation, every other crate depends on it. If
all you need are the AST types and validation, model alone is enough.
talkbank-transform brings parsing + JSON + caching.
Batchalign3-Facing Surface
If you are building Batchalign3 or another external consumer, the stable surface is usually:
| Batchalign3 need | Prefer |
|---|---|
| Canonical full-file parsing | talkbank-parser |
| Parse/validate contracts and typed model access | talkbank-model |
Alignment-aware downstream consumers (align, compare, benchmark) | talkbank-model alignment helpers plus the model AST |
| Whole-pipeline parse+validate+convert | talkbank-transform |
For batch workflows, keep parser instances reusable and keep alignment logic separate from parse semantics.
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).
JSON Output Reference
Status: Reference Last updated: 2026-10-02 (commit 2d7e886b)
This document describes the structure of JSON produced by chatter to-json.
For the formal JSON Schema, see JSON Schema.
Quick Start
# Default: parse + validate + align, pretty-printed, schema-checked
chatter to-json file.cha
# Write to file
chatter to-json file.cha -o file.json
# Skip validation (parse only, faster)
chatter to-json file.cha --skip-validation
# Skip alignment only
chatter to-json file.cha --skip-alignment
Validation and alignment are on by default. Use --skip-validation
or --skip-alignment to opt out.
Top-Level Structure
{
"lines": [ ... ]
}
A ChatFile is a flat list of lines. Each line has a line_type discriminator:
line_type | Description |
|---|---|
"header" | File header (@Begin, @Languages, @Participants, etc.) |
"utterance" | Main tier + dependent tiers + alignment |
"comment" | @Comment: lines |
Word Fields
Words are the fundamental unit. Every word in the main tier content array
carries these fields:
| Field | Type | Always? | Description |
|---|---|---|---|
type | "word" | yes | Discriminator |
raw_text | string | yes | Current CHAT spelling derived from typed structure, including markers |
cleaned_text | string | yes | NLP-ready text (shortenings restored, markers stripped) |
content | array | yes | Structured breakdown of word parts (see below) |
category | string | no | "omission", "filler", "nonword", "fragment", "ca_omission" |
form_type | string | no | Special form code: "c", "d", "f", "x", etc. |
lang | object | no | Language marker (see Language-Switched example) |
untranscribed | string | no | "unintelligible" (xxx), "phonetic" (yyy), "untranscribed" (www) |
Word content items use "content" for the text value:
{ "type": "text", "content": "dog" }
Computed Fields
raw_text, cleaned_text and untranscribed are computed views, not
independently authoritative input fields. Imported values cannot override the
typed content, category, form, language or part-of-speech markers. Import still
requires model validation; ignoring a stale display field does not admit
malformed structural fields.
raw_text is not a byte-exact source slice. In particular, recovery may retain
a partial typed word while separately reporting rejected source material.
Source-coordinate consumers must retain the original source and use its spans;
they must not index derived spelling with offsets into that source.
Rust callers use Word::new(WordText) for a nonempty plain-text model, then
explicit typed builders for markers. Nonempty construction alone is not full
CHAT admission. For external CHAT/ASR tokens, use ChatParser::parse_word_fragment
and handle rejection explicitly instead of falling back to unchecked construction.
There are no independent raw/cleaned constructor arguments and no set_raw_text.
raw_text() returns an owned string; use WriteChat::write_chat or
Display when streaming avoids an intermediate allocation. JSON serialization
streams that display projection directly. The cleaned-text cache is
invalidated by content mutation; no raw-text cache can become stale after
public marker mutation.
-
cleaned_text: ConcatenatesTextandShorteningelements fromcontent. Excludes lengthening markers (:), stress markers, CA elements, overlap points, compound markers, and underline markers. Example:sit(ting)→"sitting". -
untranscribed: Present only whencleaned_textis"xxx","yyy", or"www".
Word Examples
Simple Word
dog
{
"type": "word",
"raw_text": "dog",
"cleaned_text": "dog",
"content": [{ "type": "text", "content": "dog" }]
}
Filler
&-uh
{
"type": "word",
"raw_text": "&-uh",
"cleaned_text": "uh",
"content": [{ "type": "text", "content": "uh" }],
"category": "filler"
}
Untranscribed
xxx
{
"type": "word",
"raw_text": "xxx",
"cleaned_text": "xxx",
"content": [{ "type": "text", "content": "xxx" }],
"untranscribed": "unintelligible"
}
Compound
ice+cream
{
"type": "word",
"raw_text": "ice+cream",
"cleaned_text": "icecream",
"content": [
{ "type": "text", "content": "ice" },
{ "type": "compound_marker", "content": { "span": { "start": 0, "end": 1 } } },
{ "type": "text", "content": "cream" }
]
}
Omission
0she
{
"type": "word",
"raw_text": "0she",
"cleaned_text": "she",
"content": [{ "type": "text", "content": "she" }],
"category": "omission"
}
Nonword
&~baba
{
"type": "word",
"raw_text": "&~baba",
"cleaned_text": "baba",
"content": [{ "type": "text", "content": "baba" }],
"category": "nonword"
}
Special Form
doggy@c
{
"type": "word",
"raw_text": "doggy@c",
"cleaned_text": "doggy",
"content": [{ "type": "text", "content": "doggy" }],
"form_type": "c"
}
Language-Switched
maison@s:fra
{
"type": "word",
"raw_text": "maison@s:fra",
"cleaned_text": "maison",
"content": [{ "type": "text", "content": "maison" }],
"lang": { "type": "explicit", "code": "fra" }
}
The lang field has variants: {"type": "shortcut"} (bare @s),
{"type": "explicit", "code": "fra"} (@s:fra), and
{"type": "multiple", "code": ["eng", "zho"]} (@s:eng+zho).
A multi-word switch is an ANNOTATION on the group, not a field on each word:
<how to do it> [@s] serializes as an annotation
{"type": "code_switch", "kind": "shortcut"}, and [@s:hin] as
{"type": "code_switch", "kind": "explicit", "code": "hin"}. The words inside
keep lang: null unless they carry a suffix of their own, so a consumer
reading only lang will under-report language switches. Read
language_metadata instead, which is where the resolved answer lives for
every word regardless of which mark produced it.
Utterances
An utterance line contains:
{
"line_type": "utterance",
"main": {
"speaker": "CHI",
"content": {
"content": [ ... ],
"terminator": { "type": "period" },
"bullet": { "start_ms": 0, "end_ms": 3042 }
}
},
"dependent_tiers": [ ... ],
"alignments": { ... },
"utterance_language": { "status": "resolved_default", "code": "eng" },
"language_metadata": { ... }
}
Key structural points:
- The utterance body is under
"main", not"utterance". content,terminator, andbulletare nested insidemain.content.terminatoris an object with atypefield ("period","question","exclamation", etc.), not a bare string.bullet(utterance-level timing) is insidemain.content, omitted when absent (not present asnull).dependent_tiers,alignments,utterance_language, andlanguage_metadataare top-level siblings ofmain. Emptydependent_tiersandalignmentsare omitted when there is nothing to report.
language_metadata
One entry per WORD of the main tier, in in-order traversal order, under
language_metadata.word_languages:
"language_metadata": {
"tier_language": "zho",
"word_languages": [
{ "languages": { "single": "zho" }, "source": "default" },
{ "languages": { "single": "eng" }, "source": "word_shortcut" }
]
}
Three properties worth knowing before consuming it:
- Every word at any depth, including words inside quotations, phonological groups, sign groups and retraces. A retraced word was spoken and has a language, so it gets an entry.
- The produced form only. For
dog [: cat]the entry describesdog, what the speaker actually said, not the correction. sourcesays which mark decided the language. A span and a word suffix can resolve to the identical CODE, so a consumer asking “did the transcriber mark this word, or the stretch around it?” can only answer from this field. Precedence is innermost-first: the word’s own mark beats an enclosing span, which beats the utterance. The values are enumerated with a description each inschema/chat-file.schema.json, generated from the enum; they are deliberately not copied here, because a copy would go stale.- Position is the array subscript, and nothing else. There is no index
field: read it with the equivalent of
enumerate(). In particular this order is not an alignment index. The tier domains disagree about what they count (%morexcludes retraces,%phocounts them), so correlating with%moror%grapositions must go throughalignments, not through this order. There is noword_indexfield, because it would be derivable and misleading.
Content Items
main.content.content is a heterogeneous array. Each item has a type discriminator:
| Type | Description |
|---|---|
"word" | A word token (see Word Fields above) |
"event" | Non-verbal action (&=laughs) |
"pause" | Timed or untimed pause ((.), (0.5)) |
"group" | Bracketed group (<word word>) |
"separator" | Tag markers, linkers, etc. |
Dependent Tiers
When present, dependent_tiers is an array of tagged objects:
"dependent_tiers": [
{
"type": "Mor",
"data": {
"tier_type": "Mor",
"items": [
{
"main": { "pos": "pron", "lemma": "I" }
},
{
"main": { "pos": "verb", "lemma": "want", "features": ["Fin", "Ind", "Pres"] }
}
],
"terminator": "."
}
},
{
"type": "Gra",
"data": {
"tier_type": "Gra",
"relations": [
{ "index": 1, "head": 2, "relation": "NSUBJ" },
{ "index": 2, "head": 0, "relation": "ROOT" }
]
}
}
]
type | Tier | Description |
|---|---|---|
"Mor" | %mor | Morphological analysis (POS tags, lemmas, features, clitics) |
"Gra" | %gra | Grammatical relations (dependency arcs) |
"Pho" | %pho | Phonological transcription |
"Sin" | %sin | Syntax tier |
"Wor" | %wor | Word-level timing (items with inline_bullet) |
| Other | %xxx | User-defined dependent tiers |
%wor Tier
The Wor tier contains word items with timing:
{
"type": "Wor",
"data": {
"items": [
{
"kind": "word",
"raw_text": "hello",
"cleaned_text": "hello",
"content": [{ "type": "text", "content": "hello" }],
"inline_bullet": { "start_ms": 100, "end_ms": 300 }
}
],
"terminator": { "type": "period" }
}
}
Note that %wor items use "kind" instead of "type" for their discriminator,
since "type" is used by the tier envelope.
Alignment Data
When validation runs (the default), the alignments object contains:
units: per-tier index arrays (for internal bookkeeping)- Named tier pairs (e.g.,
mor,gra) with alignment mappings
"alignments": {
"units": {
"main_mor": [{"index": 0}, {"index": 1}],
"main_pho": [{"index": 0}, {"index": 1}],
"main_sin": [{"index": 0}, {"index": 1}],
"main_wor": [{"index": 0}, {"index": 1}],
"mor": [{"index": 0}, {"index": 1}]
},
"mor": {
"pairs": [
{ "source_index": 0, "target_index": 0 },
{ "source_index": 1, "target_index": 1 }
],
"errors": []
}
}
Alignment links each main-tier word (source_index) to its corresponding
dependent-tier item (target_index) by position. errors contains any
alignment-level diagnostics (count mismatches, etc.) and is [] when
alignment validated cleanly.
Headers
Headers use the header object with a type discriminator:
| Type | Header | Key Fields |
|---|---|---|
"utf8" | @UTF8 | , |
"begin" | @Begin | , |
"end" | @End | , |
"languages" | @Languages | codes |
"participants" | @Participants | entries (speaker_code, name, role) |
"id" | @ID | language, corpus, speaker, role, age, sex, … |
"media" | @Media | filename, media_type, status |
"comment" | @Comment | text |
"date" | @Date | date |
"options" | @Options | options (array of strings) |
See the JSON Schema for the complete list of header types and fields.
Timing
Utterance-level timing appears in main.content.bullet:
"bullet": {
"start_ms": 1234,
"end_ms": 5678
}
Word-level timing (from %wor tier) appears in inline_bullet on individual
words within the Wor dependent tier.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
JSON Schema
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
This repository generates JSON Schema from Rust-owned types with
schemars for the ChatFile transcript model used
by chatter to-json.
Keeping that schema generated from the Rust source of truth lets cross-language integrations consume a stable contract without re-deriving the shapes by hand.
Available schemas
| Schema | Canonical URL | Repository | Generator |
|---|---|---|---|
ChatFile transcript model | https://talkbank.org/schemas/v0.1/chat-file.json | schema/chat-file.schema.json | just schema-gen |
The generated schema declares both $schema (JSON Schema 2020-12) and $id
(the canonical URL above). External consumers that want to track the
current transcript-model version should follow the v0.1 URL; there is
no /latest/ alias in the generated artifacts.
Transcript schema: ChatFile
chatter to-json converts CHAT transcripts into a structured JSON form backed
by the same ChatFile model used by the parser, validator, and serializer.
How chatter to-json uses it
By default, chatter to-json:
- validates the CHAT input,
- checks dependent-tier alignment unless
--skip-alignmentis passed, and - validates the emitted JSON against the schema unless
--skip-schema-validationis passed.
These controls are independent. --skip-schema-validation skips only the
JSON Schema check after CHAT parsing and validation. It retains the input
filename, so filename-dependent rules such as E531 (@Media name mismatch)
still run for both single files and directory conversions. A directory run
prints individual parse/validation diagnostics and exits unsuccessfully if
any file fails, while retaining successfully converted siblings.
Library callers that know the transcript name can use
chat_to_json_with_schema_policy with TranscriptName,
JsonLayout::{Pretty, Compact} and JsonSchemaPolicy::{Validate, Skip}. The
layout and the policy select serialization only; they cannot alter the parse
options or discard the transcript name.
Useful flags:
chatter to-json input.cha --skip-validation
chatter to-json input.cha --skip-alignment
chatter to-json input.cha --skip-schema-validation
chatter from-json deserializes JSON back into the internal ChatFile model
and re-serializes it to CHAT format. The input should conform to this schema.
Roundtrip expectations
The CHAT-to-JSON-to-CHAT pipeline is intended to preserve the ChatFile model:
chatter to-json input.cha -o intermediate.json
chatter from-json intermediate.json -o output.cha
diff input.cha output.cha
Both directions go through the same typed model. When changing the parser, serializer, or schema generation, confirm roundtrip behavior with the existing roundtrip test suites rather than assuming byte-for-byte identity.
Using the schema externally
Validate JSON in Python
import json
import jsonschema
import urllib.request
schema_url = "https://talkbank.org/schemas/v0.1/chat-file.json"
schema = json.loads(urllib.request.urlopen(schema_url).read())
with open("transcript.json") as f:
data = json.load(f)
jsonschema.validate(data, schema)
IDE autocompletion
{
"$schema": "https://talkbank.org/schemas/v0.1/chat-file.json",
"lines": [],
"participants": {},
"languages": [],
"options": []
}
Generate types from the schema
Tools like quicktype, json-schema-to-typescript, and datamodel-code-generator can generate typed structs or classes from the schema for TypeScript, Python, Go, and other languages.
Regenerating the schema
After changing transcript-model types in talkbank-model:
cd chatter
just schema-gen
This writes the checked-in schema artifact in schema/. CI already checks that
generated artifacts stay in sync.
Code references
schema/chat-file.schema.json: generated schemacrates/talkbank-transform/src/json.rs: schema loading and validationcrates/talkbank-model/src/model/: Rust data modeltests/integration/generate_schema/: shared schema generation helpers
Schema generation is an explicit operation: just schema-gen selects the
ignored generator, and just regen includes it. Ordinary tests only check
currency, so they do not rewrite the schema while checking the version embedded
at compile time. Identical output preserves the file’s modification time to
avoid invalidating builds that embed it. A currency mismatch reports the repair
command without dumping the entire schema into the test log.
The generator preserves schemars’ Draft 2020-12 structure. In this dialect,
$ref permits sibling keywords,
including the tag constraints for internally tagged enums. No allOf rewrite
is required: a rewrite that recursed through the schema would also traverse
literal const values and could change their meaning, so the generator
leaves the structure as emitted. The generated schema regression verifies
both the tag and the referenced payload constraints.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Diagnostic and JSON Output Contract
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
This page documents the machine-readable JSON surfaces currently exposed by the
top-level chatter CLI.
Stability policy
- Treat field names documented here as the public contract.
- Treat additional fields as additive unless this page says otherwise.
- Treat message wording as human-facing text, not a stable machine contract.
- Every record is serialized from one typed model, so a field is spelled
the same in every record that carries it. Key order is fixed (
typefirst, then a record’s own tag such asstatus) but is not part of the contract: read records as objects.
chatter validate ... --format json
Both chatter validate FILE --format json and
chatter validate DIR --format json emit newline-delimited JSON
(NDJSON) on stdout, with the same record shapes in both modes:
- exactly one file record per file the run accounted for, then
- one final summary record.
A single-file invocation still emits a file record followed by a summary record; it is not a single-object surface.
Per-file records
Valid files (no diagnostic of severity Error; a warnings array is
present when the file showed warnings):
{"type":"file","status":"valid","file":"/path/to/file.cha","cache_hit":false}
{"type":"file","status":"valid","file":"/path/to/Session.cha","cache_hit":false,"warnings":[{"code":"W110","severity":"Warning","message":"..."}]}
Invalid files: error_count counts the diagnostics of severity Error, and
the errors array holds every diagnostic shown, warnings included, each with
its severity. The note field is appended when the validator stopped
further checks because of structural errors:
{
"type": "file",
"status": "invalid",
"file": "/path/to/file.cha",
"error_count": 1,
"errors": [
{
"code": "E502",
"severity": "Error",
"message": "Missing @End header at end of file"
}
],
"note": "Some additional checks may not have run because of structural errors. Fix the structural errors first, then re-validate."
}
Files whose roundtrip check failed (--roundtrip) use
"status":"roundtrip_failed" with a reason string, a diff (the first
differing lines, null for a verdict read from the cache) and, when the
file showed warnings, a warnings array; they count among the summary’s
invalid files.
Read-failure files use "status":"read_error" with an error string, and
count among the summary’s invalid files. A file the parser cannot make
sense of is an invalid record carrying its parse diagnostics: both parsers
always produce a model with diagnostics, so there is no separate parse-failure
status.
Internal tool failures use "status":"internal_failure" with an error
string and the retained errors array. They do not also emit an invalid
record. E001 belongs to DiagnosticKind::InternalFailure, not CHAT invalidity.
The attempt determines neither validity nor invalidity, even if other findings
were collected. It is not cached as a validation verdict and exits unsuccessfully.
Report the failure; do not alter CHAT merely to accommodate a producer bug.
Suppression or severity downgrades cannot admit the failed attempt.
This also applies to producer faults during --roundtrip reparsing: the attempt
contributes to neither roundtrip counter and writes neither cache verdict.
Retained fault spans from that reparse refer to serialized intermediate text,
not the original file; they are not emitted as original-source annotations.
Earlier input findings are withheld if that roundtrip attempt fails internally,
so they cannot publish a conflicting invalid-file verdict.
Summary record
{
"type": "summary",
"directory": "/path/to/dir",
"total_files": 2,
"valid": 1,
"invalid": 1,
"internal_failures": 0,
"outcome": "complete",
"cache_hits": 0,
"cache_misses": 2,
"cache_hit_rate": 0.0,
"cache_errors": 0
}
outcome says how the run ended: "complete" (every discovered file was
accounted for), "nothing_found" (the input named no transcript; this
summary carries only directory and outcome), "stopped" (a stop record
precedes the summary) or "incomplete" (an incomplete record precedes it).
Only a "complete" summary counts the whole input, so a consumer reads the
ending before it reads any total.
When --roundtrip is set, the summary also includes
roundtrip_passed and roundtrip_failed counters. cache_errors counts
cache reads and writes that failed (a locked or corrupt database); those
files were validated without the cache, so their results stand.
cache_hits and cache_misses count only files that consulted a cache. A
file that could not be read, or any file of a run whose cache did not open,
is neither, and cache_hit_rate (hits over consultations, a percentage) is
null when nothing consulted the cache. A file record’s cache_hit is
true only when the verdict that decided its status (validation or
roundtrip) was served from the cache.
Stop record
Emitted once, before the summary, when the run stopped with files left unvalidated:
{"type":"stop","reason":"max_errors","limit":50,"unprocessed_files":120}
{"type":"stop","reason":"cancelled","unprocessed_files":3}
reason is "max_errors" (--max-errors N was reached; limit is N) or
"cancelled" (Ctrl-C). It comes from the runner’s terminal event, so it is
emitted only when files really were left: a limit reached by the run’s last
file stopped nothing and is not reported. A stopped run exits 1. Text mode
prints the same fact on stderr, as
Stopped after reaching the error limit (N); M file(s) were not validated.
or Validation cancelled; M file(s) were not validated.
Incomplete and aborted records
{"type":"incomplete","lost_files":2,"total_files":500,"cause":"worker_faults","detail":"1 worker(s) failed with an internal error."}
{"type":"aborted","reason":"..."}
incomplete precedes the summary of a run that lost files nobody asked to
skip. cause is "worker_faults" (workers panicked, could not create their
parser, or could not be started; detail lists each) or "unexplained" (no
worker failed and no stop was requested, a defect in the validator).
detail is a sentence for a person, not a format to parse. aborted
replaces the summary of a run that died before producing totals. Both exit
1.
Notice record
A fact about the run, said before it starts:
{"type":"notice","notice":"suppressing","codes":["E736","E737"]}
{"type":"notice","notice":"deprecated_flag","flag":"--check-xphon"}
{"type":"notice","notice":"interrupt_unavailable","error":"..."}
interrupt_unavailable says the Ctrl-C handler could not be installed, so
an interrupt ends the process without a stop record. Text mode prints the
same notes on stderr (note: suppressing 2 code(s): E736, E737). An input
path that cannot be read is not a notice: it is a file record with
"status": "read_error", and it fails the run. A run that found no
transcript at all ends with a "nothing_found" summary and exits 1, since
it validated nothing. Ctrl-C in JSON mode prints nothing on stderr; the run
ends with a "cancelled" stop record.
Cache record
Emitted only when cache maintenance actually did something: a prune
reclaimed rows, --force cleared entries, or maintenance failed and
the run continued without it.
{"type":"cache","action":"clear","entries_cleared":1}
{"type":"cache","action":"prune","rows_deleted":12,"versions_deleted":2}
{"type":"cache","action":"warning","operation":"initialize","error":"..."}
A warning’s operation is one of "initialize" or "clear" (the cache
operations a validate run performs before it starts); error is a
sentence for a person.
They are records rather than stderr sentences because JSON mode leaves
stderr empty (see below), and rather than silenced because they are
results a caller can act on, so they belong on the stream in a form a
reader can parse. A consumer should dispatch on type and ignore values
it does not recognise.
Contract notes
-
The
typefield is stable, and its values are"file","summary","cache","stop","incomplete","aborted"and"notice". Treat an unknowntypeas ignorable rather than as an error: new record kinds may appear. -
Stderr is not part of the JSON surface and is empty in JSON mode. Anything a run wants to tell you arrives as a record on stdout.
-
For file records:
fileandstatusare stable;cache_hitandwarningsare stable forvalidrecords.error_countanderrorsare stable forinvalidrecords;reason,diffandwarningsforroundtrip_failedrecords. -
For summary records:
directory,total_files,valid,invalid,internal_failures,cache_hits,cache_misses,cache_hit_rate,cache_errors, andoutcomeare stable. -
For stop records:
reason("max_errors"or"cancelled"),unprocessed_files, and for"max_errors"limit, are stable. -
statusvalues:valid,invalid,roundtrip_failed,read_error,internal_failure. New status values may appear. -
Errors do not include a byte-offset
locationfield in the NDJSON surface; for byte-offset diagnostics use the LSP or the non-JSON renderer. -
The
notefield on invalid file records is human-facing guidance and may be added or omitted between releases. -
Exit code
0means all files validated successfully; exit code1means at least one file failed or an I/O error occurred.
chatter validate --audit FILE
An audit writes one JSON object per line to FILE, one record per
diagnostic (errors and warnings alike), and prints its summary to stdout:
{"file":"/path/to/bad.cha","code":"E502","severity":"Error","message":"...","line":12,"column":1}
{"file":"/path/to/ok.cha","code":"W110","severity":"Warning","message":"...","line":4,"column":7}
Contract notes:
file,code,severity,message,lineandcolumnare stable, and every record has all six.severityis"Error"or"Warning", spelled as the--format jsonrecords spell it. A warning does not fail its file; filter onseverityto read errors only.lineandcolumnare 1-based, andnullwhen the diagnostic has no line position.- A file with no diagnostic writes no record; the summary counts it.
- A file that cannot be written whole fails the run (exit 1).
chatter to-json
chatter to-json emits the full ChatFile JSON model rather than a diagnostic
summary. The authoritative contract for that output is the JSON Schema
documented in JSON Schema.
Practical notes:
- The JSON itself is the contract, not any validation status lines printed by the CLI.
- Use
-o/--outputif you want only the JSON in a file. - Use
--skip-validation,--skip-alignment, or--skip-schema-validationonly when you explicitly want to bypass those checks.
chatter cache stats --format json
Cache statistics emit one JSON object on stdout, tagged by database, which
says what is in the cache directory. The command only reads: it creates,
migrates and writes nothing.
{
"database": "current",
"total_entries": 743,
"cache_dir": "/Users/example/Library/Caches/talkbank-chat",
"cache_size_bytes": 274432,
"last_modified": "2026-03-09T13:05:31.000Z"
}
{ "database": "absent", "cache_dir": "/Users/example/Library/Caches/talkbank-chat" }
{ "database": "older_schema", "cache_dir": "/Users/example/Library/Caches/talkbank-chat" }
Contract notes:
databaseis"current","absent"(the directory holds no database; the command exits 0, as having no cache yet is a legal state) or"older_schema"(a database an older build left, which the nextchatter validateorchatter cache clearmigrates; its entries are not counted, since migrating can remove some). A database a newer build wrote fails the command (exit 1).absentandolder_schemacarry onlycache_dir.currentcarriestotal_entries,cache_dir,cache_size_bytesandlast_modified.cache_diris the cache directory as a string, ornullfor an in-memory cache, which has no directory. A directory whose path is not valid UTF-8 fails the command rather than being written with replacement characters.cache_size_bytesis the database file’s length in bytes, ornullwhen there is no database file (an in-memory cache, or a database removed between the command’s look and its read).0means a file that exists and is empty, never “no file”.last_modifiedis RFC 3339 in UTC with exactly three fractional digits (2026-03-09T13:05:31.000Z), so string order is time order. It isnullexactly whencache_size_bytesis, because there is no database file.- A database file that exists but whose size or modification time cannot be
read, including a modification time outside the years -9999 to 9999, fails
the command; it is never reported as a zero size or a
nulltime.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
What a Version Bump Promises
Status: Current Last modified: 2026-10-02 (commit 2d7e886b)
If you depend on chatter, two different things can move under you and they move independently:
- the Rust API, which decides whether your code still compiles;
- the validation verdict, which decides whether files you already have still pass.
A release can be perfectly source-compatible and still change which CHAT files
validate accepts. That is not a defect; it is the point of the project. But
it means the version number alone cannot tell you whether a bump is safe for
your corpus, so this page says what each half promises.
The Rust API follows SemVer
Ordinary Cargo expectations, with one qualifier: chatter is pre-1.0, so a minor
bump may break the API. Pin exactly (=0.10.0) if you need stability, or track
the minor and read the changelog.
The validation verdict follows a different rule
Any release may change which files validate, including a patch release. Validation is not a stable interface and will not become one before 1.0. The project exists to move the boundary of what counts as valid CHAT, and freezing verdicts would freeze that.
What you get instead is a guarantee that the change is announced:
Every release whose validation verdicts move opens its
CHANGELOG.mdentry with a bold Validation behaviour note, naming the codes that were added, removed, retired or made stricter, and what kind of file is affected.
If a release has no such note, its verdicts did not move. If it has one, read
it before upgrading a pipeline that gates on validate.
Which direction a change can go
Both, and they are not symmetric:
- Stricter (a new code, or an existing one reaching more inputs) means files that used to pass now fail. This is the common direction, and the failing files are usually genuinely wrong; the corpus they came from has simply not been cleaned yet.
- Looser (a code retired, or a false positive fixed) means files that used to fail now pass. Retired codes are never reused for a different rule.
Neither direction is a breaking change in the SemVer sense, because neither touches the API.
Retiring a code is not a clean removal
A code names a rule at a moment in time, and downstream records cite it: an adjudication log, a repair ledger, a review note all say “this edit was made because E754 fired”, and that remains true after the rule is withdrawn. So a retired code keeps two obligations. It is never reused, as above. And a consumer holding historical citations should NOT validate them against the current code set, because doing so makes a correct old record un-loadable over a rule that was withdrawn for reasons that have nothing to do with that record.
If you keep such citations, validate them as a closed list that includes the retirements you know about, so a typo still fails and a NEW retirement fails loudly until someone records why. That is the check worth having; checking against today’s live set is not.
Example: a ledger that cites E754 (LetterFormMultipleLetters, a retired
code) for a repair that is still correct: a digit zero typed for the letter o in 0@l. The rule went away
because it counted characters and a digraph is one letter written with two;
the repair it surfaced was right either way.
What to do about it
If you gate on validate in CI over a fixed corpus, treat a chatter upgrade
the way you would treat a linter upgrade: pin it, upgrade deliberately, and
review the verdicts over your own files rather than assuming. Chatter’s
acceptance evidence comes from reviewed specifications and a finite reference
corpus exercised through public workflows, including deliberate invalid
variants. A full production-corpus differential run is not a release gate.
New findings require CHAT-policy adjudication; matching another validator’s
diagnostic code, count or wording is not the acceptance criterion. See the
CHECK assessment
for its scope and reopening rules.
If you only consume the parsed model and never call validate, only the SemVer
half applies to you.
Why this page exists
An integrator pinning chatter found that the two halves were not distinguished
anywhere, and hit the case that makes the distinction concrete: a release that
was perfectly source-compatible (their adapter compiled unchanged, their whole
suite passed) while validation moved in both directions at once, one file
newly rejected in a bad corpus and one newly rejected in a good one. The
practice of announcing verdict changes already existed by then and had been
followed for several releases; it was simply not written down anywhere a
consumer would look.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).
Merge Override File Format
Last modified: 2026-10-02 (commit 2d7e886b)
Status: Draft
The merge override file is the typed, human-readable record of
operator decisions in the chatter speaker-id → structural assembly
pipeline. It serves three purposes:
- Persistence: operator adjudications made for one batch can
be replayed on later runs without re-prompting (
chatter speaker-id --override-file <FILE> --session-id <ID>). - Audit trail: each entry records who decided what, when, and on the basis of which Jaccard scores. Years later, a researcher can answer “why was PAR0 labeled INV in this session?” by reading the file.
- Interchange: an adjudication UI (CLI, future web app) and the batch pipeline share the same file format; UI tools can be added or replaced without changing the on-disk contract.
This page is the authoritative reference for the file’s schema.
For the usage contract (which commands read/write it, when, why),
see chatter speaker-id.
File location and naming
The file’s location is caller-chosen. The convention is one file per donor batch, named for the batch:
batch-2026-05-27-childes-eng.overrides.toml
batch-2026-06-15-fluency-pilot.overrides.toml
batch-2026-08-22-aphasiabank-bilingual.overrides.toml
Pipeline operators pass the path explicitly via --override-file;
no implicit search of a default location.
File format
UTF-8 TOML. The file has exactly one top-level key,
schema_version, followed by zero or more session entries, each
keyed by a session ID.
schema_version = 2
[<session_id_1>]
mode = "auto"
# ... fields per entry ...
[<session_id_2>]
mode = "explicit"
# ... fields per entry ...
The session ID is the table name. It is a free-form stable string,
typically the basename stem of the CHAT file the entry applies to
(s12-t1, Corpus2024-session-07, etc.). The TOML parser treats it
as a key; CHAT-conformant identifiers fit the unquoted-key grammar
and need no escaping, but any string is permitted if it conforms
to TOML key syntax (use quoted keys like "unusual_session-id"
if the ID contains non-bare-key characters).
Top-level fields
| Field | Type | Required | Meaning |
|---|---|---|---|
schema_version | unsigned integer | yes | The schema version this file conforms to. Currently 2. Readers refuse files with any other value. |
The reader refuses files with schema_version absent or
unknown, returning a typed error
(OverrideFileError::UnsupportedSchemaVersion). There is no
implicit version, no fallback, no auto-migration. Operators of a
file written by a newer version of chatter must upgrade their
binary; operators of a file written by an older version that the
current binary no longer supports must re-adjudicate. This policy
is documented in
architecture/merge-domain-types.md §6;
its rationale is to keep the schema honest and avoid premature
migration code that might silently misinterpret old data.
Per-session entry fields
Each [<session_id>] table contains the fields below. Required
fields must be present and well-typed; optional fields may be
omitted; unknown fields cause a parse error.
Required fields
| Field | Type | Meaning |
|---|---|---|
mode | string enum | One of "auto", "explicit", "override". How the decision was made; see “Mode semantics” below. |
adult_roles | table of donor code → inline table | The CHAT identity assigned to each speaker whose mapping action is "rename", keyed by that speaker’s donor code. Every "rename" key in mapping must have a matching key here. Each inline table has fields: code (string, CHAT speaker code), tag (string, CHAT role-tag), specific_role (string, optional, CHAT specific-role label such as First_Investigator, set only when two adults in the entry share tag). |
mapping | inline table | Map from input speaker codes to actions. Keys are speaker codes; values are "rename" or "drop". Every speaker that exists in the input CHAT file must appear in mapping. |
operator | string | Free-form identifier of the person who created the entry (username, initials, email prefix). Recorded as audit trail. |
decided_at | RFC 3339 datetime | When the decision was made. chatter writes a quoted string in New York time, whole seconds, with its offset ("2026-05-27T08:41:00-04:00"). It reads any RFC 3339 time with an offset (Z, +00:00, -04:00, any fraction), quoted or as a native TOML datetime; a time without an offset is refused. |
Optional fields
| Field | Type | Default | Meaning |
|---|---|---|---|
scores | inline table | {} | Per-speaker Jaccard scores recorded at decision time. Keys are speaker codes; values are floats in [0.0, 1.0]. Populated when the decision was based on a reference-mode auto attempt (even if the final mode is "explicit" because the operator overrode a low-confidence result). |
margin | float | absent | The decisive margin (winner-score / loser-score). Finite values serialize as TOML floats; the winner-takes-all case (loser score = 0, winner > 0) serializes as the TOML float inf; when no margin is meaningful (both scores 0) the field is omitted. (The shipped on-disk form is numeric; see override_file.rs and the merge-domain-types architecture page. Whether to switch the unbounded case to a string sentinel is an open contract question, not current behavior.) |
note | string | "" | Free-text operator note. Strongly recommended for "explicit" and "override" modes, captures why the operator made the call. |
flags | array of strings | [] | Operator-supplied flags marking unusual situations. Known values listed in “Flag vocabulary” below; unknown strings are preserved verbatim (treated as Custom). |
engine | string enum | "deterministic" | Which engine produced the decision. Always written on new entries; absent only in pre-provenance files, which read as "deterministic". One of "deterministic" (Jaccard reference-mode, spreadsheet, or operator adjudication) or "llm" (language-model judgment). |
judgment | inline table | absent | LLM audit trail. Present only when engine = "llm"; omitted for deterministic decisions. Sub-fields documented below. |
judgment sub-table fields
The judgment inline table records the audit trail for LLM-produced
decisions. It is present if and only if engine = "llm".
| Field | Type | Required | Meaning |
|---|---|---|---|
model | string | yes | Model identifier used for the judgment (e.g. "deepseek-v4-flash"). |
endpoint | string | yes | OpenAI-compatible base URL the judgment was made against. |
prompt_version | string | yes | Prompt-template version tag (e.g. "v2", the current template). Bumping this marks older entries as produced by a prior template. |
confidence | inline table | no (omitted when empty) | Per-field model confidence in [0.0, 1.0]. Keys are decision field names (e.g. "mapping", "roles", "merge_applicable"). Omitted entirely when no confidence values were reported. |
reasoning | string | yes | One or two sentence model rationale for the decision. |
Mode semantics
The mode field records how the decision was made and is
informational only at read time, every mode applies the same
mapping deterministically. Distinguishing modes matters for
audit purposes.
| Mode | Set when | Operator confidence |
|---|---|---|
"auto" | chatter speaker-id ran in reference mode, Jaccard margin was at or above --confidence-threshold, and the operator did not intervene. | High; the algorithm picked. |
"explicit" | The operator supplied --mapping directly, typically after a prior reference-mode attempt failed at the confidence threshold. | Operator made the call; confidence depends on what evidence they used (listening to audio, contributor data sheet, prior knowledge). |
"override" | The entry was created by reading a prior override file (replay). | Inherited from whichever prior decision the entry was first stamped with. The mode is updated to "override" whenever a replay re-writes the entry. |
The reader does not enforce mode → field correlations (e.g., it
does not require scores to be present when mode = "auto"). The
writer follows these conventions:
"auto"entries always includescoresandmargin."explicit"entries includescoresandmarginIFF a prior reference-mode attempt produced them; otherwise they are absent."override"entries preserve whateverscores,margin, andnotewere in the source file.
Mapping semantics
Each entry in mapping is one of:
"rename": the speaker is renamed per its own entry inadult_roles(looked up by the speaker’s donor code), toadult_roles[<donor_code>].codewith role tagadult_roles[<donor_code>].tag, and specific-role labeladult_roles[<donor_code>].specific_roleif present, in the output CHAT file. Every utterance for this speaker has its*CODE:prefix rewritten; the@Participantsentry for this speaker has its code + role-tag (+ specific-role label, if set) rewritten; the@IDrow’s code (field 3) and role (field 8) are rewritten."drop": the speaker’s utterances are removed from the output entirely. The speaker’s@Participantsentry and@IDrow are also removed.
Precondition. Every speaker that appears in the input CHAT
file must appear in mapping. There is no defaulting; omission
is rejected with
SpeakerIdError::SpeakerNotInMapping { speaker }. This is by
design: every decision must be explicit, so a future reader
knows that no speaker was silently passed through.
The reader rejects:
- Mapping entries whose key is not a speaker present in the input
(
SpeakerIdError::MappingSpeakerNotInInput). - Mapping values other than
"rename"or"drop"(TOML parse error from the typed deserializer).
Flag vocabulary
The flags array contains zero or more string values. The
following are recognized vocabulary; consumers MAY treat them
specially:
| Flag | Meaning |
|---|---|
"diarization-mixed" | The ASR diarization label being renamed actually contains multiple real-world speakers (e.g., clinician + parent collapsed). The rename is the best available approximation; downstream consumers should know the output is imperfect. |
"best-guess" | The operator could not confidently determine which speaker is which (e.g., from audio alone). The mapping is recorded as best-guess and merits review by a domain expert before publication. |
Any other string is preserved verbatim as a contributor-specific
flag (Custom(String) in the Rust type). Consumers SHOULD NOT
crash on unknown flags but MAY surface them in audit-trail
displays.
The order of flags within an entry is not semantically meaningful; duplicates are tolerated but considered noise. Tooling that modifies the list SHOULD deduplicate.
Reader semantics
OverrideFile::read_or_default(path) is the canonical reader (the
only public reader; used by chatter speaker-id --write-override). Its
behavior:
- If
pathdoes not exist, returnOverrideFile::default()(empty, current schema version). Otherwise: - Open
pathUTF-8. - Parse via
toml. - Refuse if
schema_versionis absent or not equal to the binary’sCURRENT_SCHEMA_VERSION(currently2). Error:OverrideFileError::UnsupportedSchemaVersion { found, supported }. - Parse all
[<session_id>]tables intoMergeOverridevalues; reject unknown fields. - Return
OverrideFile { schema_version, entries }.
OverrideFile::get(&session_id) retrieves a single entry;
returns None if absent.
Writer semantics
OverrideFile::write(path) serializes the file deterministically:
- Top-level field order:
schema_versionfirst. - Entries ordered by session ID alphabetically (
BTreeMapdefault). - Per-entry field order:
mode,adult_roles,mapping,scores,margin,operator,decided_at,note,flags,engine,judgment. - Optional fields omitted when empty / absent.
- Atomic replace: writes to
<path>.tmpthen renames over<path>to avoid leaving a partial file on crash.
chatter speaker-id --write-override <path> appends a single
entry: it reads the file (or starts empty), inserts/updates the
entry for the current session, and writes back. The session ID
defaults to the input CHAT file’s basename stem unless
overridden via --session-id.
Example: minimal auto-mode entry
schema_version = 2
[session-101-t1]
mode = "auto"
adult_roles = { PAR0 = { code = "INV", tag = "Investigator" } }
mapping = { PAR0 = "rename", PAR1 = "drop" }
scores = { PAR0 = 0.1931, PAR1 = 0.7347 }
margin = 3.81
operator = "alice"
decided_at = "2026-05-27T08:41:00-04:00"
The reader reconstructs: child speaker was PAR1 (high Jaccard
match with reference’s CHI); auto-decide succeeded with margin
3.81×; PAR0 becomes INV:Investigator in the output.
Example: operator-adjudicated entry
After a low-confidence refusal, the operator listened to the
audio, confirmed the call, and re-ran with --mapping:
[session-102-t1]
mode = "explicit"
adult_roles = { PAR1 = { code = "INV", tag = "Investigator" } }
mapping = { PAR0 = "drop", PAR1 = "rename" }
scores = { PAR0 = 0.6286, PAR1 = 0.3457 }
margin = 1.82
operator = "alice"
decided_at = "2026-05-27T11:15:00-04:00"
note = "Auto refused at 2.0× threshold. Listened to first 60 seconds; PAR0 produces child-content matching the hand transcript. PAR1 introduces herself as the clinician."
The scores from the prior auto attempt are preserved; the note captures why the operator was confident in the call despite the close margin. Years later, a researcher can verify by listening to the same 60 seconds and confirming the operator’s observation, the audit trail is reproducible.
Example: diarization-mixed parent sample
[session-103-t1-parent]
mode = "explicit"
adult_roles = { PAR0 = { code = "MOT", tag = "Mother" } }
mapping = { PAR0 = "rename", PAR1 = "drop" }
scores = { PAR0 = 0.3727, PAR1 = 0.6940 }
margin = 1.86
operator = "alice"
decided_at = "2026-05-27T11:22:00-04:00"
note = "Parent sample. Per contributor data sheet: mother. PAR0 contains clinician intro + parent mixed (Batchalign diarization limitation)."
flags = ["diarization-mixed"]
The flags = ["diarization-mixed"] warns downstream consumers
that the renamed MOT speaker is not a clean parent-only stream
the first ~15 seconds were the clinician giving setup
instructions before leaving the room. The note captures the
specifics for future review.
Example: replayed entry
The same file run on a different day from the override file:
[session-102-t1]
mode = "override"
adult_roles = { PAR1 = { code = "INV", tag = "Investigator" } }
mapping = { PAR0 = "drop", PAR1 = "rename" }
scores = { PAR0 = 0.6286, PAR1 = 0.3457 }
margin = 1.82
operator = "alice"
decided_at = "2026-05-27T11:15:00-04:00"
note = "Auto refused at 2.0× threshold. Listened to first 60 seconds; PAR0 produces child-content matching the hand transcript. PAR1 introduces herself as the clinician."
mode becomes "override" whenever the entry is re-applied by
reading the file. The other fields (including the original
operator and decided_at) are preserved, the override file
is the audit trail of the original decision, not of the
replay.
TOML grammar reference
For consumers writing the file by hand or generating it from other tools, the grammar is standard TOML 1.0 (toml.io) with the following domain-specific conventions:
- Datetimes use RFC 3339 with explicit time zone. UTC offset
Zand offsets like-04:00are both accepted. - Floats: standard TOML float syntax. The
marginfield accepts either a float or the string"unbounded". - Tables vs inline tables: top-level
[<session_id>]tables may use either standard or inline syntax; the writer emits standard tables for readability. - Comments: TOML
#line comments are permitted anywhere; the reader ignores them. The writer does not preserve comments across read-modify-write cycles (toml, nottoml_edit); hand-edited comments may be lost on subsequent--write-overrideruns. If preserving comments becomes important, the writer can be swapped fortoml_editin a future release.
Future schema changes
Schema version increments appear here under “Migration” with the
version-to-version diff and migration instructions. The policy is
strict refuse-with-clear-error on any schema_version value this
binary does not recognize; there is no auto-migration.
Migration: schema_version 1 -> 2 (adult_roles map, 2026-07)
schema_version bumped from 1 to 2 when the single per-entry
inserted_role field was replaced by adult_roles, a map from
donor speaker code to InsertedRoleSpec. The old field could only
name one CHAT identity per entry, so a session with two distinct
adult speakers (two different roles, or two speakers sharing one
role) had no way to record more than one of them. adult_roles
keys each InsertedRoleSpec by the donor code it applies to, so
every "rename" speaker in mapping gets its own role assignment;
InsertedRoleSpec also gained an optional specific_role field for
the CHAT manual’s First_Investigator/Second_Investigator-style
disambiguation when two adults in one entry share a role.
This is a breaking, non-migrating version bump: a
schema_version = 1 file is refused with
OverrideFileError::UnsupportedSchemaVersion, not auto-converted.
Operators holding a pre-bump override file must re-adjudicate those
sessions. The pending-adjudications.toml format bumped its own
schema version in lockstep for the same reason; see
Adjudication Workflow.
2026-06 additive fields: engine and judgment (no version bump)
The engine and judgment fields were added in 2026-06 to record
decision provenance (deterministic vs LLM). This addition did NOT
increment schema_version because both fields are backward
compatible in both directions:
- Old reader, new file: TOML
deny_unknown_fieldsis not set globally; older binaries that parse a file containingengineandjudgmentwill silently ignore the unknown keys. The decision itself (mode, mapping, adult_roles) is unaffected. - New reader, old file:
enginehas#[serde(default)]and defaults to"deterministic";judgmenthasskip_serializing_if = "Option::is_none"and is absent, which deserializes asNone. Pre-provenance files are therefore readable without error and are treated as deterministic decisions.
A future version bump would be warranted only if a change makes old files unreadable or misinterpretable, neither of which applies here.
Relationship to JSON Schema
The Rust OverrideFile type is implemented (in
talkbank-transform, src/speaker_id/override_file.rs) and drives
the override-file replay workflow today. What is not yet built is its
JSON Schema export: OverrideFile does not yet derive
schemars::JsonSchema, so no schema is generated, and the canonical
URL https://talkbank.org/schemas/v0.1/merge-overrides.json is
reserved but not yet published. Exposing it follows the same
schemars-based generator pattern documented in
JSON Schema.
The TOML form is the on-disk format; JSON Schema is the
machine-readable spec for external tooling. Both describe the
same OverrideFile Rust type.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).