Transform Pipeline
Status: Current Last updated: 2026-10-07 (commit 5e895791)
The talkbank-transform crate provides high-level pipelines that compose parsing, validation, and serialization into reusable workflows.
Core Pipelines
Transcript construction
build_chat assembles a typed TranscriptDescription into a mutable CHAT
document; validation remains a separate step. Media names accepted by
MediaFilename as HTTP/HTTPS references are preserved verbatim, including
quotes, extensions and Unicode spelling. They are not filesystem paths and
are never fetched. Complete input spelling must pass media representability
admission before local basename/extension reduction; malformed directory text
cannot be hidden by normalization. The resulting local stem is admitted again.
Parse + Validate
The most common pipeline: parse a CHAT file and validate it.
use talkbank_transform::parse_and_validate;
let result = parse_and_validate(source, &parser, &error_collector);
This:
- Parses the source text into a
ChatFileAST - Runs validation (alignment checks, header consistency, etc.)
- Collects all errors and warnings into the
ErrorSink
Source-bound preservation versus replacement
TreeSitterParser::admit_planned_tiers selects a tier-removal policy from the
headers of the same producing parse. It fully validates the retained document;
this does not certify original bytes when replacement was selected.
AdmittedReplacement::into_disposition consumes that result and distinguishes:
AdmittedDisposition::Preserved(AdmittedPreservation): no replacement was selected and no concrete tier was removed. The capability binds complete admission to the exact original source, including its formatting.AdmittedDisposition::Replaced(AdmittedReplacement): only the retained document is admitted. Even a selected replacement that matched no tier cannot manufacture original-source admission through this transition.
AdmittedSourceChat::from_preservation transfers the opaque preservation
capability directly to an unchanged-output proof without parsing again or
accepting independently supplied source/model arguments. Editing still consumes
admission and requires fresh checked construction before output. Recovery and
internal failures cannot produce either successful admission state.
TreeSitterParser::admit_word_timing_plan additionally supports
WordTimingPlan::PreferRetained. Its producer defers concrete %wor candidates
and their own diagnostics while lowering every other component normally.
Complete original admission retains valid timing, including partial timing and
legal stale sidecars; reusability is a separate downstream decision. If original
admission fails, the plan physically omits only the concrete word tiers the
evidence names (their own lowering failed, or a validation error lies inside
their source span), binds those diagnostics to each removal receipt, and
validates the reduced document again; errors located in no word tier remove
every remaining word tier. It requires complete retained admission before
returning a word-timing replacement receipt, never clears recovered-model taint
and never filters diagnostic codes. Internal producer failures cannot select
regeneration. book/src/architecture/wor-timing.md describes the loop.
Components parse and lower once. Word-tier entries move into the document
rather than being copied, and a rejected attempt returns the document with its
diagnostics (ValidationFailure::into_rejection), so no model copy is made;
removal finds each entry again by its own source span. WordTimingPlan::Preserve
gives no exemption and is appropriate for actual pass-through or preservation
policies.
The plan’s phases are types. Lowering consumes a PlanEntry, either already
decided or a header callback, and hands back the decided plan beside the
producing source, so nothing can route a tier or read a selection before the
decision. Each decided plan implements TierRouting: RetainAll (plain
parsing) and TierRemoval (nothing, or a fixed selection with its removal
receipts) never defer a tier, and their deferred type is uninhabited, so their
utterances carry no word candidates by type;
WordTimingDecision::PreferRetained defers word tiers and makes no removal
receipt during lowering. Removal receipts and deferred candidates therefore
never coexist in one plan.
stateDiagram-v2
[*] --> Decided: admit_replacing_tiers, plain parsing
[*] --> AfterHeaders: admit_planned_tiers, admit_word_timing_plan
AfterHeaders --> Decided: every header lowered, callback runs once
Decided --> RetainAll: plain parsing, every tier lowers
Decided --> TierRemoval: route removes selected domains
Decided --> WordTimingDecision: route defers word tiers
TierRemoval --> AdmittedReplacement: complete validation
WordTimingDecision --> AdmittedReplacement: Preserve, complete validation
WordTimingDecision --> WordTimingAdmission: PreferRetained, candidate loop
Source-bound utterance partitions
utterance_split::UtteranceSplitPlan selects morphology-domain or original-word
timing-domain assignments against one borrowed utterance. Execution accepts no
second model or assignment vector. It refuses count drift, disjoint child runs,
boundaries inside indivisible groups/replacements and separator-stranding
boundaries. Child-label magnitudes do not determine allocation sizes.
Execution returns SplitOutcome::Unchanged when every slot names one child, so
the caller keeps the utterance it holds and nothing is copied; only
SplitOutcome::Split rebuilds children.
The shared partition policy retains free-text dependent tiers on the first
child, preserves only count- and lexically corroborated %wor, and derives a
child’s main timing only from complete measured word timing. Original-turn
analysis and uncorroborated word timing are returned as explicit, source-bound
invalidation receipts, not merely logged. No full parent interval is assigned
to a partial child. Rebuilt children still need checked construction before
output; structural partition admission is not whole-file validity.
utterance_split::WordSpeakerSource admits complete source timing before
requesting acoustic inference and keeps its immutable source borrow through
that request. bind_timeline then establishes WordSpeakerSplitPlan, without
rematching timing or accepting another source. The convenience split-plan
constructor performs those same transitions when evidence already exists.
The split plan adds measured word-level speaker ownership to that same partition
owner. It requires complete, lexically corroborated %wor, uses
union-of-held-time attribution from rediarize, and refuses uncovered or tied
words. It never picks the nearest turn or carries a previous speaker through a
gap. Returning to a prior speaker creates a new contiguous child run rather than
reordering words. A single run is WordSpeakerPartition::Relabeled, which names
the owning track and leaves the source for the caller to relabel; only several
runs rebuild children. Execution retains ownership evidence and invalidation
receipts; participant identities, headers and final output admission remain the
caller’s responsibility. The existing whole-turn rediarization API and its
contestation reporting are unchanged.
CHAT → JSON
Convert a CHAT file to its JSON representation:
use talkbank_model::ParseValidateOptions;
use talkbank_transform::{JsonLayout, chat_to_json};
let json = chat_to_json(source, ParseValidateOptions::default().with_validation(), JsonLayout::Pretty)?;
The JSON follows the schema at schema/chat-file.schema.json, and is checked
against it before it is returned.
JsonLayout is Pretty (indented, the CLI’s default) or Compact (one line,
chatter to-json --compact, which chatter’s clap parsing turns into the
value). It was a pretty: bool, which every caller had to
spell as a bare true or false. chat_to_json_named adds the transcript’s
name (so E531 can run), chat_to_json_with_schema_policy adds a
JsonSchemaPolicy, and chat_to_json_unvalidated skips the schema check; all
four take the layout the same way, and the schema policy and the layout are
matched as one pair, so every combination has exactly one serializer.
JSON → CHAT
The JSON produced by chat_to_json is schema-conformant and
round-trips. Deserialize it back into a ChatFile with serde_json
(the model derives Deserialize), then serialize through WriteChat
to reproduce CHAT text:
let chat_file: talkbank_model::ChatFile = serde_json::from_str(json_str)?;
let chat_text = chat_file.to_chat_string();
The chatter from-json command wraps this path
(crates/chatter/src/commands/json.rs, json_to_chat).
CHAT → CHAT (Normalize)
Parse and reserialize to normalize formatting:
use talkbank_transform::normalize_chat;
let normalized = normalize_chat(source, &parser)?;
normalize_chat lives in
crates/talkbank-transform/src/pipeline/convert.rs.
Selective name pseudonymization: planning API under development
pseudonymize::NameMap admits a private caller-supplied mapping.
TranscriptNames::admit_document binds that mapping to parsed, validated source,
including alignment checks. PseudonymizationInput::plan_words creates sensitive
review previews without mutating the input. PseudonymizationInput::prepare_output
can produce an admitted in-memory PseudonymizedDocument, but this is
not yet a user-facing de-identification command. Private receipt persistence,
CLI integration and broader policy/corpus acceptance remain unfinished.
Output preparation refuses unsafe lexical, morphology, timing or pronunciation plans. It applies only their source-bound edit ranges, copying every intervening byte unchanged; applied private receipt entries are built during that same operation. The rewritten source must pass parsing and alignment-aware validation under the input’s filename and rule context. The output parse is reused for a follow-up plan, which must propose no further changes. A refusal retains original review findings but exposes no partial output text. Accepted output still is not a promise of complete de-identification: report-only findings remain visible in its private review and callers must protect both output and receipts.
Lexical reviews retain a producer-assigned WordLocation: zero-based utterance
and original-word indices, plus WordSpelling distinguishing spoken material
from an indexed editorial replacement target. Retraced words count; separators
and targets do not advance the original-word index. These are review coordinates,
not morphology or phonology alignment indices. Aligned morphology, corroborated
timing and pronunciation refusals carry the selecting word’s location through
their existing bindings. Applied lexical edits retain it in EditOrigin; their
EditKind is derived from that origin rather than stored independently. Multiple
source fields removed from one shortening retain the same word location.
Metadata findings and edits carry HeaderFieldLocation (header index plus typed
field); prose uses ProseLocation, which distinguishes a document header from
an utterance’s dependent tier. Header indices count all headers, including
structural ones; dependent-tier indices count all tier kinds. Independently
reported morphology uses LemmaLocation: utterance, morphology-item index and
LemmaPart::Main or a specific post-clitic. It never invents a main-tier match
for an unmatched lemma. Exact source ranges remain available alongside this
context. All indices are zero-based. Persistent CLI receipts are being integrated.
The CLI publication boundary is being implemented separately from the transform.
It stages only admitted PseudonymizedDocument bytes beside an explicit output
destination and uses no-clobber publication. Existing files, directories and
symlinks are not replaced; a collision after staging also refuses publication.
Abandoning preparation removes its temporary file, not the input or destination.
Unix staging files are owner-only; other systems require a protected destination
directory’s inherited ACL. This does not establish a cross-file transaction or
crash-durable directory update.
The receipt owner encloses the low-level publisher. Its PreparedPublication
owns both the staged file and a committed private SQLite receipt; callers cannot
separate that evidence from its output or call the enclosed publisher directly.
The receipt records source/output BLAKE3 identities, the exact applied edits,
typed locations and report-only findings in queryable tables, not JSON blobs.
It starts as prepared, becomes written only after publication and file sync,
or records publication_failed on a no-clobber failure. An interruption or failed
completion update can leave prepared; that state means uncertain publication,
not success or proof of absence. Existing receipt files are never overwritten.
SQL dependencies live in the CLI, not the transform core. The user command
remains unavailable pending command routing and end-to-end acceptance. Refused
plans have separate output_refused receipts: proposals are labeled proposed,
never applied, and output coordinates and identity are absent. Input admission
refusals and missing map entries are distinct states. A derived summary view
distinguishes a mapped document with no findings from an unmapped document.
The authored word-features/pseudonymizer-source.cha / pseudonymizer-expected.cha
pair tests exact output, every untouched gap, cross-tier changes, metadata,
possessive retention, deterministic application and idempotence. Existing
pronunciation, morphology and timing mismatch references also exercise whole-
document refusal. A placeholder that reintroduces a mapped word at a Unicode
boundary is refused during map admission, before document planning. Placeholders
must first be plain CHAT lexical tokens; invalid tokens are refused separately.
The output stability check remains an independent boundary safeguard.
Admission retains the producer-owned ParsedSource alongside the validated
model. TreeSitterParser::parse_chat_file_with_source returns both from one
parse; its diagnostics still require review and its model still requires
validation. The retained CST enables generated source-bound field traversal
without a second parse, raw-line searches, or assuming whole-header serialization
preserves untouched formatting. CST ownership alone is not a validity proof.
Admission refuses both errors and warnings. Canonical controls exercise empty
input, missing participant roles, forbidden controls inside free text, and a
warning-only non-NFC media name paired with its canonical spelling. Even rejected
empty input can retain a CST; its existence never authorizes transformation.
Header-field review covers whole Unicode-bounded names within participant names and
typed @ID group, education and custom fields, including repeated names within
longer metadata text. It uses the same matcher and report-only possessive
policy as prose, not a second whole-field replacement path. Generated field
projections supply exact source ranges; a binding failure is a refusal, not
permission to search raw lines. ParticipantWordRoles owns the name/role
partition for both parser lowering and source-bound review. Speaker codes,
participant roles, spacing and structural pipes are not selected. Each match
has exactly one metadata or prose owner; these are still review findings, not
complete writable pseudonymization.
The current planning policy uses exact-case typed lexical components, reports
case near misses, and derives dependent-tier changes from selected main-tier
words. Main-tier previews retain exact generated text/shortening edits;
each shortening owns its parentheses and inner segment, so it cannot create
overlapping edits. Marker bytes, suffixes, annotations and unselected compound
partners remain outside those ranges. Source association and lexical-piece
corroboration are required: failure creates a refusal instead of an unbound
proposal. The word-features/selective-name-fields.cha reference exercises
lengthening and names on either side of a compound boundary.
Free-text planning uses unicode-segmentation word boundaries on typed prose
payloads, with apostrophe-s possessives such as Rose’s and Rose's reported
rather than partially rewritten. A mapped name may span several segments, such
as Rose-Marie. At each boundary the longest recognized spelling takes
precedence, including report-only near misses and possessives: mapping both
Rose and Rose-Marie never partly rewrites Rose-marie as Rose.
The headers/unicode-name-boundaries.cha reference exercises this policy in
participant names, ID metadata and prose. LexicalPlan::free_text returns private
source-bound exact matches, case near misses and possessive review findings.
The CST supplies prose segments within descriptive headers and free-text
dependent tiers (including user-defined tiers); timing bullets, identifiers,
configuration headers and structured tiers are not treated as prose. This does
not tokenize raw CHAT structure. Header-field review above reuses this matcher
while retaining its own typed metadata classification.
The headers/free-text-names.cha reference covers accented names, punctuation,
continuation lines, timing, possessives and several dependent-tier kinds.
Dependent-tier proposals are driven by the selected main-tier
words. Morphology requires established lemma correspondence. Edit locations
are not inferred from text searches: each morphology proposal retains
the generated, source-bound main-lemma field from the same admitted parse.
Tier item counts and lemma values corroborate the association; failure refuses
the proposal. Post-clitic fields are not mistaken for subsequent main items.
The canonical tiers/mor-name-source-fields.cha example exercises repeated
names, both shortening forms, uneven spacing and a preceding post-clitic item.
This establishes exact edit locations, not writable document output.
Unchanged mapped lemmas, including post-clitics, are reported separately as
exact matches or case near misses; the lemma scan cannot authorize an
independent replacement. Compound
components without proven lemma-part correspondence produce a refusal.
Timing uses the %wor corroboration transition. Words within a
timing tier receive the selected main word’s existing component decisions, not
another name-map lookup. Equal cleaned spellings with different compound
structure refuse the change. Exact timing-word lexical edits preserve markers
and leave timing bullets outside their ranges; a failure refuses the tier’s
proposal rather than exposing a partial list. The selective-name reference
also covers a name that occurs only in the timing tier: it is not independently
replaced.
Words within a phonological group share the group’s alignment position.
Pronunciation findings retain the
original %pho/%mod item and refuse automatic output; structured Phon
companions require correspondence review. No tier is silently dropped and no
orthographic placeholder is treated as a pronunciation. These review objects
contain protected transcript content and must not be logged by default.
Validation + Roundtrip Cache Lifecycle
The following diagram shows the full validation and roundtrip pipeline, including the cache layer:
flowchart TD
file["CHAT file"]
cache{"Cache\nhit?"}
parse["Parse\n(tree-sitter → AST)"]
validate["Validate\n(per-file → per-utterance →\nmain tier → dependent tiers)"]
rt{"Roundtrip\nflag?"}
ser1["Serialize → CHAT text"]
reparse["Reparse CHAT text"]
ser2["Serialize again"]
cmp{"Two\nserializations\nmatch?"}
store["Store in cache\n(SQLite)"]
pass["Pass"]
fail["Fail"]
cached["Return cached result"]
file --> cache
cache -->|miss| parse --> validate --> rt
cache -->|hit| cached
rt -->|yes| ser1 --> reparse --> ser2 --> cmp
rt -->|no| store --> pass
cmp -->|yes| store
cmp -->|no| fail
Streaming Parse
For large files or interactive use, the transform crate supports streaming parse where utterances are processed incrementally rather than loading the entire AST into memory.
The shared validation runner (every frontend, one engine)
All bulk validation, whatever the frontend, flows through the
validation_runner module’s two streaming entry points in
crates/talkbank-transform/src/validation_runner/:
validate_directory_streamingwalks a directory and feeds every CHAT transcript to a worker pool;validate_files_streamingruns an explicit file list through the same worker pool.
The one worker pool and the one walk
The pool is talkbank_transform::worker_pool::fan_out, and it is the only
pool: chatter to-json fans its directory conversions out through it too.
flowchart LR
items["items (caller's iterator)"] -->|"calling thread feeds"| queue["bounded queue, 2 x width"]
queue --> w1["worker 1<br/>16 MiB stack"]
queue --> wn["worker n<br/>16 MiB stack"]
w1 --> join["join every worker"]
wn --> join
join --> run["PoolRun { results, outcome }"]
- Width is
--jobs, else the machine’s parallelism, never zero; serial work is a pool of width one, not a second code path. - Every worker runs on a
CHAT_THREAD_STACK_BYTES(16 MiB) stack, the size of the CLI’s program thread, because parsing and validation recurse with the data. - Each worker RETURNS its result (the validation runner’s per-worker tally,
to-json’s counts); the caller combines them after the join. Nothing reads
shared counters whose exactness rests on a comment. Workers borrow what
they share (the event sender, the cancellation latch, the cache and the
configuration) through a
WorkerContextof references, because the pool’s threads are scoped: noArcclones, no per-worker config copy. - A worker that unwinds is joined and counted as
PoolOutcome::SomeUnwound; a thread that cannot be started isPoolOutcome::CouldNotStart, with nothing fed. The caller measures what was lost against what it fed. - Stopping early is the caller’s iterator (the runner feeds
take_while(not cancelled)); workers stop by returning.
Discovery is talkbank_transform::paths. Every walk descends every level.
walk_files returns every file found WITH its path relative to the walked
root (a FoundFile), built as the walk descends, plus every entry that could
not be read. keep is given each entry’s file name, so a rejected entry
never has its path built. walk_transcripts returns each transcript as a
FoundTranscript: its relative path and its StoredTranscript, whose name
is the listing entry that found it, so nothing lists the directory again to
learn the stored name (a stem that is not UTF-8, which validation could not
name, is a failure of the walk). What happens at a symbolic link is the
caller’s choice, a Links value:
Links::Follow(every reading walk): a link is walked as what it points to, a directory reached twice is walked once, and a link whose target is gone is a failure whatever its name: nothing says whether it was a file or a directory, and a link to an unmounted volume’s subcorpus hides every transcript under it.Links::Skip(to-json --prune, the one walk whose results are deleted): links are neither walked nor kept. Following them would let--prunedelete JSON outside--output-dirthrough a linked directory, and then remove that directory’s emptied parents.
A directory the walk cannot list is always a failure, never skipped, even an
operating system’s own folder at a volume root (.Spotlight-V100,
System Volume Information, lost+found): nothing can say whether it held
transcripts. Walk the corpus directory, not the volume root.
expand_transcript_arguments is the one expansion of command-line paths:
each file argument is resolved to its stored name by one
StoredNameResolver (each parent directory listed once), and each directory
argument contributes its walk’s transcripts, so the result is
StoredTranscript values. The runner’s work queue carries those values,
and validate_files_streaming resolves a plain path list the same way
before any worker starts, so no worker resolves a name. Nothing is skipped
silently, and there is one policy for what cannot be
read: it is a FileStatus::ReadError in the run’s own results, counted in
its totals, so the run fails and says which paths. The directory entry
point does this for its walk, and validate_arguments_streaming for
command-line arguments (chatter validate), so the unreadable path reaches a
JSON consumer as a record. Commands that are not streaming runs (fix,
to-json, the debug commands) refuse an incomplete input before
processing anything.
Both share one worker loop, so every consumer gets identical rule
coverage (including the file-stem-dependent checks such as the @Media
filename match), identical stats accounting, and the same on-disk cache.
The chatter CLI, the TUI, and the desktop app all call these
entry points. The invariant to preserve: no frontend grows its own
validation orchestration; a file must validate identically whether
selected alone or reached by a directory walk.
What a run checks: typed, end to end
ValidationConfig says what each worker checks with typed values, never
booleans: alignment: AlignmentValidation (Structure or
IncludeTierAlignment) and roundtrip: RoundtripCheck (Skip, the default,
or Run). A frontend parses its flags into these once, at its boundary:
chatter’s clap arguments are Flag<AlignmentValidation> and
Flag<RoundtripCheck> (see “Flags parsed into modes” below), and the desktop
translates its request’s checkbox where it deserializes it. The CLI’s
ValidationRules, the TUI and the desktop runner then carry the same values
the worker matches on; the libraries have no bool-to-mode constructors, and
there is one roundtrip type from flag to worker.
The CLI’s side: one presentation, one renderer, one ending
chatter validate resolves its output flags once, in dispatch, into a
ValidationPresentation, then runs one event loop:
flowchart TD
flags["--format, --quiet, --audit, --tui-mode"] --> resolve["ValidationPresentation::resolve\n(conflicts are usage errors)"]
resolve -->|Tui| tui["TUI loop\n(same runner, same error limit)"]
resolve -->|"Streamed(Lines / Json / Audit)"| renderer["one ValidationRenderer"]
renderer --> loop["event loop: discovering, started,\none FileComplete per file (status + diagnostics)"]
loop --> end["the run's RunEnding\n(Aborted(NoEnding) if the stream closed without one)"]
tui --> end
end --> finish["renderer.finish(&RunEnding): exhaustive"]
end --> exit["ValidationOutcome::failed(): !RunEnding::passed()"]
The runtime holds no output channel of its own: every fact a run has to say
(a stop, a loss, an abort, an input with nothing in it) is a RunEnding the
renderer’s finish matches, so JSON mode can keep stderr empty by
construction rather than by remembering to. The TUI shows the same ending
and returns it with the session (InteractiveEnd::Closed(RunPhase)), so
the exit status is the run’s on every surface: only RunEnding::passed
exits 0, and a session closed before its run ended fails. The TUI lists
every file that failed (its diagnostics, or why it could not be read, why
its roundtrip failed, or the tool failure) and shows the run’s notices and
cache events, as the other surfaces print them.
Flags parsed into modes
A presence flag selects one of two values of a typed mode. chatter’s
cli/args/flag_modes.rs has one generic clap argument, Flag<M>, for any
M: FlagMode (the flag’s name, help, optional short form, and the value
when it is absent or present), so a command’s field is
#[command(flatten)] alignment: Flag<AlignmentValidation> and its handler
receives the mode. FlagMode is chatter’s own trait, so it can be
implemented for library types (AlignmentValidation, RoundtripCheck,
JsonLayout, JsonSchemaPolicy) as well as chatter’s (FixMode,
ClearMode, CacheRefreshMode, JsonRefresh, OrphanJson,
AlignmentView). The translation from flag to mode therefore lives in the
CLI, and the library crates expose no constructor from a bool.
Two flags that select one value have their own Args impls, which refuse
the combinations that would leave one flag without effect: ToJsonCheckArgs
(--skip-validation conflicts
with --skip-alignment) and NormalizeCheckArgs (--skip-alignment
requires --validate), each yielding one CheckLevel.
How a run ends, and who decides
While its receiver remains connected, every stream ends with exactly one
ValidationEvent::Finished(RunEnding), and the runner decides which:
Complete(stats): at least one file was discovered and every one was accounted for. The only basis for a claim about the whole input, though still not “all files valid”:RunEnding::passedadds that no file was invalid, unreadable or a tool failure. A run that was told to stop after its last file had already been taken also ends here, because nothing was left to stop.NothingFound: the input named no transcript and nothing unreadable. It has no counts. The runner’sNonZeroUsize::new(total_files)is the one place this is recognised, so every snapshot’stotal_filesis non-zero.Stopped { stats, reason }: the run was told to stop and leftstats.missing_files()(non-zero) files unvalidated.reasonis aCancelReason:ErrorLimit { limit }(the run’s own error limit) orRequested(the caller’sCanceller). ItsDisplayis the one wording every surface uses.Incomplete { stats, cause }: the run reached its end without covering everything it discovered, and no requested stop explains it.causeis aLossCause:WorkerFaults, every worker failure the run observed (the pool’s ownPoolFault,UnwoundorThreadRefused, andParserUnavailable, none hiding another), orUnexplained, a runner defect.statsdescribes only what was processed.Aborted(reason): no totals at all. A drop guard on the orchestrating thread sendsAborted(Panicked)during an unwind, so a panicking run terminates its stream instead of closing it in silence; a consumer whose stream closes with no ending anyway reportsAborted(NoEnding).
passed() is the single answer to “may this run be reported as a success”:
the CLI’s exit status, the TUI’s header color and the desktop’s all-valid
claim all read it. Counted endings require producer-admitted CompleteStats
or PartialStats, each exposing read-only snapshot() counts. A snapshot
cloned from a stopped run cannot construct Complete, even when every
processed file passed. The partial payload owns its derived, nonzero
missing_files() count; callers cannot supply a contradictory shortfall.
RunEnding::stats() retains the common read-only snapshot view. Text, JSON
and desktop event formats are unchanged; Rust consumers matching an ending
use the admitted payload’s accessors instead of the removed count fields.
The ending of a run that reached its end is decided from three facts, each from its owner:
flowchart TD
cov{"stats.coverage()"}
cov -->|Complete| complete["Complete(stats)"]
cov -->|"Shortfall(partial)"| faults{"WorkerFaults::observe"}
faults -->|"Some(faults)"| faulted["Incomplete { cause: WorkerFaults }"]
faults -->|None| stop{"latched stop?"}
stop -->|"Some(reason)"| stopped["Stopped { stats: partial, reason }"]
stop -->|None| unexplained["Incomplete { cause: Unexplained }"]
A worker fault outranks a stop: files a worker took and never finished are
lost, and a worker that could not create its parser took none, whatever else
happened. Each worker returns its tally or a WorkerSetupFailure, so a
worker that could not start is a value the runner sees rather than an empty
tally indistinguishable from an idle worker. WorkerFaults::observe, the
one constructor, reads the pool’s outcome and those setup failures and
answers None when every worker started and returned, so a WorkerFaults
is never empty.
Stopping a run: the Canceller and the error limit
There are two ways to stop a run early, and both end in the same latch:
- the caller’s
Canceller(returned by both streaming entry points; the CLI’s Ctrl-C handler, the TUI’sckey and the desktop’s Cancel button hold one), whose only operation iscancel(); - the run’s own
ValidationConfig::error_limit, anErrorLimit(UnlimitedorStopAfter(n)), which the runner counts.
The error limit is the runner’s, not a consumer’s. Each worker spends
FileStatus::errors_found() from a shared ErrorBudget before taking its
next file: the file’s Severity::Error diagnostics after suppression, or 1
for a failed roundtrip. Warnings never count. The addition that reaches the
limit latches CancelReason::ErrorLimit, so the worker that crossed it stops
at once and the others at their next file. With one worker the stop is
exact: a limit of 1 stops before the second file starts.
The count lives in the runner, not in a consumer reading rendered events,
so warnings cannot spend the limit, a run that covered every file is never
reported as stopped, and the text, JSON, audit and TUI presentations of
chatter validate all honour the same limit. The desktop app sets no limit.
ValidationEvent and RunEnding are deliberately NOT #[non_exhaustive].
Adding a variant breaks external consumers on purpose: a new ending that a
consumer silently ignores is precisely the defect these types exist to
prevent, so a downstream crate gets a non-exhaustive match error and decides
for itself what the new ending means. The compiler-checked
validate_directory_streaming rustdoc example shows an exhaustive event
loop. Keep the Canceller alive while consuming the stream (dropping it is
not a cancel), and treat channel closure without an ending as
Aborted(NoEnding), never as successful validation.
Dropping the result receiver is also a stop boundary. Once a worker fails to deliver a file-completion event, it must not start another queued file. This applies to fresh validation, cached results and file-read failures alike. The already completed attempt is accounted for before delivery; disconnection does not turn an unreadable input into a parser error or a cacheable result. With multiple workers, other in-flight attempts may finish before observing closure. The reference-backed lifecycle contract synchronizes at a cache lookup and checks subsequent work counts, rather than relying on sleeps or timing.
Two design points worth keeping:
- Every short ending is a VARIANT, not a field. A
lost: usizeor acancelled: boolbeside the totals would be something every consumer must remember to check, and forgetting yields a false clean bill of health: files abandoned by a crashed worker contribute to no counter, so partial totals look immaculate. A 500-file corpus could validate 480 and report “all valid”. - Loss is DERIVED, not counted.
ValidationStatsSnapshot::coveragereconcilestotal_filesagainst the per-file counters in one place, so there is no third counter free to drift from the two it reconciles. Whether a shortfall was requested is not a count: the snapshot reports onlyRunCoverage::{Complete(complete), Shortfall(partial)}, and the runner decides stop versus loss from the latch and the pool. A requested shortfall is not lost data, and reporting it as such would make the incompleteness report routine, and therefore ignored.
Caching
The transform layer integrates with a file-system cache. Validation results are keyed by content hash, so unchanged files skip re-validation. Cache location is platform-specific: ~/Library/Caches/talkbank-chat/ (macOS), ~/.cache/talkbank-chat/ (Linux), %LocalAppData%\talkbank-chat\ (Windows).
Use --force to bypass the cache for specific paths.
Error Collection
Pipelines use the ErrorSink trait for error reporting. Callers can provide:
- A collecting sink (gathers all diagnostics for batch output)
- A printing sink (writes diagnostics to stderr in real-time)
- A custom sink (for LSP diagnostics, JSON output, etc.)
This page last changed: 2026-10-07 (commit 5e895791). The whole book last changed: 2026-10-07 (commit 5e895791).