Parser Backends
Status: Current Last updated: 2026-10-02 (commit 2d7e886b)
TalkBank has two CHAT parser implementations. Both implement the ChatParser
trait and produce the same ChatFile model type, not necessarily identical
values or recovery results.
The --parser flag selects the backend at the CLI boundary; everything
downstream consumes the shared model API. Backend differences remain observable
in supported input, diagnostics and recovery:
flowchart TD
cli["chatter validate --parser <backend>\n(ParserBackend enum,\nchatter cli_types.rs)"]
sel{"which backend?\n(ParserKind,\ntalkbank-transform\nvalidation_runner/config.rs)"}
ts["TreeSitterParser\n(talkbank-parser:\nGLR, incremental)"]
re2c["Re2cParser\n(talkbank-parser-re2c:\nre2c DFA + chumsky)"]
trait["ChatParser trait\n(talkbank-model\nparser_api/chat_parser.rs)"]
model["ChatFile\n(shared model type;\nbackend-specific recovery)"]
cli --> sel
sel -->|"tree-sitter (default)"| ts
sel -->|"re2c"| re2c
ts -->|"ParserDispatch::TreeSitter\n(worker.rs) implements"| trait
re2c -->|"ParserDispatch::Re2c\n(worker.rs) implements"| trait
trait --> model
ParserDispatch::new(kind) (in validation_runner/worker.rs) is the single
place that constructs the chosen backend from a ParserKind; both variants
wrap a ChatParser implementor, so the validation runner never branches on
backend again.
The shared ChatParser trait
Backend compatibility mandate
This section is the authoritative policy for re2c/tree-sitter compatibility. Backend measurements and regression baselines record evidence; they do not override this contract.
Both implementations must pursue the same independently justified CHAT syntax, meaning and validity rules idiomatically within their own architectures. Neither parser is an oracle for the other. For supported valid input, the goal is semantic agreement in the shared model, not identical internal representation.
For developers, these are mandatory constraints:
- Use re2c lexer states, rich tokens, parser-owned recovery and typed, source-bound evidence. Use typestate, ownership and validated constructors to preserve admission and recovery distinctions.
- Do not emulate tree-sitter’s CST, MISSING nodes, recovery traversal or recovered model shape merely to satisfy an equality test. Do not add a parallel parser, reparse reconstructed text, or scan raw input after parsing to manufacture another backend’s diagnostics.
- Report the fault the available evidence supports, at a truthful source location. Preserve useful content where sound, but never fabricate valid structure, discard faults silently, or treat a recovered result as valid. A broader honest diagnostic is preferable to invented specificity.
- Adjudicate differences as a genuine semantic defect, an acceptable architectural difference, or an explicitly unsupported experimental feature. Retain valid controls and invalid-input rejection tests. Do not weaken a CHAT rule or change a baseline solely to make a comparison pass.
- Exact diagnostic codes, wording, counts, ordering, rejection stage and recovery output are not cross-backend requirements. Fix an inaccurate diagnostic for its user impact, not because the other backend differs.
For users switching --parser:
| Aspect | What to expect |
|---|---|
| Supported valid CHAT | The intended meaning and shared-model semantics should agree. A disagreement needs investigation. |
| Invalid input | Diagnostic detail, number, order and highlighted recovery region may differ; identical error reports are not promised. |
| Recovered content | Partial models and retained fragments may differ. Recovery is not validity and is not a portable data-repair contract. |
| Experimental limitations | re2c is incomplete and may reject supported CHAT or miss invalid input. A clean re2c result is not a substitute for default tree-sitter validation. |
Use the default tree-sitter backend for production validity decisions and the LSP. Consumers must not depend on cross-backend diagnostic identity or interchangeable recovered output. Report concrete validity or semantic disagreements with a minimal input and parser version; a diagnostic difference alone does not establish a bug.
Before 1.0, re2c improvements are welcome when justified and affordable, but remain lowest priority. Full backend equivalence is not a release prerequisite; re2c may remain explicitly experimental and incomplete. This does not waive truthful diagnostics or the prohibition on architecture-distorting workarounds.
API shape
Both backends implement talkbank_model::ChatParser directly (the
tree-sitter impl landed 2026-07-24 in
talkbank-parser/src/api/chat_parser_impl.rs; the re2c impl has carried it
from the start). The trait is the parser-agnostic API for every
granularity: whole files, headers, utterances, main tiers, %mor/%gra
and the other dependent tiers, down to single words and relations. Each
method takes (input, offset, errors) and returns a ParseOutcome;
diagnostics stream through the caller’s ErrorSink.
Downstream consumers should bind on the trait, not on a concrete backend:
fn analyze<P: ChatParser>(parser: &P, text: &str) { /* ... */ }
selects the backend with one generic bound, including cross-target setups
(tree-sitter natively, pure-Rust re2c on wasm, where compiling
tree-sitter’s C runtime is undesirable). No facade or cfg-gated dispatch
module is needed on the consumer side. The wasm half of that contract is
pinned in CI: the wasm job in ci.yml checks talkbank-model and
talkbank-parser-re2c for wasm32-unknown-unknown on every push.
Two notes on the trait’s shape:
- The trait has generic methods (
errors: &impl ErrorSink), so it is not dyn-compatible; runtime backend selection uses a small enum such asParserDispatchrather thanBox<dyn ChatParser>. - On
TreeSitterParser, every trait method delegates to the matching inherentparse_*_fragmentmethod, so trait-path and inherent-path behavior are identical by construction. The conformance gate istalkbank-parser/tests/chat_parser_trait.rs.
TreeSitterParser (default)
For isolated words, parse_word_with_context (or the inherent
parse_word_fragment_with_context) applies the enclosing document’s effective
@Options: CA to the admitted typed word. A standalone parenthesized word then
has the same CA-omission interpretation as whole-document parsing; mixed lexical
shortenings remain shortenings. This transition preserves raw spelling and
source coordinates and does not retry rejected input. The context-free word API
does not infer file options: callers that know them should supply the context.
- Crate:
talkbank-parser - Technology: tree-sitter GLR parser
- Grammar:
grammar/grammar.js→ generated C parser - Strengths: Incremental reparsing (LSP), robust error recovery (GLR), CST-level diagnostics
- Weaknesses: Slower on batch workloads,
!Send + !Sync(one parser per thread)
Used by the LSP, the default CLI, and all production validation.
Re2cParser
- Crate:
talkbank-parser-re2c - Technology: re2c DFA lexer + chumsky parser combinators
- Grammar: Translated from
grammar.jsrules → re2c conditions + chumsky combinators - Strengths: 4-8x faster,
Send + Sync, zero constructor cost, independent experimental implementation - Weaknesses: No incremental reparsing, incomplete diagnostic parity, and it is not ready to judge CHAT validity (see below)
Used for parser parity testing and performance benchmarking.
Source ownership and participant recovery
Parsed values borrow the caller’s source; token storage and temporary
recovery buffers are released after parsing. The backend leaks no memory
(no Box::leak).
File parsing receives a LexedSource that privately owns tokens and their
lexer locations alongside the borrowed source. Its only constructor lexes
that source, preventing callers from pairing unrelated token and location
arrays. Participant lists consume those located tokens through one parser
shared with the fragment entry point:
flowchart LR
source["Source text"] --> lexed["LexedSource: tokens and locations"]
lexed --> parser["Participant list state machine"]
parser --> entries["HeaderParsed::Participants: recovered entries"]
parser --> errors["ErrorSink: located diagnostics"]
entries --> model["Header::Participants"]
The list distinguishes its initial state, a nonempty entry, and a consumed comma awaiting another entry. A trailing comma therefore reports E550 while preserving the preceding participants. Conversion receives parsed entries instead of reparsing raw header tokens, and header fragments forward the same diagnostics with the caller’s offset. The internal AST snapshot records this distinction; it does not define a serialized CHAT format change.
The re2c newline token represents one LF, CRLF or lone CR, matching the canonical grammar. It does not fuse consecutive breaks, so blank-line structure is kept. Source-aware file dispatch reports an unconsumed blank newline at its lexer span. Generated error fixtures preserve their exact line-ending bytes in Git; published Markdown normalizes display line breaks and labels that presentation.
Annotation categories survive conversion
Token classification produces ParsedAnnotation::Scoped(ScopedAnnotationParsed)
for annotations that decorate content. Retraces, replacements, language codes
and postcodes remain distinct outer variants. The model converter accepts only
ScopedAnnotationParsed and returns a ContentAnnotation directly: structural
markers cannot enter that conversion and be silently discarded through None.
Replacement lookup likewise returns its payload rather than an index requiring
a second match or an unreachable branch.
The file-level E757 spacing check uses the same classified categories for
closing annotations and retraces. It reads adjacent tokens from LexedSource
and reports the following word’s complete lexer span when the code is glued to
that word. The specification includes both glued examples and a spaced control;
the cross-backend gate checks those generated cases. The same located-token pass
rejects replacements glued to rich or reconstructed words with E375/E316, matching
word_with_optional_annotations and CHECK 161. The bracket-location boundary
test compares both backends against the violation and its spaced legal control.
Canonical closing-bracket recovery excludes absorbed trailing whitespace from
its highlight and builds its context from the original source. The internal AST snapshot
changes to show the category, while reference-corpus model equivalence guards
serialized CHAT behavior.
Postcode admission and diagnostic offsets
A PostcodeToken owns its lexer’s full span and a private payload state:
nonempty content after trimming trailing whitespace, or recoverable missing
content. The lexer preserves leading payload whitespace. TierBody::postcodes
contains only these tokens, so lowering cannot accidentally treat another token
kind as a postcode. Missing content emits E363 and contributes no model postcode;
valid content retains its source span. Main-tier and utterance lowering require
an error sink explicitly.
The re2c trait implementation streams diagnostics through one offset adapter. File diagnostics, utterance fragments, main-tier lowering and header/participant fragments therefore use the same rebasing operation as their models. This removes temporary diagnostic collection in header fragments and the discarded utterance diagnostics. Remaining fragment entry points that do not yet produce diagnostics are still a separate parity gap. The postcode boundary test loads the authored E363 examples and checks actual token spans, nonzero offsets, recovered tier content and canonical-parser normalization.
Morphology admission and recovery
The morphology lexer distinguishes the stricter first lemma character from its
continuation characters and requires content after each feature separator.
The parser splits an admitted token without inventing an empty lemma fallback.
On failed %mor parsing, RejectedMorTier retains raw tokens and reports E600
at construction, alongside the primary syntax diagnostic. It cannot convert to
a model tier or masquerade as an unsupported dependent tier. Utterance lowering
retains morphology taint so alignment does not treat the dropped tier as clean.
The authored E316 examples and a legal lemma/feature control exercise this path.
This does not change the grammar’s allowance for angle brackets inside a lemma;
it rejects the forbidden leading angle bracket shown by the source examples.
Dependent-tier prefix admission
The lexer distinguishes a complete TierPrefix, including its required colon
and tab, from an IncompleteTierPrefix recovered from a label. File dispatch
recognizes both forms. Dependent-tier recovery consumes this classification
rather than inferring malformed syntax from an empty body and a suffix check.
It reports E602 over the original complete line and retains the recovered
content for inspection. A complete prefix with no body remains a separate
content-validation question (E756). The E602 boundary test loads both malformed
specification examples and the valid colon-tab control and checks source spans.
Separator provenance belongs to lexical admission
PrefixToken owns the matched payload and separator provenance. Prefix lexer
rules consume spaces after the required tab, so those spaces never become
header or tier content. The same carrier covers ordinary headers, embedded
speaker headers, dependent prefixes, and main-tier separators. AST header lines,
main tiers and dependent entries retain the admitted TierSeparator through
model lowering. Both backends use the shared file validator and its CA
policy; there is no main-tier-only whitespace scan or separate CA probe. Separator spans are omitted from serialized AST/model metadata, and
CHAT serialization writes the canonical tab in both CA and non-CA files.
The source-spec boundary test covers all E758 examples, padded CA headers and tiers, canonical non-CA controls, exact byte spans and nonzero source offsets. It compares canonical serialized CHAT with tree-sitter. The lexer prefix payload and header/dependent-entry AST shapes change; the file inspection snapshot records that API change.
Lengthening counts preserve their source
The grammar admits a nonempty run of colons with no 255-character limit.
WordLengthening::count and the re2c AST carry NonZeroUsize, measured at the
parser boundary. Neither backend narrows source length to u8, so a run of
any length neither overflows nor silently wraps. A zero count cannot be constructed or decoded
from JSON, and both the default constructor and omitted JSON count mean one
colon. Serialization does not repair zero counts with max(1).
The Rust count field and with_count argument are NonZeroUsize. JSON retains the integer count field, omitted for one colon,
but accepts longer runs and rejects zero. The schema describes that boundary.
The public-parser regression checks both source roundtrip and semantic equality
at the former u8 boundary (255 colons); equality alone would allow both
backends to lose the same information.
Remaining parity limits
The dated re2c measurements
include known silent invalid cases; do not treat zero-silence results as
guarantees. A clean --parser re2c run is not a general validity
guarantee. Backend disagreements
remain in diagnostic specificity, extra or missing diagnostics, and source
locations. The per-case authority is
tests/integration/error_parity/baseline.rs; run its gate for derived counts.
Both backends feed the shared model validator, but source information discarded before lowering cannot be checked there. Rejected morphology preserves taint; other recovery paths and remaining dummy diagnostic locations still need review.
CLI Usage
# Default: tree-sitter
chatter validate corpus/
# Opt into the experimental re2c backend
chatter validate --parser re2c corpus/
# Roundtrip with re2c
chatter validate --parser re2c --roundtrip corpus/
The --parser flag accepts tree-sitter (default) or re2c. Cache entries
are parser-specific, switching parsers does not invalidate the other’s cache.
Parity Status
The reference-corpus equivalence and roundtrip gates compare actual parsed
models and serialized output. The error-spec gate
backends_diverge_only_where_recorded separately compares diagnostic code
sets against a named, bidirectional baseline: a newly divergent case fails,
and a resolved case must be removed from that baseline. E550 is not in the
baseline because file and fragment participant recovery agree, and neither is
E747: both lexers preserve single logical line breaks, and both parsers locate a
blank line under LF, CRLF and lone-CR endings while retaining its surrounding
utterances.
A passing baseline means that disagreements are accounted for, not that both backends meet every spec. The harness distinguishes backend agreement from each backend’s conformance to the declared spec. Run its report with:
cargo test -p talkbank-parser-re2c --test integration backends_diverge_only_where_recorded --locked -- --nocapture
No wild-corpus percentage or performance figure is recorded here; measure the current build rather than relying on a stored number.
Performance
Run benchmarks: cargo bench -p talkbank-parser-re2c --bench parse_comparison
When to Use Which
| Use Case | Recommended Parser | Why |
|---|---|---|
| LSP / editor integration | tree-sitter | Incremental reparsing |
| Batch validation (>100 files) | tree-sitter | re2c is faster but is not a validity authority |
| CI validation | tree-sitter | The two backends are not interchangeable validity authorities |
| Error diagnostics (user-facing) | tree-sitter | More specific E3xx codes |
| Parser comparison testing | Both | Disagreements require adjudication against the specs; neither backend is an oracle |
| Profiling / benchmarking | re2c | DFA lexer gives a performance floor |
Shared Model Infrastructure
Both parsers convert to the same talkbank_model::ChatFile type and share
post-hoc promotion logic:
TierContent::extract_terminal_bullet(): trailing InternalBullet → utterance bulletparse_bullet_node_timestamps(): structured bullet CST → (start_ms, end_ms)
CA intonation arrows are not promoted to terminators at the
parser/model boundary; both parsers leave them as Separator items.
See CA Terminator Resolution.
Detailed Parity Report
See crates/talkbank-parser-re2c/docs/parity-report.md
for the full gap analysis, divergence categories, and remaining work items.
This page last changed: 2026-10-02 (commit 2d7e886b). The whole book last changed: 2026-10-07 (commit 5e895791).