Skip to content

Commit e119166

Browse files
thomas-manginclaude
andcommitted
docs(ai): add learned summary 767-tokenizer-no-escape
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1 parent 33342a3 commit e119166

3 files changed

Lines changed: 39 additions & 1 deletion

File tree

‎ai/LEARNED-INDEX.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -47,6 +47,7 @@ Buffer-first, zero-copy, attribute pools, UPDATE building, NLRI parsing.
4747
- [722](plan/learned/722-spec-bgp-4-aspa-policy.md) -- ASPA policy enforcement: override ordering (ASPA reject wins over origin accept), re-validation via validateCh, origin policy is hardcoded not configurable
4848
- [764](plan/learned/764-attr-flags-json.md) -- Attribute flags in Ze native JSON: includeFlags parameter on shared formatter, static flags for pool RIB, unwrap at extractRoutes for LG consumers
4949
- [765](plan/learned/765-gc-pressure-reduction.md) -- GC pressure reduction: [256]bool for attr codes, inline FNV-1a, NextHopAddrs inline struct, LargeCommunities dedup fast path, clear() map reuse; stack arrays that escape via closure/interface/return are NOT optimizations
50+
- [767](plan/learned/767-tokenizer-no-escape.md) -- Command tokenizer: removed backslash escape handling, backslash is a normal character; no per-byte escape scan on every command
5051

5152
## Plugin System
5253

‎plan/learned/.counter‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1 +1 @@
1-
766
1+
767
Lines changed: 37 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,37 @@
1+
# 767 — Command Tokenizer: No Escape Sequences
2+
3+
## Context
4+
5+
Both command tokenizers (`plugin/server/command.go:tokenize` and
6+
`cli/model_commands.go:tokenizeCommand`) scanned every byte/rune for
7+
backslash escape sequences (`\"`, `\\`). Commands never contain
8+
backslashes in practice; the escape logic was defensive code with no
9+
real use case.
10+
11+
## Key Decisions
12+
13+
### Backslash is a normal character
14+
15+
Rather than adding a fast path (`IndexByte` check to skip escape logic),
16+
we removed escape handling entirely. Backslash is treated as a literal
17+
character, the same as any other. This removes two branches per rune
18+
from the hot loop and simplifies the tokenizer to: split on whitespace,
19+
respect quote delimiters.
20+
21+
### joinTokensWithQuotes drops escape encoding
22+
23+
The round-trip encoder `joinTokensWithQuotes` no longer escapes
24+
backslashes or quotes. Since `tokenizeCommand` uses `"` as a delimiter
25+
(never as content), no token from the tokenizer can contain `"`, so
26+
there is nothing to escape.
27+
28+
### Web tokenizer was already correct
29+
30+
`web/cli.go:tokenizeCommand` never had escape handling. The two
31+
"full-featured" tokenizers now match its simplicity.
32+
33+
## Traps
34+
35+
- Removing escape support means there is no way to embed a literal `"`
36+
inside a quoted value. This is acceptable: Ze command values (peer
37+
names, descriptions, paths) have no use case for embedded quotes.

0 commit comments

Comments
 (0)