A CLI for building, inspecting, dumping, repacking, and inferencing compact trigram-model bundles.
lexicore turns text corpora into a versioned, single-file .lexi
bundle containing a shared vocabulary and one or more trigram models.
It can inspect, unpack, repack, and perform reproducible seeded
generation from those bundles.
cargo build --release
lexicore <command> [args]
Commands:
tokenize— read text files, tokenize them, build models, and write one.lexibundleparse— inspect a.lexibundle and print its section tabledump— dump a.lexibundle into split binary files on diskpack— repack a dumped split layout back into a.lexibundleinfer— load a.lexibundle and generate text using two named models
Note
Currently, only inference of two models is supported, and is to be generalized in the future.
For two named models:
cargo run -- tokenize \
--input ./corpus_a.txt --model alpha \
--input ./corpus_b.txt --model beta \
--output ./bundle.lexi
cargo run -- parse --input ./bundle.lexi
cargo run -- dump \
--input ./bundle.lexi \
--output-root ./dumped
Resulting shape:
dumped/
├── vocab.bin
├── vocab.idx
├── alpha/
│ ├── context_keys.bin
│ ├── offsets.bin
│ ├── next_ids.bin
│ ├── counts.bin
│ ├── start_ids.bin
│ └── start_counts.bin
└── beta/
├── context_keys.bin
├── offsets.bin
├── next_ids.bin
├── counts.bin
├── start_ids.bin
└── start_counts.bin
cargo run -- pack \
--input-root ./dumped \
--output ./bundle2.lexi \
--model alpha \
--model beta
Generate a fixed number of tokens:
cargo run -- infer \
--input ./bundle.lexi \
--primary alpha \
--secondary beta \
--alpha 0.03 \
--beta 0.25 \
--mode tokens \
--size 200 \
--seed 42
Generate a fixed number of sentences:
cargo run -- infer \
--input ./bundle.lexi \
--primary alpha \
--secondary beta \
--alpha 0.03 \
--beta 0.25 \
--mode sentences \
--size 20 \
--seed 42
--primary <name>— first model--secondary <name>— second model--alpha <p>— probability of switching from primary to secondary before sampling a token--beta <p>— probability of switching from secondary back to primary before sampling a token--mode tokens|sentences--size <n>--seed <u64>
The generator keeps a current active model:
- if active is
primary, switch tosecondarywith probabilityalpha - if active is
secondary, switch back with probabilitybeta
At each step:
- maybe switch model
- sample from the active model
- if that context is missing, try the other model
- if both fail, emit
</s>
The tokenizer is a manual single-pass Unicode scanner.
- a word starts with a Unicode letter or digit, and continues as such
- internal joiners are allowed only when already inside a word and followed by another word character:
'’-
- whitespace separates tokens
- selected punctuation is emitted as standalone tokens:
. , ! ? ; : ( ) " „ ” « » … -
- anything else is ignored
See FORMAT.md for the binary layout of bundle files.