gitctx trains a small model for one narrow job: produce grounded Conventional Commit messages from Git context.
Input:
- repository metadata;
- repository instructions or detected commit conventions;
git status;- diffstat;
- unified diff;
- changed file paths;
- optional issue, pull request, or changelog context.
Output:
- Conventional Commit subject;
- optional body;
- optional
BREAKING CHANGEfooter; - confidence and warnings.
| Rung | Size | Purpose |
|---|---|---|
| GCTX-0 | rules/templates | deterministic baseline |
| GCTX-1 | 60M-100M | data and eval proof model |
| GCTX-2 | 150M-300M | first serious public model |
| GCTX-3 | 500M | expanded repo-operator tasks |
| GCTX-4 | 1B | only if the project grows beyond commit messages |
The project should not start with 1B. The first proof model should falsify data and evaluation assumptions cheaply.
GCTX-1 is not unlocked by the existence of a tiny DEV artifact. A local smoke
model may be trained earlier to test code paths, but it carries no quality
claim.
Current GCTX-1 proof status:
gctx1-strictis the current private proof artifact.- The strict data-readiness gate is passed: 10,299
DEVtraining records, 1,627 lockedREPORTrecords, and 1,295 reservedHELD_OUTcandidates. - The dependency-free prototype and tiny-softmax smoke paths have run on
gctx1-strict; they validate training/eval plumbing only. - The next model milestone is a real 60M-100M decoder-only proof model trained
from the reviewed
DEVsplit and evaluated on lockedREPORT.
The public proof-model config placeholder is:
configs/gctx1-proof-model.v0.json
Before training a neural model, gitctx uses a dependency-free prototype model to validate the artifact contract:
make training-smoke PILOT_ARTIFACT=nextThis trains path-type-v0 on the DEV records in a reviewed SFT artifact and
evaluates deterministic predictions on REPORT. The prototype model learns only
aggregate path-token/type statistics and emits simple Conventional Commit
messages. It is intentionally weak. Its purpose is to prove that model
artifacts, prediction artifacts, and eval reports can be produced from the SFT
artifact without changing the data lineage.
Artifacts:
artifacts/models/path-type-v0.<artifact>.v0.json
artifacts/eval/path-type-v0.<artifact>.v0.report.predictions.jsonl
artifacts/eval/path-type-v0.<artifact>.v0.report.report.json
This is a training/eval pipeline smoke, not a model-quality benchmark and not a public model release candidate.
After the dependency-free prototype, gitctx can run a tiny dependency-free neural smoke:
make neural-smoke PILOT_ARTIFACT=nextThis trains tiny-softmax-v0, a single-layer softmax classifier, on reviewed
DEV records and evaluates Conventional Commit predictions on REPORT.
Training uses deterministic gradient descent over path and diff-stat features.
It writes a checkpoint-like JSON model artifact, prediction JSONL, and eval
report.
Artifacts:
artifacts/models/tiny-softmax-v0.<artifact>.v0.json
artifacts/eval/tiny-softmax-v0.<artifact>.v0.report.predictions.jsonl
artifacts/eval/tiny-softmax-v0.<artifact>.v0.report.report.json
This is the first real neural-style training loop in the public repo, but it is still a smoke test. It is not a language model, not a model-quality benchmark, and not a public model release candidate.
Minimum GCTX-1 proof-run conditions:
- 10,000 reviewed
DEVtraining records; - 1,000
REPORTrecords; - 1,000 reserved
HELD_OUTcandidates; - at least 25 training repositories;
- at least 5
REPORTrepositories or non-overlapping report windows; - at least 5
HELD_OUTrepositories; - at least two programming-language ecosystems;
- no single repository contributes more than 25% of training records;
- deterministic, raw-teacher, historical-subject, and reviewed-target baselines are recorded.
The split policy is defined in split-contract.md.
Run make split-readiness against any GCTX-1 planning manifest and split plan
before extraction.
After reviewed SFT artifact creation, run proof readiness against the promoted artifact:
make proof-readiness PILOT_ARTIFACT=gctx1split-readiness checks whether the planned DEV/REPORT/HELD_OUT windows are
valid before extraction. proof-readiness checks the actual promoted SFT
artifact and records whether the GCTX-1 proof-run gates are met. A failed
readiness report is still a useful artifact: it identifies the exact missing
condition before expensive training starts.
For the current strict proof artifact, use the named targets:
make gctx1-proof-config-check
make gctx1-proof-readiness
make gctx1-tokenizer
make gctx1-tokenizer-check
make gctx1-proof-handoff
make gctx1-proof-handoff-check
make gctx1-proof-train-dry-run
make gctx1-proof-train-dry-run-check
make gctx1-proof-sequences
make gctx1-proof-sequences-check
make gctx1-proof-sft-smoke
make gctx1-proof-sft-smoke-check
make gctx1-proof-trainer-job
make gctx1-proof-trainer-job-check
uv venv .venv
uv pip install -e .
uv pip install torch
make gctx1-proof-lm-train PYTHON=".venv/bin/python" GITCTX_DATA_DIR="../gitctx-data" GCTX1_PROOF_LM_MAX_RECORDS=32 GCTX1_PROOF_LM_MAX_STEPS=8
make gctx1-proof-lm-train-check PYTHON=".venv/bin/python" GITCTX_DATA_DIR="../gitctx-data"
make gctx1-proof-smoke
make gctx1-proof-smoke-checkThese targets do not train the 60M-100M proof language model. They validate the
proof config contract, rerun readiness against the strict artifact, and run the
current pipeline smoke models on locked REPORT. The tokenizer target builds a
dependency-free regex-diff-v0 vocabulary from reviewed DEV records and uses
locked REPORT only for coverage measurement. The proof config check is not
just JSON parsing: it verifies the GCTX-1 proof band, DEV/REPORT/HELD_OUT split
contract, minimum record thresholds, locked-REPORT evaluation policy,
reproducibility fields, and release-card preconditions before the real training
handoff starts. The proof handoff target writes a private run manifest with the
config hash, training artifact hash, tokenizer hash, readiness gates, training
code revision, and required trainer outputs. The proof-trainer dry-run target
then verifies those handoff hashes, records local runtime capability, counts
tokens and context windows by split, and writes a no-weights checkpoint skeleton.
It also writes a deterministic sequence plan that marks each record as
use_full, use_truncated, or exclude_oversize. The first proof policy uses
one supervised trainer sequence per kept record, preserves locked REPORT
unless every REPORT example can be represented, and excludes only raw records
above the configured raw-token cap from the first proof run. It is the setup
proof for the future decoder-only trainer; it is not a model quality claim.
The proof-sequences target consumes that plan and materializes deterministic
trainer-input metadata hashes, including prefix/suffix crop decisions and
loss-mask hashes, without storing full token-id payloads in the data repository.
The proof SFT smoke target then consumes the same materializer over a bounded
DEV sample, updates a small dependency-free hashed weight vector, and writes a
resumable checkpoint. It proves that selected trainer sequences, optimizer
accounting, checkpoint state, and resume behavior are wired together while
leaving locked REPORT untouched. It is still not the 60M-100M proof language
model and not a model-quality claim.
The proof trainer job target writes the next manifest: a ready-or-blocked
contract for the actual decoder-only trainer. It records the selected
60M-100M-range model shape, sequence metadata hash, checkpoint paths, resume
requirements, and locked REPORT eval outputs before any expensive training
job starts.
The proof LM trainer target is the first real PyTorch decoder-only training
entrypoint. It reads that trainer job manifest, re-materializes the same DEV
sequences, trains only on assistant loss tokens, and writes resumable
checkpoint manifests plus a trainer report. PyTorch remains an optional runtime
dependency for the public package: environments without it receive an explicit
backend blocker instead of a silent partial run. Bounded GCTX1_PROOF_LM_*
limits can be used for CPU smoke runs; removing those limits is the expensive
proof-model training path.
- 150M-300M decoder-only model;
- 8K-16K context;
- code/diff-aware tokenizer;
- quantized local runtime;
- trained from high-quality human labels plus license-approved teacher labels.
- General chat.
- Fully autonomous coding.
- Automatic commit or push.
- Training on closed-model outputs.
- Private repository ingestion without explicit permission.