Skip to content

Commit d4964cc

Browse files
committed
agentgrep(feat[insights]): Wire the LiteRT-LM runtime for local Gemma summaries
why: The MVP could fetch a Gemma .litertlm artifact but the llm level only ran Ollama, so a downloaded model could not actually produce a summary. LiteRT-LM has an installable in-process runtime, so the llm level can run a local Gemma model end-to-end without a daemon. what: - Add a litert-lm backend to the llm level, selected by --backend, that loads the cached .litertlm via litert_lm.Engine, streams the reply through the progress sink, and provisions the model on demand. - Default the LiteRT token budget to 2048: the budget is the total prompt+output KV-cache size, and undersizing it surfaces as an opaque tensor-allocation failure rather than a clear message. - Quiet the LiteRT-LM C++ runtime to ERROR so streamed output is not buried under model-metadata logging on stderr. - Order llm backends by the requested --backend, declare the insights-llm-litert extra, keep litert_lm out of the import path, and cover the runtime and the not-provisioned path with injected-fake tests.
1 parent 180f3d1 commit d4964cc

8 files changed

Lines changed: 245 additions & 42 deletions

File tree

‎docs/cli/insights.md‎

Lines changed: 16 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -50,7 +50,7 @@ and, where relevant, the installed embedding model.
5050
| `ml` | `agentgrep[insights-ml]` | TF-IDF + KMeans topic clusters |
5151
| `embeddings` | `agentgrep[insights-embeddings]` | semantic clusters, near-duplicate detection |
5252
| `index` | `agentgrep[insights-index]` | persistent tantivy + sqlite-vec hybrid index |
53-
| `llm` | `agentgrep[insights-llm]` | local Ollama narrative summary |
53+
| `llm` | `agentgrep[insights-llm]` or `agentgrep[insights-llm-litert]` | local narrative summary via Ollama or an in-process LiteRT-LM model |
5454

5555
List the rungs and what is installed:
5656

@@ -77,13 +77,26 @@ defaults to `tantivy` + `sqlite-vec`; pass `--index-backend lancedb` after
7777
installing `agentgrep[insights-index-lancedb]` for the single-store
7878
alternative.
7979

80-
Stream a grounded local-LLM summary (shown as `text` because it requires a
81-
running Ollama daemon):
80+
Stream a grounded local-LLM summary. Two runtimes are wired: Ollama over
81+
local HTTP, and an in-process LiteRT-LM model (e.g. a Gemma `.litertlm`
82+
artifact). Both examples are shown as `text` because they need a daemon or
83+
a multi-gigabyte model that cannot run as a documentation test:
8284

8385
```text
8486
$ agentgrep insights report --level llm --backend ollama --model llama3.2
8587
```
8688

89+
```text
90+
$ agentgrep insights report --level llm --backend litert-lm --model gemma-4-e2b
91+
```
92+
93+
The LiteRT-LM artifact downloads on demand with `--auto-download-models`,
94+
or ahead of time:
95+
96+
```text
97+
$ agentgrep insights models install gemma-4-e2b --level llm --backend litert-lm --yes
98+
```
99+
87100
The summary is grounded in compact facts — counts, top terms, timeline,
88101
and open-thread titles — not raw transcripts, unless you pass
89102
`--include-text`.

‎pyproject.toml‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -59,6 +59,8 @@ insights-index = ["tantivy>=0.26", "sqlite-vec>=0.1.9", "numpy>=1.24"]
5959
insights-index-lancedb = ["lancedb>=0.33", "numpy>=1.24"]
6060
# Local-LLM summaries over the Ollama HTTP control plane.
6161
insights-llm = ["httpx>=0.28"]
62+
# In-process LiteRT-LM runtime for local Gemma .litertlm artifacts.
63+
insights-llm-litert = ["litert-lm-api>=0.13"]
6264
# Every stable level (1-4); excludes the -st embedding runtime and the LLM level.
6365
insights-all = [
6466
"jinja2>=3.1",

‎src/agentgrep/insights/enrichers/__init__.py‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -124,6 +124,12 @@ class EnricherContext:
124124
setup_command="uv pip install 'agentgrep[insights-llm]'",
125125
builder="build_llm",
126126
),
127+
BackendSpec(
128+
name="litert-lm",
129+
modules=("litert_lm",),
130+
setup_command="uv pip install 'agentgrep[insights-llm-litert]'",
131+
builder="build_llm",
132+
),
127133
),
128134
}
129135

@@ -138,6 +144,9 @@ def _ordered_backends(level: InsightsLevel, request: ReportRequest) -> tuple[Bac
138144
backends = _BACKENDS.get(level, ())
139145
if level == "index" and request.index_backend == "lancedb":
140146
return tuple(sorted(backends, key=lambda b: 0 if b.name == "lancedb" else 1))
147+
if level == "llm" and request.llm_backend not in ("auto", ""):
148+
preferred = request.llm_backend
149+
return tuple(sorted(backends, key=lambda b: 0 if b.name == preferred else 1))
141150
return backends
142151

143152

‎src/agentgrep/insights/enrichers/llm.py‎

Lines changed: 137 additions & 37 deletions
Original file line numberDiff line numberDiff line change
@@ -1,14 +1,16 @@
1-
"""Level 5 enricher: local-LLM narrative summary (Ollama over local HTTP).
2-
3-
The summary is grounded in compact facts — counts, top terms, timeline,
4-
and open-thread titles — never raw transcripts unless ``--include-text``
5-
is set. Tokens stream to the progress sink as they arrive so the CLI can
6-
render the summary live. LiteRT-LM and llama.cpp remain fetch-only in
7-
this MVP; requesting them raises a clear configuration error.
1+
"""Level 5 enricher: local-LLM narrative summary.
2+
3+
Two runtimes are wired: Ollama over local HTTP, and LiteRT-LM in-process
4+
(e.g. a Gemma ``.litertlm`` artifact loaded from the model cache). The
5+
summary is grounded in compact facts — counts, top terms, timeline, and
6+
open-thread titles — never raw transcripts unless ``--include-text`` is
7+
set. Tokens stream to the progress sink as they arrive so the CLI can
8+
render the summary live. llama.cpp remains fetch-only.
89
"""
910

1011
from __future__ import annotations
1112

13+
import contextlib
1214
import json
1315
import os
1416
import typing as t
@@ -21,7 +23,9 @@
2123
from agentgrep.insights.model import InsightsReport
2224

2325
_DEFAULT_ENDPOINT = "http://127.0.0.1:11434"
24-
_DEFAULT_MODEL = "llama3.2"
26+
_DEFAULT_OLLAMA_MODEL = "llama3.2"
27+
_DEFAULT_LITERT_MODEL = "gemma-4-e2b"
28+
_LITERT_MAX_TOKENS = 2048
2529
_MAX_FACT_TERMS = 12
2630
_MAX_FACT_THREADS = 8
2731

@@ -58,48 +62,63 @@ def _build_prompt(report: InsightsReport, *, include_text: bool) -> str:
5862

5963

6064
def build_llm(ctx: EnricherContext) -> InsightsEnrichment:
61-
"""Stream a grounded summary from a local Ollama model."""
62-
backend = ctx.request.llm_backend
63-
if backend not in ("ollama", "auto"):
64-
message = f"local LLM backend {backend!r} is fetch-only in this build; use --backend ollama"
65-
raise BackendConfigurationError(message, level="llm")
66-
65+
"""Stream a grounded summary from the selected local-LLM runtime."""
66+
prompt = _build_prompt(ctx.report, include_text=ctx.request.include_text)
67+
if ctx.backend == "litert-lm":
68+
return _run_litert(ctx, prompt)
69+
if ctx.backend == "ollama":
70+
return _run_ollama(ctx, prompt)
71+
message = f"local LLM backend {ctx.backend!r} is fetch-only in this build"
72+
raise BackendConfigurationError(message, level="llm")
73+
74+
75+
def _emit_delta(ctx: EnricherContext, backend: str, model: str, accumulated: str, text: str) -> str:
76+
"""Emit a streamed delta to the progress sink; return new accumulated text.
77+
78+
Handles runtimes that yield either cumulative or incremental chunks by
79+
treating a chunk that extends the accumulated text as cumulative.
80+
"""
81+
if not text:
82+
return accumulated
83+
if text.startswith(accumulated):
84+
delta = text[len(accumulated) :]
85+
new_accumulated = text
86+
else:
87+
delta = text
88+
new_accumulated = accumulated + text
89+
if delta and ctx.progress is not None:
90+
ctx.progress.llm_chunk(
91+
backend=backend,
92+
model=model,
93+
delta=delta,
94+
char_count=len(new_accumulated),
95+
)
96+
return new_accumulated
97+
98+
99+
def _run_ollama(ctx: EnricherContext, prompt: str) -> InsightsEnrichment:
100+
"""Stream a summary from a local Ollama model over HTTP."""
67101
httpx = ctx.modules["httpx"]
68-
model = ctx.request.model or _DEFAULT_MODEL
102+
model = ctx.request.model or _DEFAULT_OLLAMA_MODEL
69103
endpoint = _endpoint()
70-
prompt = _build_prompt(ctx.report, include_text=ctx.request.include_text)
71104
payload = {
72105
"model": model,
73106
"messages": [{"role": "user", "content": prompt}],
74107
"stream": True,
75108
}
76-
77109
if ctx.progress is not None:
78110
ctx.progress.phase("summarize", detail=f"ollama:{model}")
79111

80-
summary_parts: list[str] = []
112+
accumulated = ""
81113
try:
82-
with httpx.stream(
83-
"POST",
84-
f"{endpoint}/api/chat",
85-
json=payload,
86-
timeout=120.0,
87-
) as response:
114+
with httpx.stream("POST", f"{endpoint}/api/chat", json=payload, timeout=120.0) as response:
88115
response.raise_for_status()
89116
for line in response.iter_lines():
90117
if not line:
91118
continue
92119
event = json.loads(line)
93120
content = event.get("message", {}).get("content", "")
94-
if content:
95-
summary_parts.append(content)
96-
if ctx.progress is not None:
97-
ctx.progress.llm_chunk(
98-
backend="ollama",
99-
model=model,
100-
delta=content,
101-
char_count=sum(len(part) for part in summary_parts),
102-
)
121+
accumulated = _emit_delta(ctx, "ollama", model, accumulated, accumulated + content)
103122
if event.get("done"):
104123
break
105124
except Exception as exc:
@@ -113,12 +132,93 @@ def build_llm(ctx: EnricherContext) -> InsightsEnrichment:
113132
failed = f"Ollama summary failed: {exc}"
114133
raise BackendRuntimeError(failed, level="llm") from exc
115134

116-
summary = "".join(summary_parts).strip()
135+
return _summary_enrichment(
136+
accumulated.strip(), backend="ollama", model=model, endpoint=endpoint
137+
)
138+
139+
140+
def _run_litert(ctx: EnricherContext, prompt: str) -> InsightsEnrichment:
141+
"""Stream a summary from an in-process LiteRT-LM model artifact."""
142+
from agentgrep.insights import models as models_mod
143+
144+
litert_lm = ctx.modules["litert_lm"]
145+
model_id = ctx.request.model or _DEFAULT_LITERT_MODEL
146+
spec = models_mod.resolve_llm_model(model_id, "litert-lm")
147+
if spec is None or spec.artifact_filename is None:
148+
message = f"no curated LiteRT-LM model {model_id!r}"
149+
raise BackendConfigurationError(message, level="llm")
150+
151+
if not models_mod.is_installed(spec, ctx.model_cache):
152+
if not ctx.policy.allow_download:
153+
message = f"LiteRT-LM model {spec.model_id!r} is not provisioned"
154+
install = (
155+
f"agentgrep insights models install {spec.model_id} "
156+
f"--level llm --backend litert-lm --yes"
157+
)
158+
raise BackendConfigurationError(message, level="llm", setup_command=install)
159+
models_mod.install_model(
160+
spec,
161+
model_cache=ctx.model_cache,
162+
progress=ctx.progress,
163+
import_module=ctx.import_module,
164+
)
165+
166+
model_path = models_mod.model_cache_path(spec, ctx.model_cache) / spec.artifact_filename
167+
if ctx.progress is not None:
168+
ctx.progress.phase("summarize", detail=f"litert-lm:{spec.model_id}")
169+
170+
# The LiteRT-LM C++ runtime logs model metadata to stderr at INFO; quiet it
171+
# so the streamed summary is the only thing the user sees.
172+
with contextlib.suppress(Exception):
173+
litert_lm.set_min_log_severity(litert_lm.LogSeverity.ERROR)
174+
175+
accumulated = ""
176+
try:
177+
engine = litert_lm.Engine(
178+
str(model_path),
179+
backend=litert_lm.Backend.CPU,
180+
max_num_tokens=_LITERT_MAX_TOKENS,
181+
)
182+
try:
183+
conversation = engine.create_conversation()
184+
for chunk in conversation.send_message_async(prompt):
185+
accumulated = _emit_delta(
186+
ctx, "litert-lm", spec.model_id, accumulated, _litert_chunk_text(chunk)
187+
)
188+
finally:
189+
engine.close()
190+
except Exception as exc:
191+
failed = f"LiteRT-LM summary failed: {exc}"
192+
raise BackendRuntimeError(failed, level="llm") from exc
193+
194+
return _summary_enrichment(
195+
accumulated.strip(), backend="litert-lm", model=spec.model_id, endpoint=str(model_path)
196+
)
197+
198+
199+
def _litert_chunk_text(chunk: t.Any) -> str:
200+
"""Extract response text from a LiteRT-LM conversation chunk."""
201+
content = chunk.get("content") if isinstance(chunk, dict) else None
202+
if isinstance(content, str):
203+
return content
204+
if isinstance(content, list):
205+
return "".join(part.get("text", "") for part in content if isinstance(part, dict))
206+
return ""
207+
208+
209+
def _summary_enrichment(
210+
summary: str,
211+
*,
212+
backend: str,
213+
model: str,
214+
endpoint: str,
215+
) -> InsightsEnrichment:
216+
"""Build the enrichment payload for a generated summary."""
117217
return InsightsEnrichment(
118218
level="llm",
119-
backend="ollama",
219+
backend=backend,
120220
status="ok",
121-
message=f"summarized via ollama:{model}",
221+
message=f"summarized via {backend}:{model}",
122222
data={"summary": summary, "model": model, "endpoint": endpoint},
123-
provenance={"backend": "ollama", "model": model, "endpoint": endpoint},
223+
provenance={"backend": backend, "model": model, "endpoint": endpoint},
124224
)

‎src/agentgrep/insights/models.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -131,7 +131,7 @@ def files(self) -> tuple[str, ...]:
131131
license="Gemma",
132132
source_url="https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm",
133133
local_id="gemma4-e2b",
134-
notes="LiteRT-LM Gemma 4 E2B; fetch-only registry parity in this MVP.",
134+
notes="LiteRT-LM Gemma 4 E2B; runs in-process via agentgrep[insights-llm-litert].",
135135
),
136136
LLMModelSpec(
137137
model_id="phi-4-mini-gguf",

‎tests/test_import_time.py‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -36,6 +36,7 @@
3636
"httpx",
3737
"jinja2",
3838
"huggingface_hub",
39+
"litert_lm",
3940
)
4041

4142

‎tests/test_insights_enrichers.py‎

Lines changed: 63 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -336,3 +336,66 @@ def test_llm_enricher_streams_grounded_summary() -> None:
336336
assert enrichment.provenance is not None
337337
assert enrichment.provenance["backend"] == "ollama"
338338
assert progress.deltas == ["Worked on ", "indexing."]
339+
340+
341+
def _fake_litert_lm() -> t.Any:
342+
"""Return a fake litert_lm whose conversation yields cumulative chunks."""
343+
344+
class _Conversation:
345+
def send_message_async(self, _prompt: str) -> t.Iterator[dict[str, str]]:
346+
for text in ("Worked on ", "Worked on indexing."):
347+
yield {"role": "model", "content": text}
348+
349+
class _Engine:
350+
def __init__(self, _path: str, **_kwargs: object) -> None:
351+
pass
352+
353+
def create_conversation(self) -> _Conversation:
354+
return _Conversation()
355+
356+
def close(self) -> None:
357+
pass
358+
359+
return types.SimpleNamespace(
360+
Engine=_Engine,
361+
Backend=types.SimpleNamespace(CPU=object()),
362+
LogSeverity=types.SimpleNamespace(ERROR=object()),
363+
set_min_log_severity=lambda _severity: None,
364+
)
365+
366+
367+
def test_llm_enricher_litert_streams_from_local_model(tmp_path: pathlib.Path) -> None:
368+
"""The litert-lm backend loads a provisioned model and streams a summary."""
369+
spec = models_mod.resolve_llm_model("gemma-4-e2b", "litert-lm")
370+
assert spec is not None
371+
target = models_mod.model_cache_path(spec, tmp_path)
372+
target.mkdir(parents=True, exist_ok=True)
373+
(target / "agentgrep-manifest.json").write_text("{}", encoding="utf-8")
374+
375+
progress = _RecordingProgress()
376+
report = build_report(
377+
_RECORDS,
378+
ReportRequest(requested_level="llm", llm_backend="litert-lm", model="gemma-4-e2b"),
379+
import_module=_importer({"litert_lm": _fake_litert_lm()}),
380+
progress=progress,
381+
model_cache=tmp_path,
382+
)
383+
enrichment = report.enrichments[0]
384+
assert enrichment.status == "ok"
385+
assert enrichment.backend == "litert-lm"
386+
assert enrichment.data["summary"] == "Worked on indexing."
387+
assert progress.deltas == ["Worked on ", "indexing."]
388+
389+
390+
def test_llm_enricher_litert_errors_when_model_not_provisioned(tmp_path: pathlib.Path) -> None:
391+
"""An unprovisioned litert model yields an error enrichment with an install hint."""
392+
report = build_report(
393+
_RECORDS,
394+
ReportRequest(requested_level="llm", llm_backend="litert-lm", model="gemma-4-e2b"),
395+
import_module=_importer({"litert_lm": _fake_litert_lm()}),
396+
model_cache=tmp_path,
397+
)
398+
enrichment = report.enrichments[0]
399+
assert enrichment.status == "error"
400+
setup = next(d.setup_command for d in report.diagnostics if d.setup_command)
401+
assert "models install" in setup

0 commit comments

Comments
 (0)