Skip to content

1.35.0 - #1396

Merged
jakeaturner merged 89 commits into
mainfrom
dev
Sep 29, 2026
Merged

1.35.0#1396
jakeaturner merged 89 commits into
mainfrom
dev

Conversation

@jakeaturner

Copy link
Copy Markdown
Collaborator

No description provided.

jakeaturner and others added 30 commits September 8, 2026 11:33
Allows users to optionally specify a model to use for ancillary tasks (chat suggestions, chat titles, etc.) instead of strictly using the last used model. This means aesthetic/non-critical work can be passed to a small, lightweight model instead of a heavy reasoning model used for chats.

If a model is not selected in the AI Assistant settings (`/settings/models`), the existing behavior of using the last used model is retained.
Chat was run through Ollama's OpenAI-compatible /v1 endpoint, which has no
context-size field, so the num_ctx we sent was silently discarded and the
entire NUM_CTX_LADDER was dead code. Every conversation ran at the backend
default and was truncated from the middle (where history and
retrieved context actually sit). Measured with identical ~7k-token prompts:
2,060 tokens processed against 7,049 on the native endpoint, and with the
needed fact in the discarded middle the old path hallucinated a plausible
wrong number.

Transport: add a native /api/chat path using the already-installed ollama
package, sending a real options object (num_ctx, num_predict, temperature,
seed) plus keep_alive. The OpenAI path is still used for LM Studio and llama.cpp,
minus the ineffective num_ctx and with include_usage so both paths return
prompt token counts. /api/show is now read for trained context, parameter
size, quantization and any modelfile num_ctx. The <think> stream normalizer
moves to app/utils/think_stream.ts so both transports share it; it gains a
flush() that fixes a dangling partial tag being dropped at end of stream.

Token accounting: replace chars/N with a structural segmenter scaled by a
per-model factor learned by EWMA from the prompt_eval_count every response
already returned and threw away. End-to-end estimator error is 4.6% against
18-20% for a fixed divisor — qwen2.5 converges to k=1.26 where llama3 sits
at 1.00, so no single global constant serves both.

Context window: resolved once per model from trained context, the exact KV
cost derived from GGUF metadata (head_count_kv, not head_count — GQA makes
them differ 4x), and available VRAM, then memoized. Memoization is
load-bearing: asking for a different num_ctx makes Ollama unload and reload
the model, stalling the turn and discarding the KV cache. The floor never
overrides the model's trained ceiling — tinyllama trains to 2048, and a
blind 4096 floor pushes RoPE past anything the weights saw.

Prompt budget: planPrompt allocates explicitly instead of delegating
overflow to silent backend truncation — response reserve, then fixed blocks
(system and current query, never dropped), then RAG chunks whole and
best-first, then history in whole user+assistant turns, evicted newest-first
in blocks with an elision marker. Block eviction keeps the prefix stable for
several turns rather than shifting on every one.

KV-cache ordering: the per-turn retrieved block moves behind history to the
tail, so the stable prefix stays byte-identical turn to turn. At turn 14 of
a long conversation, prefill is 231ms at the tail against 539ms at the front
(~1ms/turn growth against ~25ms). The win is conditional — with a large
retrieved block and short history the two measure the same — so both
orderings ship behind RAG_PLACEMENT. The per-turn query rewrite also moves
to the tasks model; on a single-slot Ollama it was evicting the chat model's
cached prefix on every single turn.

Also: OLLAMA_CONTEXT_LENGTH and OLLAMA_KV_CACHE_TYPE=q8_0 on the container
(q8_0 is treated as headroom, never as a budgeting assumption, since
unsupported architectures fall back to f16 silently), a Context Window
setting in Settings > Models, and budget fields on PipelineTrace.

Eval: golden suites (--suite) so the new long_context goldens cannot move
the aggregates the committed baselines were recorded against, plus budget
and estimator-error metrics through the report path. Retrieval is unchanged
and eval:compare reports no metric outside the tolerance band; the
long_context suite scores correctness 1.000 with historyElided 0.500 (the
two eviction cases, by design). Unit tests 337 -> 386 passing, same 7
pre-existing failures.

Ingest-side CHAR_TO_TOKEN_RATIO is deliberately untouched: changing it
re-chunks the corpus and invalidates retrieval baselines, which belongs in
its own measured change.
…query (#1254)

ThinkTagSplitter was wired into both streaming paths but neither
non-streaming one, and every ancillary call goes through non-streaming
chat(). So a reasoning tasks model put its reasoning transcript into the
sidebar title, into the suggestion chips, and — worst — into the string
that gets embedded and sent to Qdrant, silently corrupting retrieval for
the rest of the conversation. splitThinkTags() already existed and was
already unit tested; it just had no production callers.

Fixed in three layers:

1. Strip at the transport. New normalizeNonStreamed() in think_stream.ts
   merges structured reasoning (native `thinking`, /v1 `reasoning`) with
   anything the splitter pulls out of inline tags, and both _chatNative
   and _chatCompat now return through it. That covers every caller at
   once, including the legacy controller path and the eval harness.

2. Suppress at the source. _chatNative and _chatStreamNative forwarded
   `think` only when truthy, so `think: false` was never sent and Ollama
   left reasoning on for capable models. Guard is now definedness. This
   also fixes the ai.autoThinking=off setting, which had no effect on
   the native streaming transport.

3. Guard the call sites. Title, suggestions, and query rewrite now pass
   think:false plus thinkingCapable (from the memoized capability check,
   so no extra round-trip), and handle an empty-after-strip result. The
   rewrite matters most: QUERY_REWRITE_MAX_TOKENS caps it at 120, which
   a reasoning model burns entirely on thinking, so stripping alone
   would have turned a corrupt query into an empty one. It now falls
   back to the raw user message instead of embedding "".

Verified: think_stream.spec.ts 17/17 (6 new cases, including the
truncated-mid-thought case that drives the fallback), typecheck clean,
full test:unit failures byte-identical to the pre-change baseline.
Ollama's format parameter turns a formatting request into a decoder
constraint, and NOMAD's three ancillary calls — chat title, suggestion
chips, RAG query rewrite — were the ones paying for not having it.
The rewrite is particularly important: whatever it returned got embedded
and sent to Qdrant, so a malformed response degraded retrieval for the
rest of the conversation.

Three explicit design choices made:
- Native-only. The compat path sends nothing new and keeps its string
parsing, so a non-Ollama backend that rejects response_format can't
take out title generation entirely.
- parseStructured returning null is the control flow. It's how each
caller falls back to the parser it already had, which is why it's written
to be total and never throw.
- {queries: []} not {query: ""} for the rewrite (tee-ing up future
improvements)
…er it

There was one cutoff, Qdrant's score_threshold at 0.3, and nomic scores
unrelated text well above it, so every question retrieved *something*
(nonEmptyRateOnRefusal 1.0) and rag_context wasted four lines asking a 3B
model to "silently judge" relevance. That is a classification a number
should decide, not a micro model.

Adds `RAG_MIN_FINAL_SCORE`, a floor on the post-rerank score, applied before
source diversity. The floor judges relevance, the 0.85^n penalty judges
redundancy. Nothing clears it and nothing is injected. `score_threshold` stays
at 0.3 as the candidate net. Exposed as `rag.minRelevance` with a module-
scoped TTL cache invalidated from `updateSetting`.

The floor is calibrated and the harness was describing the wrong number:
aggregate() bucketed semanticScore, so its percentiles measured raw cosine
while the new knob filters the reranked score. Both axes now, and the gap
matters: recall falls at 0.60 on the semantic axis, 0.66 on the final one.

  floor  recall@5  precision@5  emptyOnAnswerable  nonEmptyOnRefusal
  0.00     0.991      0.221            0%               100%
  0.62     0.991      0.524            0%                60%
  0.66     0.991      0.695            0%                60%
  0.67     0.972      0.704            1%                60%

Default 0.62, not the last-lossless 0.66: the refusal win is complete at
0.62 and 28 documents is a thin basis for spending margin. 0.60 is the
honest floor. Three of five refusal goldens are near-corpus by
construction, where declining is the generation tier's job.

Both eval tiers pin the floor to the constant and ignore the setting, for
the same reason the harness never sets skipRetrieval.

4 metrics improved against baseline, no regressions.
Adds a security policy so vulnerability reports are routed to the private
advisory form instead of public issues, and defines scope so the intentional
no-auth LAN design is not re-reported as a vulnerability.

Also adds a security contact link to the issue template chooser (blank issues
are disabled, so there was previously no route for a private report), and
extends .gitignore to cover .env variants, key/cert files, .npmrc/.netrc and
local copies of the deployed compose file. A local management compose with real
generated DB passwords was committed to a branch once before; secret scanning
push protection now covers that case too.

No tracked file is affected by the new ignore rules.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The history table rendered nomad_score for every row, so a v2 run showed its
legacy score while the Benchmark Details card directly above it showed the v2
one. The same run read as 65.0 in the table and 1036.1 in the card.

Rows predating v2 have no v2 score and can only show the legacy number, so the
two scales necessarily share a column. They differ by more than 10x, and
without the scale on screen a v2 run sitting above older runs reads as a
collapse rather than a rescale. Legacy rows now carry a "/ 100" suffix, which
names the scale in place and reuses the wording the details card already uses
("Legacy scale: X / 100").

This compounds the re-run guidance in the release notes: someone re-runs on
1.34 as asked, gets a real v2 score, and their own history then shows the old
number as though the machine got slower.

Frontend only. getAllResults() already returns full rows, so nomad_score_v2
was in the payload the whole time.

Verified on the NOMAD2 dev environment with a seeded history matching a real
mixed case (one v2 run at 1036.1 plus two pre-v2 runs): the v2 row shows
1036.1 with no suffix and agrees with the details card, and the legacy rows
show 71.1 / 100 and 86.5 / 100. Inertia typecheck shows the same 30
pre-existing errors as dev with none in this file.
The batch continuation gated on articles that produced chunks rather than
articles the extractor consumed. Articles whose text is empty after HTML
cleaning — redirect stubs, category listings, media and PDF wrappers —
return no chunks and so no documentIds, which made a normal sparse batch
look like the end of the archive. Ingestion stopped after the first batch
and the rest of the file was silently skipped.

Medicine LibreTexts embedded 16 chunks out of 23,171 articles: its first
50 entries are navigation pages, only 10 of which had text, so 10 >= 50
was false and no continuation was ever dispatched. Measured against the
same install, WikiMed reached 28% of its articles, Libre Pathology 31%,
MedlinePlus 84%, and CDC Travelers' Health 2.7%.

The same value was also returned as `articlesProcessed`, which
EmbedFileJob uses to advance the next batch offset. Advancing by the
content-bearing count re-read the overlap on every sparse batch, and a
batch yielding no text at all would have advanced by zero and re-run the
same window indefinitely.

Both now use the extractor's consumed-article count, which it already
computed and logged but discarded. The end-of-archive signal is the
iterator running dry (articlesProcessed < batchSize); archive.articleCount
cannot serve as a bound because iterByPath() legitimately yields more
article entries than that figure for some archives.

The log line now reports both counts (content-bearing/consumed) so a
sparse batch is visible without re-deriving it.
Fixes #1269. Fixes #1203.

`file-embeddings` and `downloads` both routinely outlive the default
5 minute lock, so BullMQ calls healthy jobs stalled on both. They need
opposite recovery policies, which is why neither was covered by the
existing drug-queue override.

Embedding is a self-continuing chain, so re-queueing a stalled job while
the original is alive runs two chains over the same file. Both write to
Qdrant with freshly minted point ids, so the second cannot overwrite the
first, and the queue's concurrency of 2 means the fork never contends
for a slot. maxStalledCount 0 fails on the first stall instead.

A download is resumable, so recovery is safe there; the default of 1 was
failing jobs outright and bypassing their own attempts: 10 policy.

The decision moves to a pure, unit-tested util alongside the other
decision helpers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LtGWCVMPYn6qPWVYiycK74
…t silently dropped

Three dispatch sites pin a deterministic jobId. `queue.add` with an existing
custom jobId returns the existing job rather than throwing, so any retained
terminal record makes later dispatches a silent no-op while still reporting
success to the caller.

- download_model_job: only `failed` records were cleared, so once a model
  downloaded successfully its retained completed job blocked every later
  request for that model. Deleting the model and reinstalling it did nothing,
  and a restart did not help because the record is persisted in Redis. Adds
  getActiveByModelName, mirroring RunDownloadJob.getActiveByUrl.
- download_service.retryFailedJob: the model branch dispatched before removing,
  so remove() targeted the job just enqueued under the same id instead of the
  old record. Flipped to remove-then-dispatch, matching the file branch.
- embed_file_job: a retained failed record made re-indexing that file a no-op,
  returning 202 "Indexing queued" with nothing enqueued.

In-flight jobs are still returned as-is, so re-clicking during an active
download or embed stays idempotent.

Diagnosis by @caweis in #1214, including the retry-ordering race and the embed
case. Reimplemented in-house rather than ported, as noted there.

Closes #1214
Ollama's /api/pull returns HTTP 200 and reports failures IN-BAND, as a JSON
line inside the NDJSON stream body:

    HTTP_STATUS=200
    {"status":"pulling manifest"}
    {"error":"pull model manifest: file does not exist"}

The stream handler only ever read completed/total/digest off each line and
silently dropped everything else, including that error line. The stream then
ended normally, on('end') resolved unconditionally, and we logged the model as
downloaded successfully. on('error') never fires for this, because at the HTTP
layer the request succeeded.

Every failed model pull, for any reason, was reported as a success. On #1071
that surfaced as an ~800 GB fp16 model "downloading" in two seconds and nothing
being installed, which had been triaged twice as disk space or a stale blob.
Those explanations produce workarounds that cannot work, since the pull never
reaches the disk at all.

Three changes:

- Check parsed.error in the data handler and destroy the stream with it, so the
  existing catch reports a real failure.
- Verify the model is actually in getModels() before returning success. This is
  the load-bearing half: it catches silent failures we have not anticipated,
  not only the ones Ollama labels. includeEmbeddings must be true or
  nomic-embed-text fails verification after a good pull, and names are
  normalised because Ollama resolves a bare name to ":latest".
- Treat an in-band "model does not exist" as permanent. It was retryable, so a
  bad model name spun through attempts: 10 on exponential backoff and the user
  watched a delayed job for hours instead of being told the name was wrong.

Verified on NOMAD3 against the shipped v1.34.0 build, which carries the same
defect:

- bogus tag, before: 'Model "..." downloaded successfully.'
- bogus tag, after:  'Failed to download model "...": pull model manifest:
  file does not exist', and the job stops at 1 attempt instead of retrying
- real model (qwen2.5:0.5b): still succeeds, passes verification, and is
  present in `ollama list`

Fixes #1273
Refs #1071

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LtGWCVMPYn6qPWVYiycK74
When an AMD GPU is detected, the install switches to the ROCm image and passes
/dev/kfd and /dev/dri into the container. _discoverAMDDevices() returns both
unconditionally, with no check that they exist. Detection reads PCI data, which
says nothing about whether the ROCm kernel interface was ever loaded, so on a
pre-ROCm card or a host without amdgpu the container is rejected:

    (HTTP code 500) server error - error gathering device information while
    adding custom device "/dev/kfd": no such file or directory

and the install dies. There is an opt-out, ai.amdGpuAcceleration, but it is a
backend-only KV key with no settings toggle, no route and no UI, so an affected
user cannot install the AI Assistant at all through the interface even though
CPU-only would work fine on their machine.

Fall back to the CPU image instead, with a message explaining why.

Two things about the shape of this fix, both measured rather than assumed:

- It hooks start(), not createContainer(). The daemon accepts a container whose
  --device path does not exist and only resolves devices on start. Verified
  against Docker 29.6.2 over the socket: create returns 201, start returns 500
  with the message above.

- It retries rather than pre-checking. The admin runs in its own container, so
  stat()ing /dev/kfd describes the admin's namespace, not the host's. Verified
  on NOMAD6, a working ROCm box (Ollama running ollama/ollama:rocm with both
  devices attached): /dev/kfd is PRESENT on the host and ABSENT inside the admin
  container. A pre-check would therefore have stripped GPU passthrough from a
  machine where AMD acceleration works.

The matcher keys on the device-gathering phrase rather than a bare "no such file
or directory", which Docker also emits for missing bind-mount sources; falling
back to CPU on one of those would hide a real failure behind a degraded install.
Checked against all three formats the daemon/dockerode produce, and against a
port-conflict message, to confirm it discriminates.

The success path is unchanged: the create payload is built by one helper used by
both attempts, and the fallback only runs after a start() failure that matches.

Fixes #1232

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LtGWCVMPYn6qPWVYiycK74
…PU model

Under WSL2 the GPU is reached through /dev/dxg rather than the real adapter, so
si.graphics() returns Microsoft's generic placeholder even while CUDA work runs
on a physical card. isUnresolvedGpuModel() caught blanks, "Device <hex>" and
"Unknown" but not this, so the placeholder was accepted as a legitimate model
name and the authoritative detection path was skipped.

The reporter on #1218 has an RTX 3090 doing real work (92% utilisation, 11.6 GB
VRAM, 352 W, 127.9 tok/s) that reaches the public leaderboard labelled
"Microsoft Basic Render Driver". A wrong label is worse than none: it fragments
per-hardware grouping, which is the same failure #1165 fixed for raw PCI ids.

Adds Microsoft's two placeholder adapter names to the same rejection list.

Scoped so it cannot affect anyone off WSL: these are Microsoft's own placeholder
names and appear nowhere outside a Windows/WSL graphics stack, so on a native
Linux host the function returns exactly what it returns today. Checked against
real product names from all four vendors, including "Microsoft Basic Renderer
Pro 9000" as a near-miss, to confirm the anchored pattern cannot swallow one.

Note this makes the model resolve to unknown rather than to "RTX 3090": the
admin container has no nvidia-smi and WSL's copy lives at /usr/lib/wsl/lib,
outside the container. Honest-unknown is the correct outcome here; naming the
card under WSL is a separate problem.

Refs #1218

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LtGWCVMPYn6qPWVYiycK74
The quickstart never said how much free space is required to install, which is
the first thing anyone plans around on a project whose whole purpose is
downloading large offline libraries. The FAQ's storage question only covered
content sizes, so it answered "how big is Wikipedia" but not "can I install
this at all".

Measured on a v1.34.0 install: admin 2.81GB, MySQL 1.1GB, plus redis, dozzle,
the disk collector and the sidecar updater at roughly 300MB combined. That is
about 4.2GB of images, so ~5GB installed and 10GB free to install comfortably.
The AI Assistant adds the Ollama image at 10.6GB plus a model (llama3.1:8b is
~4.9GB), which is where the ~25GB figure comes from.

Reported in #1271, which proposed the README sentence with the number left
blank.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LtGWCVMPYn6qPWVYiycK74
Measured on a running install rather than estimated. The README already
gave 4 GB and 32 GB; what was missing is why the gap is that wide, which
is the model rather than NOMAD, and whether it comes out of VRAM or
system RAM.

Asked for by @mk-pmb on #1271.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
Three copies of a figure that varies is worse than one. Keep the two
numbers a reader needs before running the install command, and move the
detail down to Device Requirements where it sits with the specs.

Suggested by @mk-pmb on #1271.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
_embedWithFallback() always tried Ollama's native /api/embed first and only
reached /v1/embeddings from the catch block. Against a backend that only
speaks the OpenAI embeddings API, every batch therefore paid a
guaranteed-failing round trip plus a warn line before doing the real work.

Invisible on a handful of documents. Expensive on a bulk file-embeddings
job: the reporter is indexing several million chunks through an
OpenAI-compatible proxy, which doubles request volume and log noise against
their embedding host for the entire lifetime of the ingest.

Gates the native attempt on _isNativeBackend(), matching how _doDownloadModel
already skips /api/pull. Used the probe rather than reading isOllamaNative
directly: the field is null until something populates it, and an embed job
can easily be the first thing to touch the service after a restart, which is
exactly when a long ingest is starting and the saving matters most. The probe
memoizes for the process lifetime, so this costs one request overall rather
than one per batch.

Behaviour is otherwise unchanged. A native backend still tries /api/embed
first and still falls back if it fails for a non-context reason, and
context-length errors still bubble to embed() for the truncated retry. The
response-shape validation stays: the probe establishes that /api/tags
answered like Ollama, not that /api/embed will.

Reported by @horsleyb.

Closes #1279

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#1165 taught the benchmark path to reject an unresolved PCI id as a GPU
model, but isUnresolvedGpuModel lived inside benchmark_service.ts and the
Settings > System display path never learned the same thing. It still shows
whatever lspci handed back.

The two symptoms of a stale pci.ids database are independent, which is why
the existing guard misses this. system_service already re-probes when a
discrete-GPU controller reports an implausible BAR0 reading as VRAM, and
that happens to fix the name as a side effect. But lspci can return a
perfectly plausible VRAM figure alongside a name it could not resolve, and
then nothing fires.

Confirmed live on an AMD box. The Radeon 780M reports:

  model: "Device 1900", vram: 512

512 MiB clears the 256 MiB bogus-VRAM threshold, so no probe runs and the
System page renders "Device 1900". The same box's Ollama startup log carries
description="AMD Radeon 780M Graphics", so the probe the guard declined to
run had the real name available the whole time. NVIDIA hides this because
its path tends to be reached for other reasons; AMD has no equivalent.

Adds an unresolved-name trigger alongside the VRAM one, and moves
isUnresolvedGpuModel to app/utils/gpu_model.ts so the submission path and the
display path share one definition instead of diverging again. Tests cover
both placeholder shapes and, more importantly, a set of real product names
that must never be rejected.

Stacked on #1277, so the shared helper carries the WSL placeholders too.

Closes #1196

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes #1179.

Surfaces the documents an answer was drawn from as a "Sources" list under the
assistant message, with the archive date where one is known, and persists it so
reopening a past conversation still shows what backed each answer.

Offline there is no second opinion to check an answer against, so provenance is
the only signal the user has that a claim came from a real document rather than
from the model. This is the other half of retrieval being allowed to decline:
one says when nothing relevant was found, this says what was found.

Citations are built from the chunks actually injected into the prompt, not from
everything retrieval returned. A chunk dropped by the budget planner never
reached the model, so citing it would credit the answer to a document it was
not based on -- the exact failure the list exists to prevent.

The archive title and date have been written into the vector payload by the ZIM
ingestion path all along and were simply never read back, so this works on
already-embedded content with no re-embed.

Rebased onto the RagPipelineService refactor, which landed after this was
written: the citation list now comes off the pipeline trace rather than the
inline retrieval block that used to live in the controller. Entries with
neither a path nor a title are skipped rather than shown as "Unknown source" --
a row the user cannot act on makes an answer look sourced when it is not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
Closes #1176. Salvages @gabegraves's work from #1178, rebased onto the
RagPipelineService refactor that landed after it was written.

Adds JPEG/PNG/WebP attachments to chat over multipart, leaving text-only JSON
requests on their existing path. Images are normalized server-side to bounded
JPEG with metadata stripped, capped on count/size/pixels/output, and the
temporary uploads are cleaned up on every validation path. They are attached
only to the latest user message, are excluded from RAG, titles, logs and
persisted history, and are never written to storage.

Vision support is resolved per model from what the backend actually reports --
OpenAI-compatible /v1/models metadata, then Ollama /api/show, then llama.cpp
/props -- never inferred from a model name. A model that reports no vision
support has attachments disabled with a visible explanation; a backend that
reports nothing is allowed with a warning and its rejection is translated into
plain compatibility guidance rather than a raw error.

Changes from the original:

- Prompt assembly moved into RagPipelineService in the meantime, so the
  controller half is rebuilt against the trace and the multipart-aware
  collection filter is now the only one (dev's request.input() version would
  have silently dropped the collection on every image request).
- Added a native-path converter. dev split chat into native and OpenAI-compat
  paths after this was written; Ollama's native /api/chat does not take OpenAI
  content parts, it takes a sibling `images` array of bare base64. Without this
  every image request on a native Ollama backend would fail.
- Dropped the model-recommendation change (selectRecommendedModels and the
  Qwen3-VL fallback entry). Displacing a recommended model at first-time setup
  is a product decision and is being taken separately. The "Supports images"
  label in Easy Setup is kept: it describes what is already on offer rather
  than changing it.
- Dropped the DEFAULT_QUERY_REWRITE_MODEL download prompt his branch carried,
  which dev deleted in d022c2d.
- Removed a duplicate checkModelHasThinking; kept dev's memoized getModelInfo
  path, since thinking detection is not part of this feature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
It constructs OllamaController, which imports the Adonis logger service at
module scope, so it cannot run under the plain node:test runner. It failed
identically on the original branch -- the coverage has never executed. Moving
it to the Japa runner puts it in the same bucket as the other app-importing
specs, where a booted app is available.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
Salvages @kennethbrewer3's work from #836, which had drifted onto a stale base
(GitHub showed 231 files; the real change is the map UI plus a migration).

Marker visibility (individually and Hide All), per-marker icon and custom hex
colour with contrast-adjusted icon colour, search, sorting and a tree view over
the marker list, coordinate search and fly-to with URL parameters
(?lat=&lng=&zoom=), creating a marker from typed coordinates, and markers
rendered as SVGs anchored to true coordinates rather than approximately.

Name, notes and preset colour already existed; the new columns are
custom_color, icon, icon_color and visible.

Dropped from the original as stale-base noise rather than his work: unrelated
duplicate changes to the Ollama, Docker, download, RAG and ZIM services, the
API reference and getting-started docs, README/FAQ/CONTRIBUTING, the installer
script, and a revert of AppLayout to a version predating the compact sidebar.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
Review changes from looking at the feature running on NOMAD3.

The icon picker offered every Tabler icon *and* every Font Awesome icon via
namespace imports: 160 pages, whose first page included the Amazon and Amazon
Pay logos, an "Ad" badge and text-alignment glyphs. It also took the maps chunk
from 864 kB to 5,774 kB. Replaced with a hand-written set of 36 map-relevant
icons in a 6x6 grid -- water, shelter, medical, power, comms, transport,
terrain, hazards -- with no search box or paging, because 36 fit on screen.
Maps chunk is back to 904 kB, +40 kB over dev rather than +4,910 kB, and the
react-icons dependency is gone entirely since Tabler was already present.

Icon resolution moved into marker_icons.ts so the panel and the pin share one
definition, with a caller-supplied fallback (the pin's inner glyph defaults to
a filled circle, the list to a map pin). An icon later removed from the set
still renders as the fallback rather than blanking the marker.

Dropped the tree view and the List/Tree band: grouping pins by hue or by icon
organises them by an attribute the user picked arbitrarily, and grouping by
first letter is worse than sorting. Search and sort cover the real need. "Hide
all" lived inside the tree view, so it moved up into the header where it
applies to whatever the list shows -- it is the one bulk control worth keeping.

The title bar now closes the panel, matching the "Pins" button that opens it;
clicking a title to open but having to hit a small X to close is an asymmetry
that makes the panel feel fiddly. The list also reclaims the height the removed
band was using.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8Mgg1vd2eKSs8StZiuyMV
jakeaturner and others added 24 commits September 22, 2026 16:28
On 1.34.0 every FDA ingest finished with its labels loaded but lost the
installed_resources 'dataset' row, because resource_type was still
enum('zim','map'). The enum migration in #1243 fixes new writes but can't recreate
the lost row. Re-selecting the tier doesn't fix it either, because ZimService
skips a dataset that already has rows. So Medicine never resolved its
Standard tier.

- Add DrugInstallRowProvider. On each boot it backfills the row when an
  ingest has finished (rows present, export date set, no download
  marker, no drug job running) and no row exists.
- Put the decision in decideDrugRowReconcile() so it has unit tests.
- Add DrugReferenceService.recordInstalledRow() as the one writer used
  by the ingest job, the provider, and ZimService.
- ZimService now writes the row when a Medicine tier re-select finds a
  finished ingest with no row.

Manual ingests and reset-and-reingest runs, which dispatch without
resourceMeta, also get the row on the next boot.
scanAndSyncStorage builds sourcesInQdrant and embeddableFiles in the
same pass but only ever checks one direction - is this on-disk file
already embedded? It never checks the reverse, so points left behind
by a deleted or replaced ZIM just sit in Qdrant forever and keep
showing up as search hits for content that doesn't exist anymore.

Added a reverse sweep that diffs the two sets and purges anything left
in Qdrant with no file backing it. Guarded so it's a no-op when the
disk scan comes back empty, so a filesystem hiccup can't get misread
as "delete everything."

Also has to exclude README.md/docs, since those get embedded by
discoverNomadDocs and live outside the kb_uploads/zim paths the
scanner actually walks. First pass without that exclusion would have
wiped the docs KB on every single sync - caught it by testing against
a live install where those sources were already indexed.

Verified live: found exactly one real orphan out of 29 indexed
sources (an old devdocs_en_bash zim superseded by a newer download),
purged it, and left the other 28 - including all the docs - alone.
delete() removed the file, the kiwix library entry, and the
InstalledResource row, but never touched Qdrant or KbIngestState. A
deleted ZIM kept surfacing as a search citation indefinitely. The sync
sweep added in the previous commit would eventually catch these too,
but no reason to wait for the next sync when we already know exactly
what needs cleaning up right here.
ZimService.delete() was calling purgeIndexedSource() unconditionally,
so installs without the AI Assistant (Qdrant) feature logged a false
"offline" error on every ZIM delete. Guard it on Qdrant being
installed, same as reconcileReplacedContentFile() already does.

Also point reconcileReplacedContentFile() at purgeIndexedSource()
instead of its own inline _deletePointsBySource + KbIngestState.remove
pair, so there's one shared implementation instead of two.
purgeIndexedSource() deleted Qdrant points before removing the
KbIngestState row. If the row-delete failed or the process died in
between, the source vanished from Qdrant's orphan-detection signal but
the ghost state row stayed forever — nothing else ever revisits a row
with no matching file and no matching Qdrant source. Reversing the
order means a failure instead leaves points in Qdrant, which the next
sync's reverse orphan sweep will find and retry; both steps are
already idempotent, so a retry from either partial state is safe.

Also switch scanAndSyncStorage()'s orphan-candidate filter from
denylisting known non-scanned roots (README.md/docs) to allowlisting
the two roots _discoverKbFiles() actually scans (kb_uploads, zim),
via a new shared _kbScanRoots() helper. A denylist silently treats any
future source root outside kb_uploads/zim as an orphan the first time
it's added — the same class of bug the denylist itself was written to
fix for docs.
… trip

scanAndSyncStorage()'s orphan sweep purged each orphaned source in its
own sequential await, re-running _ensureCollection and a separate
Qdrant delete call per orphan. Add purgeIndexedSources() /
_deletePointsBySources(), which use Qdrant's match.any filter and a
single whereIn() query to purge a whole batch in one round trip each,
and point the sweep at it. purgeIndexedSource() now delegates to the
batched form for a single source, so there's one implementation
instead of two.
Self-review pass over the last few commits:

- Extracted the kb_uploads/zim allowlist check out of scanAndSyncStorage()
  into a new filterOrphanCandidates() in kb_orphan_decision.ts, alongside
  the existing decideOrphans(). It was untested logic sitting inline in a
  large method; pulling it out next to decideOrphans() makes it a pure,
  I/O-free function the same way decideOrphans() already is, and lets it
  be covered directly instead of only indirectly through scanAndSyncStorage.
- Added unit tests for it: keeps sources under the scanned roots, excludes
  sources outside them (e.g. bundled docs), and a dedicated case for the
  sibling-directory false-positive a naive prefix check (no trailing
  separator) would let through.
- Also fixed two prettier violations introduced earlier in this PR (an
  over-length logger.info call and an unwrapped ternary in
  scanAndSyncStorage) that weren't caught by the per-edit lint checks I
  ran at the time.
decideOrphans guards on the disk scan coming back empty, but the failure
mode it needs to defend against is per-root. _discoverKbFiles() walks
kb_uploads and zim in a loop and skips a missing root rather than
failing, because a fresh install legitimately has no kb_uploads yet.

So if the zim root is relocated, renamed, or not yet mounted (#1050),
the scan still returns a non-empty list from kb_uploads. The empty-scan
guard never fires, filterOrphanCandidates has no way to tell a root that
scanned clean from one that was never there, and every ZIM in the index
is classified as an orphan and purged in a single batch. On a
Wikipedia-scale knowledge base that is millions of vectors and a
multi-hour re-embed. It is reachable from the UI: scanAndSyncStorage is
POST /rag/sync, so the button pressed when the knowledge base looks
wrong is the button that empties it.

_discoverKbFilesWithRoots() now reports which roots were walked, and
filterOrphanCandidates allowlists against that list instead of the
configured roots. A root that wasn't there contributes no candidates; a
root that scanned clean and held no files still yields its orphans, so
the legitimate purge path is unchanged. The empty-scan guard becomes the
degenerate case of the same rule rather than a separate one.

Also builds the unit-test paths with join() instead of POSIX literals.
Written as '/data/...' the filterOrphanCandidates cases pass on Linux CI
and fail on Windows, where sep is backslash and no source ever matches a
prefix, which is exactly where they would be run while changing this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
planPrompt() returned the response reserve as num_predict, and the
reserve was clamped by MAX_RESPONSE_RESERVE = 1024. Chat answers
stopped at 1024 output tokens regardless of window size, often
mid-sentence, and the normalized stream chunks dropped Ollama's
done_reason, so neither the UI nor API clients could discern between
a cut-off answer and a finished one.

- Decouple the generation cap from the reserve. num_predict is now the
  space left in the window after the prompt, less a 5% (min 64 token)
  margin for estimator error, with the reserve as its floor. Raise
  MAX_RESPONSE_RESERVE to 8192 so it only bounds prompt hold-back.
- Carry done_reason through all four OllamaService chat paths. Native
  forwards Ollama's done_reason; compat maps OpenAI's finish_reason.
- Log length-limited stops in OllamaController with num_predict,
  num_ctx and tokens generated.
- Show a truncation notice on cut-off answers, with a Continue button
  on the latest one that sends a normal follow-up turn.

Generation evals before and after this change are not length
comparable.

Closes #1342
The post-rerank relevance floor rejects noise but passes coherent
off-topic questions, and it gets worse as the library grows. Adding
Gutenberg prose, the Python tutorial and MDN pages to the eval corpus
took the off-topic leak rate from 22% to 67% at the 0.62 floor.

Score cutoff does not fix it. The weakest correct top hit scores 0.666
while "write me a haiku about autumn leaves" pulls Walden at 0.705.
A cosine floor and a top-vs-median contrast test both cost recall
faster than they cut leaks.

- Add rag.relevanceCheck (default off). After the floor, the tasks
  model, or the chat model when none is set, gives one grammar-constrained
  on_topic verdict for the whole retrieved set. Off topic means no
  context and no citations. Any failure keeps the chunks unjudged.
- Add a Settings > Models toggle recommending an 8B+ tasks model.
  On llama3:8b the check cuts off-topic leaks from 67% to 6% for about
  5 points of core recall@5 and ~0.6s per turn. On 3B models it rejects
  most valid questions, which is why it is opt-in.
- Add 6 distractor documents and an off_topic suite (18 unanswerable
  questions, 7 distractor controls) to the eval corpus.
- Report per-tag leak rates and leaked cases with top scores in
  eval:retrieval. Add --judge-model to measure the check and --dump to
  export pre-floor candidate pools for offline gate simulation.
- Pin the check off in eval:generation so a local setting cannot move
  scores.
- Record the distractor sweep on RAG_MIN_FINAL_SCORE. The 0.62 default
  stays: 0.65 halves leaks at no measured recall cost but leaves only
  0.016 of margin.

The corpus fingerprint changes to b4ebe5dce8699f5d. Retrieval reports
from before this commit are not comparable, and a new baseline is
committed.

Refs #1341
GpuPassthroughRemediationProvider returned early unless the NVIDIA
runtime was registered with Docker. AMD doesn't register one, so the
log-based CPU-fallback detection added in v1.35.0-rc.1 never ran on
AMD boxes. On a Phoenix1 iGPU running ollama:rocm with /dev/kfd and
/dev/dri passed through, Ollama logged library=cpu while gpuHealth
reported status "ok".

The gpuHealth side had its own bug. hasNvidiaRuntime and
hasRocmRuntime were set only inside the lspci probe branch. When lspci
returned a plausible controller, the else branch reported "ok" without
probing and left both flags false. A blind reinstall would not have
helped either: gfx1103 needs HSA_OVERRIDE_GFX_VERSION=11.0.0, and
without the gfx marker or ai.amdHsaOverride a reinstall rebuilds the
same CPU-bound container.

- Add resolveExpectedGpuVendor(). It reads the nomad_ollama container
  config first (nvidia DeviceRequest, ROCm image tag, /dev/kfd), then
  the nvidia runtime, then the AMD GPU type when no container exists.
  /dev/dri alone is not treated as AMD. A CPU container on an AMD box
  is trusted as deliberate (#1232 fallback, acceleration disabled).
- Gate the provider on the expected vendor. NVIDIA auto-reinstall is
  unchanged but now skips when the container requests a GPU with no
  runtime registered. AMD is detect-only: the provider logs whether a
  reinstall would change the container and never reinstalls itself.
- In gpuHealth, set the runtime flags up front and let Ollama's
  startup log decide health whenever a GPU is expected, whatever lspci
  returned. A CPU-only log now reports passthrough_failed and skips
  the nvidia-smi exec. lspci stays the display source when clean.
- Parse the first GPU "inference compute" line instead of the last
  line, which could be the CPU fallback device, and recognize Vulkan.
  Move the startup-log reader into ollama_compute so both callers
  share it.
- Report amdReinstallWouldChange, amdGfxTarget, currentHsaOverride and
  suggestedHsaOverride on AMD passthrough_failed. The gfx target is
  parsed from Ollama's "amdgpu is not supported" line (best-effort).
- Add GpuPassthroughAlert, shared by Settings > System and Settings >
  Models. When a reinstall would rebuild the same AMD container, it
  offers "Fix: Set GFX Override" instead of a reinstall. The dialog
  saves ai.amdHsaOverride and then reinstalls.
- Expose ai.amdHsaOverride through the settings API. The value lands
  in the container env, so only x.y.z, "none" or empty is accepted.
- Silence the HSA override resolver's logging on the gpuHealth poll.

The #1325 case where no GPU runtime is expected and lspci reports a
controller still returns "ok" unprobed.

Fixes #1344
Refs #1325, #1232
The Command Center header, settings sidebar, chat sidebar and README still
showed the old dotted "PROJECT N.O.M.A.D." badge with the retired backronym.
The earlier naming sweep removed the text strings but could not catch a
raster asset.

Swap in the vector primary badge from the brand kit (nomad-primary.svg) and
add object-contain so the hexagon keeps its true 0.866 aspect instead of
being stretched to fill the square boxes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Upgraded AMD iGPU installs sometimes carry a working
HSA_OVERRIDE_GFX_VERSION on nomad_ollama, set by hand or by an older
install path, but have no ai.amdHsaOverride KV and no .nomad-amd-gfx
marker. The CPU-fallback banner offered Fix: Reinstall, and the
reinstall rebuilt the container without the override, so a 780M
(gfx1103) stayed on CPU.

- Add pickAmdHsaOverride to utils/amd_hsa_override.ts with the
  resolution order KV, marker, existing container override, none.
  A marker still wins even when it maps to none. KV "none" still
  disables the override.
- forceReinstall inspects nomad_ollama before removing it and saves a
  container-only override to ai.amdHsaOverride. Saving it keeps the
  value through a reinstall that fails after the removal and shows it
  in the Set GFX Override modal.
- The update path passes the old container env to the resolver, so a
  routine Update keeps the override too.
- The GPU health check passes the container env, so
  diffAmdOllamaConfig compares against what a reinstall would build.
  On the reported box, the only drift is OLLAMA_IGPU_ENABLE=1.
- diagnoseAmdCpuFallback suggests the container's current override
  when one is set. With an override applied, Ollama logs the coerced
  target (gfx1100), which maps to none and would suggest removing a
  working value.
- Remove the unused _mapGfxToHsaOverride wrapper.

parseAmdGfxTargetFromLogs is unchanged. It still matches only
gpu_type=. The "dropping integrated GPU ... compute=gfx..." line needs
a full sample from an affected box before we can write a regex for it.
We'll address in a later PR.
Boot runs ensureDirectoryExists() on the zim root, so a separate volume
that failed to mount leaves an empty mountpoint behind. The scan walks it
without error and reports it as scanned. One upload in kb_uploads keeps
the overall scan non-empty, so the next manual sync purged the vectors
for every ZIM under that root.

decideOrphans() now decides per scanned root and returns
{ orphans, withheld }. A root keeps its indexed sources when:

- it holds no embeddable files (kiwix-library.xml does not count)
- the purge would remove more than 50% of its indexed sources and at
  least 5 of them, which points to a wrong or stale mount rather than
  hand deletions

scanAndSyncStorage() logs each withheld root and appends a note to the
sync result message, so the user sees a likely unmounted drive instead
of losing hours of embeddings.

Replaces the unit test that asserted the old empty-root purge with
regression coverage for #1378 and both guards.

Closes #1378
setFileActive updated the Qdrant payload but skipped the DB write when
a source had no kb_ingest_state row. ZIMs hit this during embedding
(markIndexed runs only after the final batch), after a lost row, or on
pre-RFC installs. The API returned success, getStoredFiles still showed
the file as active, and the switch could never turn it back on. All
chunks stayed excluded from retrieval.

When the row is missing, create it the same way the scanner backfills:
indexed if Qdrant has chunks for the source, pending_decision if not.
Then save the flag.

Creating the row does not affect retrieval, since search filters only
on the Qdrant payload. It also leaves failed or partial embeds no worse
off:
- A failed job already leaves a row via markFailed. Only active changes.
- A crashed worker leaves partial chunks and no row. The next scan
  writes the same indexed row through backfill_indexed.
- For an in-flight job, markIndexed/markFailed reuse the row and keep
  active, so later batches stop stamping chunks active: true.
- The UI already treats a null state as indexed, and stall/zero-chunk
  warnings come from Qdrant counts, so Re-embed still appears.

A partial ZIM with no row is still labeled indexed, as before.
markStalled has no callers, so stalls are not tracked in the row.
embedAndStoreText read the file's active flag, then upserted its
chunks. A toggle landing between those two steps set its payload
before the new points existed, so the new points kept the old value.
Retrieval and the KB panel then disagreed.

Both toggles (setFileActive and setKnowledgeCollectionActive) now write
kb_ingest_state before calling Qdrant setPayload, and restore the rows
they changed if setPayload fails. embedAndStoreText re-reads the flag
after its upsert and sets active on its own point IDs if the flag
changed. Whichever order the two run in, one of them writes the final
value to the new points.

Also coerce the flag to a boolean before stamping it into the payload.
MySQL returns tinyint(1) as 0/1, and a raw 0 got past the search
filter's must_not active == false, so an inactive file that was
re-embedded became searchable again.
The vision attachment feature landed in edbfe1a with Gabe Graves as the
commit author, but under the legacy gabegraves@users.noreply.github.com
address rather than the ID-prefixed noreply GitHub associates with the
account. GitHub therefore did not count it toward his contributions.

This empty commit carries the correct trailer so the credit is restored
once dev fast-forwards onto main. No files change.

Original implementation: #1178. Integration and rebase: #1288.

Co-authored-by: Gabe Graves <37971265+gabegraves@users.noreply.github.com>
@jakeaturner jakeaturner reopened this Sep 29, 2026
@jakeaturner jakeaturner changed the title Dev 1.35.0 Sep 29, 2026
@jakeaturner
jakeaturner merged commit ddb0235 into main Sep 29, 2026
2 checks passed
@cosmistack-bot

Copy link
Copy Markdown
Collaborator

🎉 This PR is included in version 1.35.0 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.