Flowkey survives a Python upgrade. Removing the interpreter Flowkey's virtual environment was built against broke every hotkey with a modal Windows dialog, and re-running the source installer repaired nothing. Found on a live machine.
- "Python venv launcher is sorry to say ... did not find executable" on every hotkey. A virtual environment's
Scripts\pythonw.exeis not an interpreter — it is a ~250 KB stub that re-execs the interpreter recorded inpyvenv.cfg. Upgrading Python 3.13 to 3.14 uninstalls that interpreter but leaves the stub on disk, so the resolver's existence check still passed and every action spawned a dead launcher that hung on a modal dialog. The virtual environment is now accepted only while its base interpreter still exists, verified by readingpyvenv.cfgrather than by running the stub — running it to find out is precisely what raised the dialog. - Locating Python no longer assumes an install layout. The rung below the virtual environment was the bare name
pyw.exe, which the PSF Python Manager installer does not ship at all (it installspythonw.exeunder%LOCALAPPDATA%\Python\bin), so deleting the stale environment would only have moved the failure. Discovery now walks the PEP 514 registry entries every conformant Windows Python writes, newest 3.11+ first, and resolvesPATHby hand so the zero-byte Microsoft Store alias stubs that shadow real installs are rejected rather than launched. - FastFlowLM reported as "not installed" on a machine where it was installed. A process only ever sees the environment block built when its session started, and Flowkey normally launches at logon — so FastFlowLM installed (or repaired) afterwards appended itself to the machine
PATHwhere Flowkey could never see it.flmanswered fine from any new shell whiledoctorsaidfastflowlm_cli: not foundand every hotkey failed. Provider CLIs are now located by an explicit resolver —PATH, thenPATHas the registry currently holds it, then the known install directories — andargv[0]is passed as an absolute path. That last part matters on its own: Windows resolves a bareargv[0]against the parent process'sPATH, so repairing the child's environment does not affect the lookup. Detection and execution now share one resolver, sodoctorcan no longer contradict the running app. flm validateran on the raw inherited environment. It was the onlyflmcall site that did not go throughflm_env(), so it alone missed both thePATHrepair above and the 2.5.2FLM_MODEL_PATHrepair.install.ps1repairs an unhealthy virtual environment instead of reporting success. It carried the same existence-is-health assumption, so re-running the installer on an affected machine printed "venv already present" and fixed nothing. It now checks that the environment's base interpreter exists, rebuilds it when it does not, and finds the interpreter to rebuild with using the same rung order as the app.
The local server starts after a reboot, and start-with-Windows works again. Both were silent failures with no error anywhere; both were found on a live machine.
- FastFlowLM could not start when launched by Flowkey. FastFlowLM's installer stores its model directory as a machine-scope
REG_EXPAND_SZcontaining%USERPROFILE%\.flm. Machine-scope variables are expanded in the SYSTEM context when a login session's environment block is built, so every process inheriting it sawC:\Windows\system32\config\systemprofile\.flm— a directory it may not create — andflmexited immediately withcreate_directories: Access is denied. This is why the failure appeared only for the app (explorer → AutoHotkey → daemon → flm all inherit that block), never from an interactive shell, survived a reboot, and was not fixed by upgrading FastFlowLM to 1.0.3. Flowkey now repairs the value for everyflmchild it spawns rather than trusting the machine environment. Upstream bug; Flowkey is now immune to it. - Logon autostart silently launched nothing. The autostart command fell back to the bare string
AutoHotkey64.exewhenever the installed layout was absent — which is every source tree, where AutoHotkey lives undervendor\ahk. AutoHotkey is bundled rather than onPATH, so Windows resolved nothing at logon while the Run entry still read as enabled. The command is now always a resolved absolute path,get_autostart_statereports avalidflag, and the daemon repairs an enabled-but-unlaunchable entry at startup (repair only — it never creates one the user did not enable).
- Re-download uses FastFlowLM's own
--forceinstead of remove-then-pull. 2.5.1 deleted the model first becauseflm pullalone skips models that are already present;flm pull --forcere-downloads without deleting, so the existing copy now survives a failed re-download and the destructive warning is gone.
Provider failures are diagnosable, and an installed model can be re-downloaded. All three items came out of a real debugging session on a live machine.
- Re-download… button in Config → Models. Re-fetches a model that is already installed — needed after a FastFlowLM upgrade invalidates local weights, where the model still lists as installed but the runtime rejects it. For FastFlowLM this is remove-then-pull, because
flm pullonly downloads a model that is missing, so a plain re-pull is a silent no-op. The UI confirms first, and a failed forced pull states plainly that the old copy is gone and the pull must be retried. Ollama is left to its ownpull, which already re-fetches when the remote digest changes.
- Provider startup errors are no longer discarded.
server.log_to_filedefaults to false, and on that path the local server's stdout/stderr went nowhere — so a refused start surfaced only asFastFlowLM server exited early (exit 1)with nothing to act on. Provider output is now always captured (to a scratch file that is deleted if the server comes up, or to the persistent log when logging is enabled), and the last lines are appended to the error.log_to_filenow controls only whether the log persists, never whether a failure can be diagnosed. - A stale version check is no longer presented as current. The non-blocking FastFlowLM update read could serve a cache days past its TTL and render it as fact; the network-failure fallback also served cached values while reporting itself as uncached. Stale readings are now flagged, labelled in the UI, and refreshed once in the background.
Notes now works like a notepad, sticky-note wall, and vision board instead of a file browser. Capture first, shape the note in an editor, and organize it without leaving the Notes tab.
- Living Notes workspace. Notes render as responsive sticky cards with dedicated types for notes, tasks, ideas, links, and read-later items. The composer/editor supports title, body, category, tags, color, status, due date, source link, pinning, Archive, and Trash.
- Smart organization. Search, type views, Pinned/Archive/Trash views, category facets, and tag facets work over an indexed vault feed.
- Vision Board. Editable sections and drag-and-drop card ordering are persisted separately in
.flowkey/board.json, keyed by stable note IDs. Removing a board placement never mutates or deletes the note. - Recoverable removal. The default delete action moves a note to Trash. Restore is immediate; permanent deletion is available only inside Trash and requires explicit confirmation.
- Safe schema-v2 migration. Existing Markdown notes gain a stable
note_idand richer metadata without losing their body, source path, or unknown frontmatter. A backup is written to.flowkey/backups/v1/before the first rewrite. - Conflict-aware, atomic writes. Notes and board state use atomic replacement and revisions, so a stale editor reports a conflict instead of overwriting newer work.
- The Notes tab contains notes only. Vault, category, extraction, and local-model controls moved to Config → Notes & capture and save with the rest of Config.
Ctrl+Alt+Nopens the Notes composer. A fresh selection prefills the draft; no selection opens a blank composer. Prior clipboard contents are never used as an implicit fallback, and capture no longer saves before review.- Local-model enrichment cannot rewrite authored content. It may fill only blank or generated metadata, preserving user-written titles and bodies.
- Legacy note actions remain compatible. Existing callers can continue reading or moving by relative path while the new workspace uses stable IDs.
- Date-only due dates keep their calendar day. A value such as
2026-08-04displays as August 4 in every timezone instead of shifting to the prior day west of UTC. - Note migration no longer races concurrent edits or vanishes notes. Upgrading a pre-2.5 note to the new schema is now serialized against other note writes, and a note whose file can't be rewritten (e.g. a read-only location) still appears in search and listings instead of silently disappearing.
- "Move to Trash" respects the same conflict check as every other note action. If a note changed since you opened it, trashing it now reports the conflict instead of acting on stale data.
- A Quill hiccup no longer breaks Meetings. A transient error from Quill during a scheduled digest run, the Overview widget, or a meeting Q&A now degrades gracefully instead of surfacing as an error or silently skipping a whole batch.
prompt: handles vague requests properly. A request too thin to break into requirements now produces a prompt that names what is missing, instead of handing the request back reworded.
- A vague request no longer returns itself. Given "develop a app that allows to perfrom a full schedule for my meeting a proper PM would",
prompt:emitted that same sentence as the task, again as its only real constraint, and again as the output format, padded with two generic scope lines. It was structurally correct and invented nothing, so every automated check passed — but it told the reader nothing they had not just typed. When a request cannot be decomposed into distinct requirements, the output now names the open questions (platform, scope, data source, success criteria) and instructs the agent to settle them before building. It still adds no requirements: each line asserts only that something was not specified. - Typos and article agreement are corrected in the output. Because the prompt is built from your own wording, "a app" and "perfrom" were repeated in every section. Surfaced text now fixes article agreement and a fixed list of common misspellings, without changing meaning. Well-known exceptions ("a user", "a unique", "a one-off") are preserved.
- New rubric item R8 — no section may restate the task. The eight-item rubric now fails any output whose constraints contain nothing beyond a restatement of the task plus boilerplate, or whose output format re-quotes the request. R8 is disqualifying on its own, so a structurally valid echo can no longer pass the release gate at any score. It is machine-checkable — no judge required.
- The reported request is now a permanent case in the fixed evaluation set, and R8 was validated against all 14 cases with zero false verdicts (it fails the old echo and clears every legitimate output, including multi-clause requests whose first bullet naturally restates the opening sentence).
Clear failures instead of cryptic ones. A benchmark that cannot fit in memory is now refused with an explanation, keep-warm no longer fights an in-progress benchmark, and errors the local server reports are passed through instead of being replaced by a generic message.
- Benchmarks that can't fit are refused up front. Running
flm benchon a model too large for the machine failed with a raw driver code —Failed to submit command to hw queue (0xc01e0200): … the video memory manager could not page-in all of the required allocations— which said nothing about the cause. The Benchmark tab now checks the model's measured footprint plus room for the 1k–32k context sweep against available memory first, and explains: "needs about 28.3 GB (~24.3 GB of weights plus room for the 32k-context sweep) but only ~25.6 GB is usable". A model with no published footprint is never blocked. - Keep-warm no longer collides with a benchmark. The background warmup reloaded the active model on its own schedule with no knowledge of benchmarks. Since a benchmark takes 10–20 minutes and the default keepalive is 15, a warmup reliably landed mid-run and competed for the same NPU memory. Warmups are now skipped while a benchmark is running (and logged as such).
- Errors from the local server are shown as-is. FastFlowLM reports some failures as an HTTP 200 response carrying an error message — for example
Failed to load <model> model!when a model's weights don't fit. Both the hotkey path and chat only looked for a completion, so that message was discarded and replaced with "Local LLM returned no usable text". The real message now reaches you. A warning alongside a valid completion is ignored as before.
Model picker fixes. The Models card is one card again, the picker follows the app's styling, no model is hidden from it, and an active model invalidated by a FastFlowLM upgrade is now called out instead of failing silently.
- No model is hidden from the picker any more. Suggestions used to drop every candidate the size heuristic judged unfit, which made real catalog models invisible —
qwen3.6-moe:35b-a3bnever appeared and had to be pulled by hand. Every model FastFlowLM offers is now listed; oversized ones are shown with a "too large" marker and still ask for confirmation before downloading. - Mixture-of-Experts models are sized correctly.
…:35b-a3bwas read as a 35B model when only ~3B parameters are active per token. Sizing now prefers FastFlowLM's own measured footprint (sogemma4-it:e4b, which reads as "4B" by name but is really 8B / 9.1 GB, is also judged correctly) and falls back to active-parameter counts for MoE tags. Fit labels now show the real memory cost, e.g. "~24.3 GB needs most of ~25.6 GB usable". - A model broken by a FastFlowLM upgrade is now reported. FLM 0.9.45 rejects models whose local weights were stamped for an older version and reports them as not installed; nothing surfaced that, so the next hotkey failed with an opaque provider error. Config → Models now shows a warning naming the model and the remedy.
- "Installed models" and "Pull a new model" are a single "Models" card with
Installed/Add a model/FastFlowLM runtimesections, instead of two separate cards. - The model picker is styled by the app. It replaced a native
<datalist>dropdown and a native<select size>list box — browsers draw datalist popups in their own chrome and ignore page CSS entirely, so that control could never match the dashboard. The new combobox and list are ordinary elements: they follow the light/dark theme, and support arrow-key/Enter/Escape navigation and type-to-filter.
Streaming chat. The dashboard Chat tab renders replies token-by-token instead of waiting for the whole answer, so the response starts appearing at warm time-to-first-token rather than after the full completion.
- The Chat tab streams replies as they generate. Instead of a blank "Thinking…" wait until the whole answer lands, the assistant's reply now fills in token-by-token over a Server-Sent-Events stream — first text typically appears at ~1.6 s (warm TTFT) rather than after the full completion. Measured live on FastFlowLM
qwen3.5:4b: first token at 1.52–1.58 s versus a ~2 s (short) to tens-of-seconds (long) full-completion wait. Prompt-mode/grammar hotkeys and the AHK paste path are unchanged (they still return whole output). A daemon without the streaming endpoint, or any failure to open the stream, transparently falls back to the previous one-shot request, so nothing regresses. Works on both FastFlowLM and Ollama (shared OpenAI-compatible SSE). A mid-stream provider drop or a client disconnect still saves the partial reply.
- Streaming persistence re-reads the thread store under the daemon write-lock (atomic read-modify-write), so a concurrent chat write or thread delete during a multi-second stream is never clobbered by a stale snapshot; a client disconnect (
GeneratorExit) still persists the partial turn. Both were caught by an adversarial review pass and pinned by discriminating tests (SPEC V37 / T27).
Prompt mode, faster and better grounded. The default prompt: path now emits a compact copy-paste-ready agent prompt in a few seconds without filling gaps with conventional-but-unstated requirements.
- Prompt v2 is the new default. One short local-model draft is finalized into exactly four ordered sections (
<task>,<context>,<constraints>,<output_format>) using only source clauses plus fixed scope guards. The finalizer prevents raw-model guesses about libraries, files, arguments, formats, tests, platforms, defaults, or error behavior from surfacing. The legacy prompt remains available immediately through Dashboard → Config → Prompt builder → v1 — legacy rollback. - Prompt decode is bounded. Prompt short/medium/long strategies now cap at 240/320/420 tokens, and default v2 makes exactly one draft call rather than entering anti-echo/rescue retries before its grounded finalizer.
- FastFlowLM stays warm. The daemon performs a best-effort background warmup on startup and at a configurable keepalive interval (15 minutes by default;
0disables periodic warmups). Failures are logged and never block daemon startup. The dashboard exposes both controls.
- Fixed set: 12 realistic requests covering implementation, debugging, review, refactor, data, vague, long, and trap inputs; one warmup plus five timed generations per style/input on FastFlowLM
qwen3.5:4b. - Speed: warm p50 18.18 s → 3.38 s, p90 22.69 s → 4.67 s, median completion tokens 222 → 26. v2 is 18.59% of v1 p50; no ordinary v2 input exceeded 25 seconds.
- Quality: manual GPT-5 source review scored v2 median 7/7 versus v1 0.5/7; v2 passed 12/12, clean-section rate was 100%, and invented-requirement failures were 0 (v1: 12).
- Warm-model probe: first post-restart request 20.88 s wall versus immediate warm request 3.70 s (5.65× wall-time improvement).
- Evidence:
data/benchmarks/prompt_v2_ab_2026-07-10.json,prompt_v1_frozen_2026-07-10.json,prompt_v2_judge_2026-07-10.json, andprompt_v2_cold_warm_2026-07-10.json.
- FastFlowLM force-restart now waits for the old socket to close before spawning the replacement, preventing a dying instance from being mistaken for
already_running. - Added a reproducible speed/quality evaluator with usage-duration aliases, manual/LLM judgment rescoring, frozen-v1 replay/export, cold/warm probing, and five-output side-by-side review evidence.
Maintenance. Ships the History visibility controls and clears the remaining SPEC audit backlog.
- History visibility controls. The History tab now defaults to a privacy-safe Telemetry view, adds an Exposed view for request/result text that was actually stored, and surfaces the redacted/visible capture setting in context. Rows captured while redacted remain text-free.
- Removed dead release-install shims (
setup/install_release.{cmd,ps1,sh},setup/bootstrap_release.sh) — superseded byscripts/install.py. - Installer README layout diagram corrected to the flattened
{app}bundle (matchesinstaller.iss). - Provider roadmap updated — the provider selector + per-provider status UX shipped in 2.0; only side-by-side dual-provider and multi-device sync remain future work.
- New README↔dashboard tab-count parity test guards against tab-list doc drift.
Prompt builder controls. Flowkey's prompt: mode can now target Claude Code or generic chat tools without exposing raw built-in system-prompt editing.
- Prompt builder settings for
prompt:mode. The Config tab now exposes bounded prompt-output controls (target agent, action, detail, structure, acceptance criteria, verification, output expectations, and a capped user suffix) without allowing raw edits to the locked built-in system prompt. The default remains Claude Code-compatible and keeps the existing prompt-mode system prompt byte-identical;generic_chatadds a production Markdown adapter with target-aware validation and deterministic fallback/preview. Live FastFlowLM eval onqwen3.5:4b:claude_code2/2,generic_chat10/10.
Maintenance. Fixes a recurring meetings-batch error and clears a batch of small audit/doc items.
- Meeting batch no longer retries content-less meetings forever. A meeting Quill has neither minutes nor a transcript for (aborted/duplicate stub recordings) used to re-queue and re-fail on every batch run, so "Run batch now" and the after-hours scheduler kept reporting errors (e.g. "processed 0 of 5, 5 errors"). Such meetings are now marked skipped (
data/meeting_skips.jsonl) and excluded from future queues — but only when the meeting is older than 2 days, so a just-ended meeting whose transcript is still syncing keeps retrying. LLM/provider failures are never skip-marked. The batch result now reports askippedcount, and per-meeting error reasons remain inlogs/daemon.log. - The
open_chathotkey default is nowCtrl+Alt+C(^!c) everywhere. The retiredCtrl+Shift+T(^+t) still lingered as a fallback in the first-run wizard, the daemon's config snapshot, and the dashboard hotkey field, so a fresh config could surface the colliding old binding.
- First-run wizard text updated ("Open chat popup" → "Open chat") — the tkinter chat popup was retired in 2.0.
- Installer bootstrapper (
bootstrap.cmd) no longer prints a hardcoded stale version in its build banner. - README dashboard tab list corrected (added Benchmark → 8 tabs).
- New seed-vs-schema drift guard test (
test_config_seeds) freezes the known delta between the shipped first-run seed andDEFAULT_CONFIG.
Meetings, on your own time. Flowkey can now read your local Quill meetings and answer questions about them on the local model — and an after-hours scheduler pre-computes each meeting's digest (summary / goals / action items) during your idle window, so daytime reads are instant.
- Meetings (Quill integration) + after-hours digest processing. Flowkey can connect to the local Quill note-taking app over MCP to search your meetings and answer questions about them — entirely on the local model. Because asking a model about a full transcript costs real prefill time (~15–17 s of time-to-first-token for a ~7k-token transcript on the NPU), a background scheduler pre-computes a digest (summary / goals / action items) for each meeting during a configurable idle window (default 17:00–21:00, only when the machine has been idle), caching it in
data/meeting_digests.jsonlso daytime reads are instant. New Meetings dashboard tab (search → read cached digest, "Process now", or "Ask about this meeting") and a Config card for all the settings (enable, Quill MCP URL, content source, schedule window, idle gating, max-per-run) with a "Run batch now" button. Off by default (opt-in). New modulesffp_quill(stdlib MCP-over-HTTP client) andffp_meetings(digest store, batch worker, scheduler logic, idle detection); new config blockmeetings; new daemon actionsquill_status/quill_search_meetings/meeting_digest_get/meeting_digests_list/meeting_process/meeting_batch_run/meeting_batch_status/meeting_ask. The Quill MCP URL is validated loopback-only; the meeting actions write a separate cache file under their own lock, so a long after-hours batch never blocks config saves or notifications. - Action-item review board. A weekly/monthly board on the Meetings tab aggregates the action items parsed from your meeting digests, each markable accepted / rejected / pending (status persisted in
data/meeting_action_status.jsonl). New daemon actionsmeeting_actions_list/meeting_action_set_status. - Weekly review. A one-click roll-up of the week's processed meetings (highlights / themes / open items) generated on the local model from the cached digests — pick the week and Generate. New daemon action
meeting_week_summary. - Meeting hours on the Overview tab — today / this-week meeting counts and hours, pulled from Quill (new daemon action
meeting_overview). - Digest quality flags + strict re-digest. Each meeting digest is checked for low-substance / social-filler / too-short / trivial-meeting signals and flagged in the Meetings tab; a "Re-digest (strict)" button re-runs the summary with a stricter prompt. New daemon action
meeting_redigest.
- Benchmark history no longer shows a blank row for an interrupted/empty run.
- "Run batch now" persists the current Meetings settings (including the Enable toggle) before running, so it reflects the form rather than the last-saved config.
- The Meetings results list is now scrollable and shows a meeting counter.
- Packaging now ships all runtime modules.
ffp_meetings,ffp_notifications, andffp_quillwere missing frompyproject.tomlpy-modulesand the PyInstaller spechiddenimports, so a wheel / frozen installer could omit them and crash on import even though source-tree tests passed. All three are now declared, and a new test (test_packaging_modules) assertspy-modulesand the spec stay in sync withscripts/*.py. - The Telemetry time-of-day chart now renders only active hours — zero-activity hours are dropped instead of drawn as empty bars (and an empty history shows "No activity yet").
- Autostart unified to a single per-user entry. Three independent autostart registrations had drifted out of sync: the daemon wrote
HKCU\...\Run\FastFlowPrompt(what the dashboard toggle reads/writes), a source install (install.ps1) wrote a different value name (HKCU\...\Run\Flowkey) the toggle couldn't see, and the packaged installer optionally wrote a third, machine-wideHKLM\...\Run\Flowkeyentry — enabling the dashboard toggle after either install path could add a redundant entry and launch the app twice at logon. All three now agree on the one per-userHKCU\...\Run\FastFlowPromptentry the daemon owns; the packaged installer no longer offers a machine-wide autostart option, and its uninstaller now removes the per-user entry. Guarded by a newtest_installer_autostartregression test.
- Removed the unused
ffp_tools.pytool-calling prototype (no runtime caller). - Unified the config seed templates: the dev example (
config/) and the shipped first-run seed (setup/defaults/) had drifted (the seed was missing thellmblock) — they're now identical, enforced bytest_config_seeds. - Docs refreshed for the current UI/build: README dashboard tabs (Chat + Meetings), the
open_chathotkey (Ctrl+Alt+C), a new "Supported surfaces" section (web chat only, installer-vs-source install, per-user autostart), and the installer README (3 exes, dropped retiredffp-chat.exereferences).
The dashboard release. Two things change what Flowkey is in 2.0:
-
It no longer requires an AMD NPU. Flowkey now runs on Ollama (any CPU/GPU) as a first-class secondary provider alongside FastFlowLM (AMD Ryzen AI NPU). A provider-neutral
llmconfig block routes chat, grammar modes, model management, and server startup to whichever backend you pick — and falls back to the other automatically if the configured one isn't available. The first-run wizard detects what's installed and recommends a provider, and model suggestions are now hardware-aware (they scale to your RAM/VRAM and hide models that won't fit). -
The web dashboard is the home for everything. The browser dashboard served by the local daemon at
127.0.0.1:52650is now where you chat, browse and organize notes, manage models, run benchmarks, tune settings, and control notifications. The standalone tkinter chat popup and the old native AHK dashboard are both retired. Chat is a tab (with notes-grounded answers), Notes is a full organizer (read / re-file / delete), and a new Notifications panel plus a Telemetry feed give you per-event control over desktop toasts (with quiet hours, Do-Not-Disturb, dedupe, and a log of everything shown or muted).
Everything still runs locally — no cloud, no analytics, no telemetry leaves the machine.
- Ollama as a secondary LLM provider. Machines without an AMD Ryzen AI NPU can now run Flowkey on Ollama: a provider-neutral
llmconfig block (with per-provider profiles underproviders.*) routes chat, grammar modes, model list/pull/remove, and server startup to FastFlowLM or Ollama. If the configured provider is unavailable the daemon falls back to the other one automatically. - The web dashboard is provider-aware: a Provider selector in Config shows installed/running status for both backends, a "Start server" button starts the active provider (including
ollama serve), model cards relabel per provider, and FastFlowLM-only controls (runtime update check, performance modes, benchmark) hide or disable when Ollama is selected. The Overview tab shows the active provider, including fallback ("FastFlowLM (fallback from Ollama)"). - First-run wizard detects both providers, recommends an available one, and offers per-provider model defaults.
- New daemon action
provider_status;statusoutput now includes the provider. - Pull any Ollama model from the dashboard. "Pull a new model" is now a free-text field with per-provider suggestions — type any name from the Ollama library (e.g.
mistral:7b) and download with live progress. FastFlowLM keeps its catalog suggestions. - Benchmarks work on Ollama too. The Benchmark tab runs timed generations against the running Ollama server (three prompt sizes × two passes, ~1–3 min on CPU) using Ollama's native prefill/decode counters, and records the same TTFT / prefill / decode metrics as
flm bench. The server keeps serving during the run, and history rows now show which provider produced them. - Hardware-aware model suggestions. The dashboard detects system RAM (GlobalMemoryStatusEx) and GPU VRAM (nvidia-smi, or the display-class registry's qwMemorySize — which also finds Ryzen AI iGPU carve-outs) and computes a per-provider size budget: FastFlowLM scales with installed RAM (32 GB ≈ 4B-class on the NPU, 64 GB ≈ 9B), Ollama with VRAM (e.g. 8 GB ≈ 9B) or conservatively with RAM on CPU-only boxes. The pull card shows the detected budget, suggestions hide models that don't fit (near-misses are marked "tight fit"), and free-typing an oversized model asks for confirmation. New daemon action
model_recommendations; new moduleffp_hardware. - Chat moved into the web dashboard. Chat is now a Chat tab in the daemon-served dashboard —
Ctrl+Shift+Tand the tray "Open Chat" open it, andCtrl+Shift+Asends the current selection there (prefilled). Threads, history, and the "ground answers in my notes" toggle work as before; the standalone tkinter chat popup is retired. The dashboard also deep-links by URL hash (/#chat). New moduleffp_chat; new daemon actionschat_threads_list/chat_thread_get/chat_send/chat_thread_delete/chat_stage_selection/chat_take_staged. - Read, re-file, and delete notes from the dashboard. The Notes tab is now an organizer: click a note to read its full body and source link, move it to a different bucket (the LLM's pick is no longer final), or delete it. New daemon actions
note_get/note_move/note_delete— all vault-contained and path-traversal guarded. - Notification settings + a notifications feed. A new Notifications card in Config controls desktop toasts: a master on/off, per-event toggles (action results, clipboard suggestions, settings changes, updates, diagnostics, app lifecycle, errors), a configurable dedupe window, Do-Not-Disturb, and quiet hours — errors and warnings always come through while DND / quiet hours are on. Every notification (shown or muted) is recorded and surfaced as a feed in the Telemetry tab. All policy lives in the daemon (
ffp_notifications), which classifies each toast by message pattern, so the ~45 existing notification call sites were left untouched; the AHK front-end consults the daemon and fails open (shows the toast) when the daemon is unreachable. New config blocknotifications, new moduleffp_notifications, new daemon actionsnotify_gate(decide + log) andnotifications_log(read).
-
Destructive actions use an in-page confirmation, not a browser pop-up. Deleting a chat thread or note, removing a model, deleting a custom mode, pulling an oversized model, and starting a benchmark all previously used the native
confirm()dialog; they now use a styled in-dashboard modal that matches the rest of the UI. -
The tray "Open Chat" entry shows your real hotkey. It was hardcoded to
Ctrl+Shift+Teven after you rebound the key; it now reflects the configuredopen_chatbinding (and updates when you change it). The defaultopen_chathotkey also moved offCtrl+Shift+T(which collides with the browser "reopen closed tab") toCtrl+Alt+C, mirroring the Alt-based note-capture key. -
Hotkey actions no longer trample each other. Grammar fix, note capture, and ask-in-chat share the clipboard; firing one while another's model call was in flight (10–30 s) could corrupt the clipboard save/restore dance, and a re-press of the same hotkey re-ran on stale state. A busy guard now makes them mutually exclusive — a second press gets a "still busy" toast instead.
-
Your clipboard comes back immediately. The grammar hotkey used to hold the captured selection in the clipboard for the whole model call; it is now restored within milliseconds (the result still lands in the clipboard for the paste), restores happen on every path including mid-capture errors, and a busy clipboard at restore time is retried instead of silently losing your copy.
-
The chat window now watches its parent via a kernel wait instead of spawning
tasklistevery 5 seconds forever; a hung toast PowerShell is killed instead of orphaned; the update-available dialog auto-dismisses after a minute. -
Outputs no longer mention emoji out of nowhere. Every built-in mode prompt told the model to "preserve emoji", and small models parroted that into results — grammar fixes ended with emoji remarks and
prompt:mode invented constraints like "Include the emoji 🌟 at the very end". The prompts no longer name emoji (preservation falls out of "leave everything else exactly as written" — verified on qwen3.5:4b: emoji kept in place, no commentary), and built-in prompts are now always sourced from code at load time, so stale copies in existing config files can't resurrect old wording. Only your tone-preset choice and custom modes are kept from config. -
The bare-CLI output path crashed with
UnicodeEncodeErrorwhen the model output contained emoji (Windows pipes default to the charmap codec); stdout/stderr are now forced to UTF-8. -
Installer builds no longer ship a broken chat window and first-run wizard. The PyInstaller spec excluded
tkinter, whichchat_popup.pyandfirst_run.pyboth import — so the frozenffp-chat.exe/ffp-first-run.execrashed at launch withModuleNotFoundError. Tk is now bundled (verified:_internal/tkinter+_tkinter.pyd+ tcl/tk present in the freeze).
- Chat is now daemon-backed (
ffp_chat— thread store reusingdata/chat_threads.jsonl, provider-resolved/v1/chat/completions, optional notes-vault grounding) and rendered by the web dashboard's Chat tab. The standalonechat_popup.pytkinter app and its127.0.0.1:52640ingest socket (chat_send_selection/chat_reload/chat_restart) are removed;Ctrl+Shift+Anow stages the selection viachat_stage_selectionand the Chat tab picks it up. The PyInstaller freeze is now three executables (ffp-daemon/ffp-grammar-fix/ffp-first-run) instead of four. - New modules
ffp_provider_status(detection/capabilities) andffp_provider_runtime(model list/pull/remove routing); registered in the wheel and PyInstaller spec. Provider roadmap notes live indocs/provider-and-sync-roadmap.md. - Dead-code cleanup: removed the redundant
_deep_merge/_version_tuplewrappers ingrammar_fix.py(callers now useffp_config.deep_merge/ffp_updater.version_tupledirectly, and their unit tests moved next to the real functions), deduplicated the update-feed-URL config lookup, and dropped the shadowedchat.llm_auth_bearerkey from shipped configs —llm.auth_beareris the live setting and always wins; old user configs that still carry the chat key keep working. - CI: the AutoHotkey syntax-check job no longer fails when the Chocolatey community feed has a transient outage — the install retries Chocolatey, then falls back to the official AutoHotkey v2 release zip, and fails fast if neither yields the interpreter (so the parse-check can never silently skip and pass green).
- Tests: added coverage for the security-sensitive self-update path (
ffp_updater— zip-slip extraction guard, update-feed parsing, sha256-verified package swap with rollback) and the notes capture pipeline (capture_notestub write + URL detection, the_safe_categorytraversal guard, LLM-JSON recovery, HTML extraction, and frontmatter serialization). Suite is now 180 tests. - CI: bumped
actions/checkoutv4→v5 andactions/setup-pythonv5→v6 onto the Node 24 runtime ahead of GitHub's Node 20 removal, and added anode --checksyntax gate for the dashboard'sapp.js. - CI: added a
Build & release installerworkflow (manual dispatch orv*tag) that PyInstaller-freezes the four executables, compiles the Inno Setup installer, and uploads the artifact (attaching it to a GitHub Release on tag). Uses Node-24 actions (checkout@v5,setup-python@v6,upload-artifact@v7). - Installer build:
build.ps1 -BundleAhknow fetches AutoHotkey v2 from its GitHub release (pinned 2.0.26) instead ofautohotkey.com, which began returning a Cloudflare bot-challenge page to non-browser clients (the download succeeded butExpand-Archivefailed on the HTML).autohotkey.comis kept as a fallback. - Installer compile:
installer.issnow actually compiles end-to-end (the build workflow's first green run). Three latent bugs, each masking the next, are fixed: (1) a trailing; ...comment on theMinVersiondirective was parsed as part of the value (Inno Setup only treats;as a comment at the start of a line); (2) noSourceDirwas set, so Inno resolved every repo-root-relativeSource/SetupIconFile/OutputDiragainst the script's owninstaller\directory — nowSourceDir=.., which also lands the output at<root>\outwhere the build script and CI look for it; (3) an{app}Inno constant inside a Pascal{ }doc-comment in[Code]closed the comment early (these comments don't nest). Also shipsscripts\assets\*(the tray.icothe AHK loads at runtime), which[Files]had omitted.
Previous public release.
- Browser-based dashboard served by the local daemon at
127.0.0.1:52650(Overview, History, Telemetry, Notes, Benchmark, Config) with an auto/light/dark theme toggle. - Custom prefix modes: define your own
prefix:command (for exampletranslate:) in Dashboard → Config → Custom modes. New modes apply to the running app within a second; built-in mode prompts stay locked. - Notes-aware chat: a "My notes" toggle in each chat tab grounds replies in your notes vault and cites note titles.
- History and notes are browsable from the web dashboard (new daemon actions
recent_history,notes_list,mode_ids).
- The native AutoHotkey dashboard was removed; the tray "Dashboard" item now opens the web dashboard.
- Mode-prefix parsing hardened: invisible Unicode (NBSP/ZWSP/BOM), CR-only line endings, nested quote markers, and multi-line bodies are handled correctly.
prompt:mode no longer returns the input verbatim, and multi-line prompt bodies are no longer truncated to the first line.- Toast notifications no longer race their temporary script file; model pulls that hang silently are now killed by a watchdog.
- The daemon rejects requests with a foreign
Hostheader (DNS-rebinding defense) and serves static dashboard files only from an explicit route allowlist. - URL fetching during note capture refuses loopback, private, and link-local addresses (SSRF guard).
- Dashboard benchmark tab for running FastFlowLM model benchmarks.
- Notes setup tab for vault location, categories, and LLM note behavior.
- FastFlowLM runtime status and update check in the dashboard.
- Local note search support for chat context.
- App files now live at the repository root instead of a nested packaging folder.
- Public README now includes screenshots and first-run setup instructions.
- Default privacy mode keeps selected text redacted from history unless enabled.
- Note capture supports inbox fallback safely.
- Ask-in-chat launches without blocking the daemon.
- Multiline mode prefixes such as
prompt:parse correctly. - Production config/data/log paths resolve consistently between Python and AutoHotkey.
- Local daemon POST actions require the
X-FFP-APIheader. - Config patching is restricted to approved keys and local patch files.
- Update ZIP extraction validates paths before unpacking.