|
| 1 | +# On-device LLM / ML — design decision |
| 2 | + |
| 3 | +**Status:** Decided — SkipperKit stays heuristic-only. Reaffirmed 2026-07 after a |
| 4 | +Gemma reassessment. |
| 5 | + |
| 6 | +SkipperKit matches streaming-app controls (Skip Intro / Skip Recap / Next Episode) |
| 7 | +by reading the **accessibility node tree** — text, content-description, view-id, |
| 8 | +bounds, and clickability that the OS already publishes. It uses **no** screen |
| 9 | +capture, OCR, image recognition, or ML model. This doc records why an on-device |
| 10 | +LLM (specifically the Gemma family) is not used, so the question isn't relitigated |
| 11 | +without new evidence. |
| 12 | + |
| 13 | +## Where an LLM could even fit |
| 14 | + |
| 15 | +Never the hot path. `onAccessibilityEvent` fires continuously during playback; |
| 16 | +running inference there would drain battery and risk a mis-tap. The only plausible |
| 17 | +place is the **discovery tier** — an off-hot-path, throttled, propose-only step that |
| 18 | +classifies an unrecognized button (skip-intro / skip-recap / next-episode end-card / |
| 19 | +control-bar / dismiss / none) and offers it for the user to approve. That tier never |
| 20 | +taps on its own. |
| 21 | + |
| 22 | +## Why not Gemma (or any general on-device LLM) |
| 23 | + |
| 24 | +Reassessed 2026-07 against current Gemma releases. Three independent reasons, plus a |
| 25 | +structural blocker: |
| 26 | + |
| 27 | +### 1. Size and latency are grossly disproportionate |
| 28 | + |
| 29 | +| Model | Disk (int4, LiteRT) | Peak RAM | Note | |
| 30 | +|---|---|---|---| |
| 31 | +| Gemma 3n E2B | 3.0–3.7 GB | 2.7–3.4 GB | "effective 2B" hides ~5B real params; never GA on LiteRT | |
| 32 | +| Gemma 3n E4B | 4.3–4.9 GB | ~4–6 GB | needs 12–16 GB device RAM to be stable | |
| 33 | +| Gemma 3 1B | ~530 MB | ~1 GB (4 GB floor) | smallest "real" LLM | |
| 34 | +| Gemma / FunctionGemma 270M | ~125–250 MB | ~256 MB | smallest; non-deterministic | |
| 35 | +| Gemma 4 E2B (2026-04) | 2.0–2.6 GB | ~1.7 GB | Apache-2.0/ungated, still oversized | |
| 36 | + |
| 37 | +For comparison, the task-appropriate alternative (below) is **~28 MB / ~53 ms**, or a |
| 38 | +word-embedding classifier **<1 MB / ~0.03 ms**. Even the smallest Gemma is ~25× larger |
| 39 | +and orders of magnitude slower, in a long-running service that must coexist with a |
| 40 | +foreground streaming app on real devices. |
| 41 | + |
| 42 | +Runtime is LiteRT-LM (MediaPipe `tasks-genai` is deprecated). Models are too large to |
| 43 | +bundle in an APK, so they'd be a multi-GB **runtime download** behind HuggingFace |
| 44 | +license gating, with engine init up to ~10 s and 30–120 s NPU JIT on first run. |
| 45 | + |
| 46 | +### 2. It's the wrong tool for a button-tapper |
| 47 | + |
| 48 | +The task is short-text classification into a small fixed label set. An LLM is |
| 49 | +generative and **non-deterministic** — a hallucinated positive means a mis-tap. |
| 50 | +A deterministic classifier is smaller, faster, and safer for something that |
| 51 | +performs `ACTION_CLICK` on the user's behalf. |
| 52 | + |
| 53 | +### 3. The "official" on-device path is structurally blocked |
| 54 | + |
| 55 | +ML Kit GenAI / AICore (Gemini Nano) enforces `BACKGROUND_USE_BLOCKED`: on-device |
| 56 | +inference runs only while the calling app is the **top foreground** app (a Private |
| 57 | +Compute Core constraint). An accessibility service is never the foreground app, so |
| 58 | +this path is structurally incompatible with SkipperKit — independent of the |
| 59 | +allowlist gating that already blocks sideloaded apps. |
| 60 | + |
| 61 | +## If on-device ML is ever justified |
| 62 | + |
| 63 | +It isn't today. If field data ever shows the heuristic misses a meaningful fraction |
| 64 | +of real controls, use a **fine-tuned TFLite text classifier bundled in the APK** |
| 65 | +(e.g. MobileBERT NLClassifier, ~28 MB, deterministic, sub-53 ms) — not an LLM — and |
| 66 | +keep it behind the existing discovery flow (propose → user approves → never |
| 67 | +auto-tap) so a wrong prediction can never cause a tap. |
| 68 | + |
| 69 | +Only revisit a general LLM if all of these hold: (a) field data shows heuristic + |
| 70 | +tiny-TFLite still miss a significant fraction of nodes, (b) a zero-shot multi-locale |
| 71 | +need can't be met by retraining the classifier, and (c) output is |
| 72 | +confidence-thresholded behind the human-approval gate. |
| 73 | + |
| 74 | +## Cheaper win first |
| 75 | + |
| 76 | +Extend the heuristic with locale-aware label patterns as new services arrive through |
| 77 | +the contribution pipeline. No model, no download, fully deterministic. |
0 commit comments