Skip to content

resolveCapabilities re-reads + Ed25519-verifies the entire skill corpus on every call (no cache) — ~27s/call at 78 skills #102

Description

@jack-arturo

Summary

resolveCapabilities() re-reads and Ed25519-verifies the entire installed skill corpus from disk on every call, with no in-memory cache. The per-call cost scales with the number of installed skills, so for any consumer that calls it on a hot path it becomes a dominant latency cost.

Observed: ~25–30s per call with 78 installed skills (~3.5 MB) in ~/.autovault/skills.

Root cause

v0.2.0, dist/capabilities/:

  1. resolveCapabilities(input, db = openCapabilityDb()) (resolver.js) — when the caller doesn't pass a db, it opens the capability DB on every call.
  2. resolvedSkills(db, groups) (resolver.js) collects names via SELECT name FROM skills (i.e. all installed skills, not just matched groups), then for each one does await readSkill(name).
  3. readSkill() (storage/index.js) reads SKILL.md + the detached signature and runs verifySignatureIfPresent() → an Ed25519 verification per skill, per call.

So a single resolveCapabilities call = O(N skills) file reads + O(N) signature verifications + parsing, every time. There's no caching of the parsed/verified corpus across calls.

This is reasonable for AutoVault's intended use (a stdio MCP server an editor spawns and calls occasionally). It becomes pathological when a host calls resolveCapabilities frequently — e.g. once per assistant turn.

Impact / repro

  • Install ~75+ skills.
  • Call resolveCapabilities({ caller_id, platform, query }) repeatedly (no db arg) and time each call.
  • Each call takes seconds-to-tens-of-seconds and re-reads/re-verifies the whole corpus; back-to-back calls don't get faster.

In our case (AutoHub) this surfaced as ~27s time-to-first-token on voice turns, because we call capability/skill resolution per turn, per caller.

Suggested fixes (any of these would help)

  1. Cache the verified skill corpus in-memory, keyed by storage path, invalidated by a cheap signal (directory mtime, or a version counter bumped on install/update/remove). resolvedSkills reads from the cache instead of re-reading + re-verifying every call.
  2. Don't re-verify signatures on the read/resolve path. Signature verification is most valuable on write/install (untrusted bytes). Re-verifying the entire corpus on every capability resolution is the bulk of the cost; consider making read-path verification opt-in/lazy.
  3. Reuse a single DB handle rather than openCapabilityDb() per call (smaller win — the dominant cost is the per-skill file I/O + crypto, not the DB open).

Workaround (downstream)

We added a per-caller TTL cache around our resolveCapabilities wrapper in AutoHub, which fixed it for us — but an upstream cache (or skipping read-path verification) would fix it for every consumer and avoid each host re-implementing the same mitigation.

Environment: @autoworks/autovault@0.2.0, Node 24, macOS.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions