Dewey draws a map of one GitHub user's starred repositories. Each star becomes a point. Repositories that do similar things sit close together in named clusters, so a few thousand stars become something you can browse.
For every starred repository it fetches the metadata and README, has an LLM write a one-paragraph technical summary, embeds the summaries, reduces them with UMAP, clusters them with HDBSCAN, has the LLM name each cluster from its most central repositories and distinctive words, and writes an interactive HTML map. Hover a point for the first lines of its summary, click it to open the repository.
The words used here (star, summary, cluster, unclustered, central repository) are defined in GLOSSARY.md.
Requires mise and a GitHub token with public read access.
mise install # python, uv, ollama at the pinned versions
uv sync # the virtualenv and every dependencyTokens go in .env:
echo "GITHUB_TOKEN=$(gh auth token)" >> .env
echo "ANTHROPIC_API_KEY=sk-ant-..." >> .envmise loads .env into the shell when its activation is installed. Without it, pass the file to uv on each run: uv run --env-file .env dewey ....
uv run dewey joaofndsThis writes map.html. The first run for a user lists their stars, fetches each repository not yet under data/repos/, and writes one summary per new repository with Claude (claude-sonnet-5 by default). The star list is kept in data/starred_ids.txt, and later runs reuse it and everything else under data/. Pass --refresh-stars to list the stars again and --overwrite-summaries to rewrite every summary.
To run without an API key, use a local model through Ollama (gemma3:4b is the default, --llm-model picks another):
ollama serve
ollama pull gemma3:4b
uv run dewey joaofnds --llm ollamaThe knobs that change the map most:
| Option | Default | Effect |
|---|---|---|
--dimensions {2,3} |
2 | 2-D maps carry cluster names on the map; 3-D maps rotate |
--min-cluster-size N |
15 | smallest group HDBSCAN will call a cluster |
--min-samples N |
5 | higher values leave more repositories unclustered, but clusters get tighter |
--embedding-model NAME |
Qwen/Qwen3-Embedding-0.6B |
any sentence-transformers model |
--seed N |
42 | UMAP's random state; the same seed and data give the same map |
uv run dewey --help lists the rest.
- Stars. The user's starred repositories come from the GitHub API. Each is stored once under
data/repos/<id>/as the repository object and the README contents object, exactly as GitHub returned them, and the star list is cached indata/starred_ids.txt. - Summaries. The LLM gets the repository's name, description, language, license, size, year, topics and the first 4,000 characters of its README, and returns a dense paragraph written for clustering (the prompt is
src/dewey/prompts/summary.md). Summaries are stored beside the repository and never rewritten unless asked. Summaries rather than raw READMEs go into the embedding because READMEs vary from a badge wall to a book, and the summary normalizes them to the same register and length. - Embeddings. Summaries are embedded with a sentence-transformers model and L2-normalized. Vectors are cached in
data/embeddings/keyed by model name and text, so changing either re-embeds. - Two reductions. UMAP reduces the vectors twice: to five dimensions with
min_dist=0for clustering, and to two or three dimensions for the map. Clustering on the plotted coordinates, which is what the first version did, throws away structure that the plot does not need but the clusterer does. - Clusters. scikit-learn's HDBSCAN clusters the five-dimensional points and gives each repository a membership probability. Repositories it cannot place are unclustered and drawn in grey.
- Names. For each cluster the LLM sees the repositories with the highest membership probability and the cluster's top TF-IDF terms, computed with one document per cluster, and answers with a two-to-four-word category (the prompt is
src/dewey/prompts/cluster_name.md). Names are cached indata/names/keyed by model, prompt, clustering and texts. - Map. Plotly draws one trace per cluster and, on 2-D maps, each cluster's name at its median position.
Measured on this user's 2,872 summaries on an Apple M5 Pro on 2026-09-21 with uv run dewey-measure sentence-transformers/all-MiniLM-L6-v2 Qwen/Qwen3-Embedding-0.6B, which runs the pipeline's own reduction and clustering at the CLI defaults and prints this table:
| Embedding model | Embed time | Clusters | Unclustered | Silhouette (5-D) |
|---|---|---|---|---|
all-MiniLM-L6-v2 (the first version's model) |
10 s | 49 | 18.6% | 0.491 |
Qwen/Qwen3-Embedding-0.6B |
60 s | 57 | 21.4% | 0.571 |
Qwen3 separated clusters better at every setting tried, at six times the embedding cost, and became the default. Both models run locally. The first run downloads the weights (about 1.2 GB for Qwen3).
The LLM writes the summaries and the cluster names, so pick it by cost and writing quality. The 2,872 summary prompts total 14.8 million characters, about four million input tokens, so regenerating every summary with claude-sonnet-5 costs on the order of ten dollars. The Message Batches API halves that price but is not wired in.
data/repos/ is committed because it holds the fetched repositories and their summaries, and the summaries are the expensive part. data/embeddings/, data/names/ and data/starred_ids.txt are caches and are ignored by git. Delete a file in one of those three places to force that stage to run again.
mise run check # ruff format, ruff lint, pyright strict, pytest
mise run fix # format and auto-fixThe test suite runs the whole pipeline through in-memory fakes of GitHub, the LLM and the embedder, so it needs no network and no model weights. Libraries that ship no type information (umap, scikit-learn, plotly, sentence-transformers) have minimal stubs under typings/ declaring only what this project calls.
GITHUB_TOKEN is not set: put it in.envor export it.gh auth tokenprints one if you use the GitHub CLI.- Ollama connection errors:
ollama servemust be running and the model pulled (ollama list). - A gated embedding model (for example
google/embeddinggemma-300m) needs a Hugging Face login that has accepted its terms; Qwen3-Embedding is Apache-2.0 and needs none. - HDBSCAN
cluster_selection_epsilonis not exposed because scikit-learn 1.9.1 raises a NumPy scalar-conversion error on that path whenever it would merge clusters (scikit-learn#34243).