The DevOps Error Library is built on one durable foundation: a tree of plain Markdown files under errors/, compiled by errlib into a JSON search index. Every item on this roadmap reads that same corpus and index β no rewrites, no lock-in. We grow the content, and tooling layers on top.
We're at ~167 documented errors across 14 technologies today, architected to scale to 10,000+. Here's how we get there.
Want to help? See
CONTRIBUTING.md. The single highest-leverage thing you can do is add accurate error pages.
The priority is breadth and depth of high-quality, verified errors, plus the tooling that keeps the corpus trustworthy.
Stretch goals for the most-requested technologies (current counts in errors/README.md):
| Technology | Now | Target |
|---|---|---|
| Kubernetes | 16 | 200 |
| Docker | 12 | 150 |
| Terraform | 16 | 200 |
| OpenStack (all services) | 21 | 250 |
| Linux | 12 | 150 |
| GitLab | 12 | 100 |
| Prometheus / Grafana | 22 | 150 |
| RabbitMQ / Redis | 20 | 120 |
| PostgreSQL / MySQL | 20 | 150 |
| Ceph / Linstor | 16 | 120 |
These are directional, not gates β accuracy still beats quantity. Batches are coordinated via tracking issues (see CONTRIBUTING.md).
- Category indexes β keep
scripts/gen_indexes.pyoutput fresh on every merge (per-categoryREADME.md+ the root index). - Link checking β CI step that verifies every relative Related Errors link and external References URL resolves.
- Freshness tracking β surface pages with a stale
last_revieweddate so they get re-verified. - Front-matter & schema tightening β broaden
errlib validatechecks (tag hygiene,relatedslugs that actually exist, severity sanity). - New technologies β add the next wave of folders as contributors bring verified errors (CI/CD runners, Helm, ArgoCD, Nginx, HAProxy, etcd, cloud-provider APIs, β¦).
Once the index is rich, expose it through more surfaces β all reading errlib's JSON index.
- CLI search tool packaging β publish
errlibto PyPI (pip install devops-error-library) and ship prebuilt binaries so the local fuzzy search is a one-line install. (The CLI itself already exists viaerrlib search.) - Static docs site β generate a browsable, SEO-friendly site directly from the Markdown corpus (one URL per error), so pages rank for "Kubernetes / Terraform / Docker / OpenStack troubleshooting" searches.
- Web search UI β a hosted search front-end over the JSON index: paste a log line, get matching pages, filter by technology/severity/tag.
- REST API β a thin HTTP service over the JSON index (
/search,/errors/<slug>,/stats) so other tools and dashboards can query the library programmatically. - GitHub Action β drop-in Action that, on a failed CI job, searches the library for the failure's error string and comments the matching error pages on the run β troubleshooting where the failure happens.
With a large, structured, verified corpus and an API, build the smart layer on top.
- AI troubleshooting suggestions β grounded, retrieval-based suggestions that cite specific library pages (never ungrounded guesses), bringing the structured corpus to the AI Incident Response Assistant.
- MCP server β a Model Context Protocol server so AI agents and IDE assistants can query the library as a first-class tool (
search_errors,get_error), with citations back to the corpus. - VS Code extension β search and read errors without leaving the editor; surface matching pages from terminal output and log files.
- Community voting β let readers mark whether a page resolved their issue, feeding a quality signal back to maintainers.
- Error popularity rankings β aggregate search and voting data into "most-hit errors per technology" and trending pages, guiding where to deepen coverage next.
- Markdown is the source of truth. Every feature reads
errors/**/*.mdand the generated JSON index. If a feature needs a separate database of content, we've taken a wrong turn. - Accuracy over quantity. Growth never lowers the bar β verified logs, read-only diagnostics, real fixes.
- Offline-first. Core value (the corpus +
errlib search) always works with no network and no account. - Open by default. Content stays CC BY 4.0; tooling stays MIT.
Have an idea or want to own a roadmap item? Open an issue or read CONTRIBUTING.md. π