Autonomous code delivery between human gates.
Polako works a GitHub issue backlog to zero. It takes the lowest open issue, hands it to Claude Code, waits for you to merge the pull request, and moves on to the next one. It runs unattended, one issue at a time. It never merges anything itself. The Github issues and pull requests are reviewed by humans.
Polako is Croatian for "take it easy" or "slow and steady", which is the philosophy we follow here. We have engineered for correctness over speed. Polako takes its time so reviewers don't waste theirs.
Two halves, one release: the /implement-issue skill takes one issue
from research to a plan to a pull request, on its own or driven by the
polako binary, which supervises the queue — runs the skill, watches the
PR, repairs it when CI goes red, advances when you merge. Stdlib-only Go, no
dependencies.
Ten verbs. Four start Claude runs:
work— works the backlog to zero, one issue at a time, never merging.plan— turns a design document into a curated backlog of proposals.health— reads the repository itself and files what looks off as proposals.design— works one design request into the design documentplanreads, behind a PR you merge.
plan and health file everything behind a proposed label a human has to
lift; design writes the document first — see
Planning a backlog.
Six look after the shift, and run no model:
status— where the backlog stands and what is waiting on you, read from GitHub, plus the last shift run here.stats— what your runs cost and how they went, read from local run data.tidy— reclaims the worktrees and branches of finished issues. It only previews until you pass-apply.unpark— lists parked issues with why each stopped, and clears the ones you approve.update— brings the plugin and the binary to the published release.setup— reports whether a repository is ready for polako, and fixes what it can with-apply.
A bare polako prints this table; polako <verb> -h prints that verb's flags.
lowest open issue with no sub-issues and no `needs-human`, `proposed`
or `awaiting-answer` label
↓
claude -p "/implement-issue N" ← headless; milestones on your terminal,
↓ the full stream in a per-shift log
↓
PR opened? ──no──► issue labelled `awaiting-answer`? ──yes──► put it down, advance to the next
│ │ (re-run it when the reply lands)
│ └──no──► crashed, or left work on the branch?
│ │ resume the same session
│ └──out of attempts──► park it, advance to the next
↓ yes
wait for merge (-poll) ← rebases if GitHub reports CONFLICTING,
↓ fixes + re-pushes if the checks go red
↓ or a reviewer requests changes
close the issue, remove the worktree, advance to the next
That is work's loop. design is the same loop on one issue you name,
ending at a merged plan document instead of code. plan and health are
simpler: one claude run, filing proposals, then done — no PR, no polling.
You have two jobs, both on GitHub: answer a question when a run asks one on an issue thread, and merge the pull requests. Neither is on a clock, and nothing else needs you — see The rules it follows and docs/behaviour.md for the long version.
polako checks for all three at startup, rather than failing an hour into an unattended run.
The skill installs as a Claude Code plugin. This repository is its own marketplace, so there is nothing to clone:
claude plugin marketplace add scharissis/polako
claude plugin install polako@scharissisRestart Claude Code and /polako:implement-issue 48 is available. Claude
prefixes plugin skills with the plugin name, so the command is not
/implement-issue on this path.
Then the binary:
go install github.com/scharissis/polako/cmd/polako@latestPrebuilt binaries for Linux, macOS and Windows are attached to every release, which is easier on a machine without Go.
Both halves come from the same release and are meant to move together. Installing the skill by hand, updating, pinning a version and setting a project up for your team are all in docs/install.md.
polako updateRun it between shifts, never during one — see
docs/install.md#update for -check, pinning, and
the by-hand commands.
Check the repository is ready first — what it has for polako and what's missing, read-only:
polako setup -dir ../my-projectSee docs/setup.md for what it checks and -apply.
Look first — -dry-run resolves the next issue and prints the command it
would run, nothing else:
polako work -dir ../my-project -dry-runWork one issue and stop:
polako work -dir ../my-project -onceTrust it, and let it work the whole backlog, telling you when it needs you:
polako work -notify ~/bin/tell-meAsk where things stand, from any machine, including about a shift running somewhere else:
polako status -repo scharissis/polakoscharissis/polako
ready 3 issues — #14, #19, #23
held back 1 issue — #33 (behind #30)
awaiting you 1 issue — #9 (quiet 26h)
parked 1 issue — #5, labelled needs-human
proposed 2 issues — #27, #28, labelled proposed
containers 1 issue — #12 (2/5 closed)
next #14 — its branch already has PR #61, so it would wait on that rather than run the skill again
open prs on issue branches
pr branch issue mergeable checks review url
#61 issue-14 #14 mergeable passing clear https://github.com/scharissis/polako/pull/61
needs you: reply on #9; review and merge PR #61; decide what to do about #5 (drop needs-human to requeue); curate #27, #28 (drop proposed to queue them)
On a repository with plans under docs/designs/, a plan documents table sits
above the needs you: line: one row per plan, and how far along its issues
are. A shift ends the same way — merged, parked and why, dollars spent — see
docs/behaviour.md for a worked example.
See what the runs cost, from the records every run leaves on your machine:
polako statsA shift cleans up after itself. After a killed shift or a run by hand, preview which finished worktrees and branches are safe to reclaim, then do it:
polako tidy
polako tidy -applytidy names whatever it refuses to touch and why — see
docs/reference.md.
Somebody still has to write the issues polako works. plan runs
/plan-backlog: point it at a vision or roadmap document and it decomposes
the gap into issues sized to one PR each, groups anything cross-cutting under
an epic, and files the lot as proposals:
/polako:plan-backlog docs/VISION.md
health runs /review-health the same way, but reads the codebase itself
instead of a document — file and function sizes, duplicated helpers,
abstractions nothing uses — and proposes the outliers. Where the repo has
prompts (skills, CLAUDE.md, prompt code), it also runs Claude Code's own
prompt audit over them and proposes fixes for text written for older models.
Point it at any repository; it is not polako-specific:
/polako:review-health .
Both are documented alongside work's own flags: plan,
health.
plan decomposes a document; it does not write one. When the idea is still
a paragraph, design runs /design-plan on it: it asks its questions on the
issue thread, then opens a PR adding one document under docs/designs/, and
the PR review is where the design gets argued. Name an open issue, or hand it
the text and it files the issue for you:
polako design -issue 31
polako design -brief "a dating app for horses"
The hand-off is the merge. Once you merge the PR, polako status lists the
new document as draft, and polako plan -design docs/designs/<doc>.md
turns it into proposals — printed at the end of the run, never run for you.
work skips anything labelled design, so a request never gets coded
before it's designed. Flags in
design.
Every proposal carries a proposed label, and that label is the point:
polako work skips every issue that has one, so nothing a machine proposed
can reach an unattended run until you have looked at it. Each proposal
carries acceptance criteria, pointers into the code, what's out of scope, and
a size (Estimate: M) — the model's judgement of the work's shape, not a
price; what a run actually costs comes from your own history, via polako stats. Curation is ordinary GitHub triage, and there are three moves:
- Approve — remove the
proposedlabel. On a-label-gated repository, add the gate label in the same command:gh issue edit 27 28 --remove-label proposed --add-label ready - Reject — close the issue.
- Rework — edit the text. A run reads the issue when it picks it up, so your edits are the spec; there is no further step.
polako status lists what is waiting on you, proposals included, so a forgotten
batch surfaces rather than rots.
- One issue at a time. Never two. That is what makes the runs unable to conflict.
- Nothing merges itself. polako opens, updates and repairs pull requests. It never merges one, and it never commits to your default branch.
- All the state is in GitHub — issues, labels, comments, branches, PRs. Kill polako at any point and start it again later. It works out where things stand by asking GitHub, not by reading anything it saved.
- An issue it cannot finish is parked, not retried forever. It gets a
needs-humanlabel and a comment saying what happened, and the shift carries on with the rest of the backlog. The reason is recorded as one identifier too, sopolako statscan rank what parks issues most — see docs/run-data.md. - Your checkout is never written to. polako fast-forwards your default branch so a review has the right base, and refuses rather than rebase, reset or commit.
- Issue text is data, not instructions. On a repo that takes issues from outside your team, that text is written by strangers, and the skill is told to read it as a description of a change rather than as orders.
In the Balkans, "polako!" is what you say to someone who's rushing. Slow down, you'll get there — a good way to ship code, too: one pull request at a time is one you'll actually read. "Wouldn't ten agents be faster?" They'd open ten pull requests that all branched from a version of the code that stopped being true the moment the first one merged; polako does one, and the next starts from what you just merged.
Every run records what it spent, so these are measured rather than estimated — from one 33-hour shift on a small Go project, eight issues finished:
| Issues that merged | 7 of 8 |
| Cost per issue | $10.38 mean, $11.31 median |
| Runs per issue | 1.4 mean |
| Size of the change | +505 / −40 across 5 files, median |
| Time from PR opened to merged | 10m median, because someone was watching |
Your numbers will differ, and the ones that move them most are how big your
issues are and how often runs crash — a crashed run's resume pays to read the
context again. Run polako stats for your own figures. -max-issue-time
already defaults to a ceiling; set -max-cost or -max-session-cost for one
too.
docs/run-data.md has the whole report.
Every run records what it did, and those records are only worth keeping if something reads them, on a cadence — measure, review, change one thing, tag the next batch. That loop runs through you and the backlog, never through the supervisor reading its own telemetry. docs/continuous-improvement.md has the retro checklist, the tagging rule, and the recipes for reading run data back.
- It will not merge for you, and there is no flag that changes that.
- It is not a sandbox. The tool allowlist narrows what a run can do, but
build commands run whatever your repository's scripts contain. Point
-dirat repositories you would runmake testin yourself. - It is not finished. This is pre-1.0. Flags and defaults still change, and the release notes say when.
- It cannot tell a good issue from a bad one. A vague issue produces either a question on the thread or a park, and both cost money to find out.
- The skills' eval suite has not had a green run yet. Each skill change runs the eval cases it touches, but no full pass has come back green. See evals/README.md.
polako work takes around two dozen flags, and the other nine verbs have
their own smaller sets. docs/reference.md has work,
plan, health, design, status, tidy and unpark, together with -dry-run, -notify,
-remote and the POLAKO_* environment defaults; stats is in
docs/run-data.md, setup in
docs/setup.md, and update in
docs/install.md, each beside what it describes.
Any flag can take its default from the environment, so a preference you always
want can live in your shell profile.
An unattended run is a Claude session whose only input is issue and comment
text. On a repository that accepts issues from outside your team, anyone writes
that input. Two things bound it. The tool allowlist is enforced by Claude Code
rather than by the model behaving well, and -label means a maintainer has to
opt each issue in before polako will touch it — required outright on a public
repository, unless you pass -ungated and mean it.
Nothing you run leaves your machine unless you ask. One exception is named
outright: -post-summary, off by default, comments one line of numbers on your
own merged PR. -remote is on by default and would be the second, but no
claude CLI registers headless runs with Remote Control yet, so polako does not
pass the flag and no session content goes anywhere.
docs/security.md has the reasoning and the limits,
docs/hardening.md covers running a shift behind an egress
firewall you supply, and SECURITY.md says how to report a
vulnerability privately.
The common ones — using the skill without the binary, running it on your language, comparisons to a hosted coding agent — are answered in docs/behaviour.md, which also covers what happens when it breaks something at 3am.
| Page | What is in it |
|---|---|
| docs/behaviour.md | What polako does when a run crashes, an issue stalls, a PR goes red, or it needs a human. FAQ at the bottom. |
| docs/install.md | Every install path, polako update and its flags, pinning, and using it on another project. |
| docs/reference.md | Every flag for work, plan, health, status and tidy, plus -dry-run, -notify, -remote and environment defaults. |
| docs/setup.md | polako setup, the report on whether a repository is ready, and -apply to fix what it's missing. |
| docs/run-data.md | What each run records (work, plan and health alike), spending caps, and the polako stats report. |
| docs/security.md | The threat model, the tool allowlist, the -label gate, and what leaves your machine. |
| docs/hardening.md | Wrapping a shift in an egress firewall of your own, and why polako does not ship one. |
| docs/releasing.md | Cutting a release: the two tags, the two PRs, and what to bump. |
| CONTRIBUTING.md | Running the tests, the eval suite, and both halves from a working tree. |
MIT — see LICENSE.
