SitemapKit is a command-line XML sitemap extractor and GitHub Action. It parses URL sitemaps, follows nested sitemap indexes, decompresses .xml.gz files, and prints JSON or one URL per line. The extract command runs locally without an account. Domain discovery and the combined full workflow use the SitemapKit extraction API.
SitemapKit CLI works in shell scripts, CI pipelines, and GitHub Actions for URL inventories and content workflows.
- Node.js 18 or newer
- A SitemapKit API key for
discover,full, and the GitHub Action
Extract one known sitemap or sitemap index. This follows nested indexes up to five levels, decompresses .xml.gz responses, and needs no API key:
npx github:0nl1n1n/sitemapkit-cli extract https://example.com/sitemap.xml --max-urls 5000To discover sitemap files from a domain, create a key at sitemapkit.com/register. The free API plan includes 20 processing credits per month. Extraction uses one credit per 1,000 returned URLs, rounded up.
export SITEMAPKIT_API_KEY=sk_live_...
npx github:0nl1n1n/sitemapkit-cli discover https://example.com
npx github:0nl1n1n/sitemapkit-cli full https://example.comEach command writes JSON to standard output. Errors go to standard error and return a non-zero exit code, so the CLI works in shell pipelines and CI jobs.
Use --format urls with extract or full when a pipeline needs one URL per line instead of the full JSON response:
npx github:0nl1n1n/sitemapkit-cli extract https://example.com/sitemap.xml --format urls > urls.txtYou can test the same workflow before adding an API key. If you prefer to check by hand, follow the seven ways to find a website sitemap.
- Find sitemap files for a domain
- Count and extract URLs from a sitemap
- Validate sitemap XML and
lastmoddates
The finder starts from a domain. The extractor and checker accept a sitemap URL or pasted XML.
To prepare or compare files locally:
- Generate sitemap.xml from a URL list or CSV
- Split a sitemap at the 50,000 URL and 50 MB limits
- Compare two XML sitemaps
These tools run in the browser and do not require an API key.
The CLI is intended for on-demand and CI runs. To keep a durable URL baseline, detect new pages automatically, and ping a signed webhook, use SitemapKit Monitoring.
The free plan includes one daily monitor. Paid plans add more websites, larger sitemaps, and checks as often as every hour. Webhook payloads and HMAC verification are documented in the monitoring webhook guide.
Add the API key as a repository secret named SITEMAPKIT_API_KEY, then run the extractor in a workflow:
- name: Extract sitemap URLs
id: sitemap
uses: 0nl1n1n/sitemapkit-cli@v1
with:
api-key: ${{ secrets.SITEMAPKIT_API_KEY }}
command: extract
url: https://example.com/sitemap.xml
max-urls: 5000
- name: Upload the URL inventory
uses: actions/upload-artifact@v4
with:
name: sitemap-urls
path: ${{ steps.sitemap.outputs.result-file }}The action writes the full API response to sitemapkit-result.json by default. Its total-urls output contains the extracted URL count. Use discover to find sitemap files without extracting them, or full to discover and extract in one step.
git clone https://github.com/0nl1n1n/sitemapkit-cli.git
cd sitemapkit-cli
npm link
sitemapkit extract https://example.com/sitemap.xmlsitemapkit discover <domain-url>
sitemapkit extract <sitemap-url> [--max-urls <1-50000>] [--format <json|urls>]
sitemapkit full <domain-url> [--max-urls <1-50000>] [--format <json|urls>]
Set SITEMAPKIT_API_BASE_URL only when testing against another compatible API origin. It defaults to https://api.sitemapkit.com.
When SITEMAPKIT_API_KEY is set, extract uses the API as well. That route adds managed fetching for bot-protected sites. Leave the key unset to parse the sitemap locally.
| Plan | Processing credits/month | URLs per extraction |
|---|---|---|
| Free | 20 | 1,000 |
| Starter | 1,000 | 10,000 |
| Pro | 5,000 | 50,000 |
| Agency | 20,000 | 50,000 |
See current API and monitoring allowances on the pricing page.
MIT