Skip to content

add sitemap crawling to help maximise URLs found #299

Description

@a11ya11y

this is currently happening with 2 sites: Tax Policy Home and Tax Technical - Inland Revenue NZ

They each have 1000s of pages, but Tax Policy Home is only getting 236 pages scanned, and Tax Technical - Inland Revenue NZ even fewer. I get the same results on a 2 separate installs of CWAC on different infrastructure.

Running LibreCrawl over Tax Policy and it got a 200 response from 1000 URLs, so the site can be crawled.

Here's the output from a 300-page scan of these 2 sites, where taxpolicy.ird.govt.nz only had 236 pages scanned, and taxtechnical.ird.govt.nz had 109:

2026-07-22_11-19-35_taxpolicy-taxtecnical.log
audit_log.csv
chromedriver.log
config.json
pages_scanned.csv
progress.csv

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions