Skip to content

feat(seo): collect per-URL index state and per-family search performance from the Search Console API (weekly) #8606

Description

@koala73

Parent: #8602 · Related: #6036, which runs the manual measurement cycle. This issue automates its Google half.

Outcome

A command that answers two questions for each URL family every week, from the Search Console API rather than UI transcription:

  1. Which of our URLs are indexed, and which did Google decline and why?
  2. Which families earn impressions and clicks per indexed page?

Until this exists, every decision about growing or pruning a page family is a guess. The 2026-09-24 coverage export has counts only. We know 1,486 URLs are "Crawled, currently not indexed" but not which ones.

Human prerequisites (owner only; an agent cannot do these)

  1. Create a Google Cloud service account with the Search Console API enabled. Add its email as a Restricted user on the worldmonitor.app Domain property.
  2. Store its JSON key where loadEnvFile() finds it, under one variable such as GSC_SERVICE_ACCOUNT_JSON (base64). Add the same value as a GitHub Actions secret if the weekly job should run in CI.
  3. Comment the property ID's format here (sc-domain:worldmonitor.app or a URL prefix) without secrets.

The agent can finish everything below against recorded fixtures before step 1 happens. Only the live-run criterion waits on the owner.

Scope (agent)

  1. Collector. Add a collector under scripts/, for example scripts/seo-gsc-collect.mjs, that:
    • calls searchanalytics.query with dimensions page (and query for page-level top queries), for the 28-day and 90-day windows, paginating past 25k rows;
    • calls urlInspection.index.inspect on a stratified sample of up to 2,000 URLs a day (the API quota). Draw the sample from the three live sitemaps, with every family represented. Record coverageState, indexingState, googleCanonical against userCanonical, lastCrawlTime and pageFetchState;
    • loads credentials only through loadEnvFile() and never prints secrets;
    • keeps missing values as null with a reason, never zero. This matches the contract in scripts/seo-ai-visibility-collector.mjs.
  2. Family mapping. Reuse PAGE_FAMILIES in scripts/seo-ai-visibility-scorecard.mjs:55-76. It does not yet cover sources, compare, research, use_cases, docs_zh, markdown_twins (*.md, llms*.txt) or story_shares. Extend the domain and do not fork it. Every sitemap URL must map to exactly one family, and a test enforces this.
  3. Report. Write a dated snapshot next to the existing baselines in docs/research/seo-ai-visibility/ (for example gsc/YYYY-MM-DD.json), following their conventions. Also write a markdown summary. Per family it shows URLs declared, sampled, indexed %, top non-indexed reasons, impressions, clicks, CTR, impressions per indexed URL, and the 5 best and 5 worst URLs.
    • Feed the per-family numbers into the existing scorecard instead of starting a parallel report.
  4. Schedule. Add a weekly GitHub Actions workflow that runs the collector and opens or updates a PR with the new snapshot. Gate it on the secret, so it skips cleanly when the secret is absent.
  5. Tests. Use fixture API responses: pagination, a quota-exhausted 429, a URL with googleCanonical ≠ userCanonical, and an unmapped URL (which must fail).

Out of scope

Acceptance criteria

  • node scripts/seo-gsc-collect.mjs --fixtures tests/fixtures/gsc/ produces a snapshot and summary deterministically. The tests pass under npm run test:data.
  • Every URL in the three live sitemaps maps to exactly one family, and a test asserts it.
  • No credential, property secret or user-level data appears in committed output. A test greps the snapshot for private_key, client_email and similar.
  • After the owner completes the prerequisites, one live run is attached to this issue. It must show which families make up the "Crawled, currently not indexed" population. This answers the tracker's open question.

Guardrails

  • Use loadEnvFile(). Never add an env parser or read credentials from $HOME or an absolute path.
  • Any new file containing an external URL (googleapis.com) needs node scripts/source-attribution.mjs --write and then node scripts/docs-stats.mjs --check.
  • Never merge. Report the PR as ready.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1High priority, fix soonseoSEO, GEO, AI-search visibility, crawlability, public discovery, and technical SEO work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions