Skip to content

fix(seo): ~200 machine-readable URLs (resilience.json, *.openapi.yaml, .well-known/*) answer 200 with no X-Robots-Tag — the bulk of the actionable crawled-not-indexed population #8608

Description

@koala73

Parent: #8602 · Evidence: #8602 (comment)

Outcome

Every non-HTML URL we publish for machines — corpus dataset downloads, OpenAPI documents, .well-known descriptors — answers 200 and carries X-Robots-Tag: noindex, follow. It stays fetchable for agents and stops occupying Google's index queue.

Why

The #8602 drilldown parsed the 2026-09-24 "Crawled – currently not indexed" export (1,000 of 1,486 sampled) and re-probed every row live. After removing rows that already carry noindex or already redirect, 178 of 1,000 answer 200 today with no noindex, and they are almost entirely machine-readable files:

103  /countries/<slug>/resilience.json
 20  other .json  (countries/<slug>/cii.json, countries/resilience-ranking.json,
                   chokepoints/<slug>/reference.json, chokepoints/status.json,
                   crises/<slug>/tracker.json, sources/search-index.json,
                   docs/snapshots/*.json, openapi.json,
                   .well-known/{agent-skills/index.json,mcp/server-card.json,webhook-sample.json})
  8  /docs/api/*.openapi.yaml
 18  markdown + txt  (.well-known/agent-skills/*/SKILL.md, llms.txt, api/llms.txt,
                      pricing.md, ai-search.md, security.txt)

Zero /countries/<slug>/ or /crises/<slug>/ HTML pages appear in the sample. This is the largest actionable population the coverage export produced, and it is in no existing child issue.

Two things make it a real defect rather than cosmetics:

  1. Mintlify already solved this for its own twins and we did not for ours. https://www.worldmonitor.app/docs/pricing.md serves x-robots-tag: noindex; https://www.worldmonitor.app/pricing.md serves no such header. 25/25 probed Mintlify /docs/*.md twins carry it, 0 of ours do.
  2. The pattern is already established in our own config. vercel.json sets X-Robots-Tag: noindex, follow on /stocks, /stocks/(.*), /stocks.md and /blog/rss.xml. This issue applies the same header to the rest of the same class.

Link: <…>; rel="canonical" is already present on these files (#4999) and does not help — a self-canonical tells Google which URL is authoritative, not that the document should be left out of the index.

Why not robots.txt

robots.txt is the wrong tool here and #7660's reasoning applies directly. /countries/<slug>/resilience.json is discovered two legitimate ways, both of which a Disallow would break:

href="/countries/iran/resilience.json"                                          # visible download link
"contentUrl":"https://www.worldmonitor.app/countries/iran/resilience.json"      # schema.org DataDownload

A disallowed URL is never fetched, so the DataDownload claim in our structured data would point at something Google is forbidden to verify. noindex keeps the file fetchable and the claim honest.

Scope

  1. Data files — unambiguous, do these. Add X-Robots-Tag: noindex, follow in vercel.json headers for:

    • /countries/(.*)\.json, /chokepoints/(.*)\.json, /crises/(.*)\.json, /sources/(.*)\.json, /research/(.*)\.json
    • /docs/api/(.*)\.openapi\.yaml, /openapi.json, /openapi.yaml, /plugin.json
    • /docs/snapshots/(.*)\.json
    • /.well-known/(.*)\.json

    Match the existing entries' shape and placement. Do not touch robots.txt, robots.*.txt, sitemap*.xml, schemamap.xml or /blog/rss.xml — the sitemaps must stay plain, and rss already has the header.

  2. Agent-facing markdown and llms surfaces — owner decision required, do not guess. llms.txt, llms-full.txt, api/llms.txt, the 12 sitemap-declared .md twins, agents.md, developers.md, and .well-known/agent-skills/*/SKILL.md are the GEO/AI-citation surface. Adding noindex is correct for Google, but it is not established that OAI-SearchBot, PerplexityBot or Claude-SearchBot treat X-Robots-Tag: noindex as "do not index" rather than "do not cite". Until that is known, changing these risks trading a cosmetic Search Console count for AI-search citability, which is the opposite of the tradeoff we want.

    Deliver this half as a written recommendation on this issue, not a code change: name each surface, state what is known about each crawler's handling, and propose the call. The safe interim step that needs no crawler assumption is dropping the 12 non-HTML entries from sitemap-main.xml (scripts/build-sitemap.mjs:95-106) — a Google sitemap is a request to index, and these are not pages we want ranked. llms.txt discovery does not depend on the XML sitemap; robots.txt and the Link headers carry it.

  3. Gate. Extend tests/deploy-config.test.mjs (which already guards the X-Robots-Tag and Content-Signal contracts) so that every vercel.json headers source matching a data-file extension asserts noindex, and so sitemap-main.xml's generator cannot re-add a non-HTML <loc> without the test failing. Grep that file for the existing robots-tag helpers before adding a new one.

Out of scope

Acceptance criteria

  • After deploy, curl -sI -A "<Googlebot UA>" https://www.worldmonitor.app/countries/iran/resilience.json shows x-robots-tag: noindex, follow and still 200. Same for cii.json, a chokepoints/*/reference.json, a crises/*/tracker.json and a /docs/api/*.openapi.yaml. Paste the output in the PR.
  • The visible download link and the schema.org contentUrl on /countries/iran/ are unchanged, and the file still downloads. Structured data still validates.
  • tests/deploy-config.test.mjs passes, and fails when the noindex entry for /countries/(.*)\.json is removed. Prove it by deleting that entry locally and pasting the red output.
  • npm run typecheck and the touched npm run test:data suites pass.
  • Scope item 2 is answered in a comment on this issue with a named recommendation per surface, whether or not any code ships for it.

Guardrails

  • vercel.json is also mirrored by the nginx confs for self-hosting. Check whether deploy/ carries an equivalent robots-tag block before assuming vercel.json alone is the contract, and tests/deploy-config.test.mjs is the suite that catches the drift.
  • Editing vercel.json near the CSP entries can change an inline-script hash. Run tests/deploy-config.test.mjs locally before pushing.
  • Do not add a Disallow line. fix(seo): unbounded map-state URL space is burning crawl budget — 1,271 redirects and rising #7660 settled that robots.txt is a crawl control, not a canonicalization or de-indexing tool, and the DataDownload structured data depends on these URLs staying fetchable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Medium priority, schedule when capacity allowsseoSEO, GEO, AI-search visibility, crawlability, public discovery, and technical SEO work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions