You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix(seo): ~200 machine-readable URLs (resilience.json, *.openapi.yaml, .well-known/*) answer 200 with no X-Robots-Tag — the bulk of the actionable crawled-not-indexed population #8608
Every non-HTML URL we publish for machines — corpus dataset downloads, OpenAPI documents, .well-known descriptors — answers 200 and carries X-Robots-Tag: noindex, follow. It stays fetchable for agents and stops occupying Google's index queue.
Why
The #8602 drilldown parsed the 2026-09-24 "Crawled – currently not indexed" export (1,000 of 1,486 sampled) and re-probed every row live. After removing rows that already carry noindex or already redirect, 178 of 1,000 answer 200 today with no noindex, and they are almost entirely machine-readable files:
Zero /countries/<slug>/ or /crises/<slug>/ HTML pages appear in the sample. This is the largest actionable population the coverage export produced, and it is in no existing child issue.
Two things make it a real defect rather than cosmetics:
Mintlify already solved this for its own twins and we did not for ours.https://www.worldmonitor.app/docs/pricing.md serves x-robots-tag: noindex; https://www.worldmonitor.app/pricing.md serves no such header. 25/25 probed Mintlify /docs/*.md twins carry it, 0 of ours do.
The pattern is already established in our own config.vercel.json sets X-Robots-Tag: noindex, follow on /stocks, /stocks/(.*), /stocks.md and /blog/rss.xml. This issue applies the same header to the rest of the same class.
Link: <…>; rel="canonical" is already present on these files (#4999) and does not help — a self-canonical tells Google which URL is authoritative, not that the document should be left out of the index.
Why not robots.txt
robots.txt is the wrong tool here and #7660's reasoning applies directly. /countries/<slug>/resilience.json is discovered two legitimate ways, both of which a Disallow would break:
href="/countries/iran/resilience.json" # visible download link
"contentUrl":"https://www.worldmonitor.app/countries/iran/resilience.json" # schema.org DataDownload
A disallowed URL is never fetched, so the DataDownload claim in our structured data would point at something Google is forbidden to verify. noindex keeps the file fetchable and the claim honest.
Scope
Data files — unambiguous, do these. Add X-Robots-Tag: noindex, follow in vercel.jsonheaders for:
Match the existing entries' shape and placement. Do not touch robots.txt, robots.*.txt, sitemap*.xml, schemamap.xml or /blog/rss.xml — the sitemaps must stay plain, and rss already has the header.
Agent-facing markdown and llms surfaces — owner decision required, do not guess.llms.txt, llms-full.txt, api/llms.txt, the 12 sitemap-declared .md twins, agents.md, developers.md, and .well-known/agent-skills/*/SKILL.md are the GEO/AI-citation surface. Adding noindex is correct for Google, but it is not established that OAI-SearchBot, PerplexityBot or Claude-SearchBot treat X-Robots-Tag: noindex as "do not index" rather than "do not cite". Until that is known, changing these risks trading a cosmetic Search Console count for AI-search citability, which is the opposite of the tradeoff we want.
Deliver this half as a written recommendation on this issue, not a code change: name each surface, state what is known about each crawler's handling, and propose the call. The safe interim step that needs no crawler assumption is dropping the 12 non-HTML entries from sitemap-main.xml (scripts/build-sitemap.mjs:95-106) — a Google sitemap is a request to index, and these are not pages we want ranked. llms.txt discovery does not depend on the XML sitemap; robots.txt and the Link headers carry it.
Gate. Extend tests/deploy-config.test.mjs (which already guards the X-Robots-Tag and Content-Signal contracts) so that every vercel.jsonheaders source matching a data-file extension asserts noindex, and so sitemap-main.xml's generator cannot re-add a non-HTML <loc> without the test failing. Grep that file for the existing robots-tag helpers before adding a new one.
Out of scope
/docs/_next/static/chunks/* — 503 of the 1,000 sampled rows, all already serving noindex. Mintlify's assets, nothing to do.
Mintlify's /docs/*.md twins — already noindex.
The 428 stale 404 records and 245 stale non-www rows in the same export — already fixed upstream, they age out.
After deploy, curl -sI -A "<Googlebot UA>" https://www.worldmonitor.app/countries/iran/resilience.json shows x-robots-tag: noindex, follow and still 200. Same for cii.json, a chokepoints/*/reference.json, a crises/*/tracker.json and a /docs/api/*.openapi.yaml. Paste the output in the PR.
The visible download link and the schema.org contentUrl on /countries/iran/ are unchanged, and the file still downloads. Structured data still validates.
tests/deploy-config.test.mjs passes, and fails when the noindex entry for /countries/(.*)\.json is removed. Prove it by deleting that entry locally and pasting the red output.
npm run typecheck and the touched npm run test:data suites pass.
Scope item 2 is answered in a comment on this issue with a named recommendation per surface, whether or not any code ships for it.
Guardrails
vercel.json is also mirrored by the nginx confs for self-hosting. Check whether deploy/ carries an equivalent robots-tag block before assuming vercel.json alone is the contract, and tests/deploy-config.test.mjs is the suite that catches the drift.
Editing vercel.json near the CSP entries can change an inline-script hash. Run tests/deploy-config.test.mjs locally before pushing.
Parent: #8602 · Evidence: #8602 (comment)
Outcome
Every non-HTML URL we publish for machines — corpus dataset downloads, OpenAPI documents,
.well-knowndescriptors — answers 200 and carriesX-Robots-Tag: noindex, follow. It stays fetchable for agents and stops occupying Google's index queue.Why
The #8602 drilldown parsed the 2026-09-24 "Crawled – currently not indexed" export (1,000 of 1,486 sampled) and re-probed every row live. After removing rows that already carry
noindexor already redirect, 178 of 1,000 answer 200 today with nonoindex, and they are almost entirely machine-readable files:Zero
/countries/<slug>/or/crises/<slug>/HTML pages appear in the sample. This is the largest actionable population the coverage export produced, and it is in no existing child issue.Two things make it a real defect rather than cosmetics:
https://www.worldmonitor.app/docs/pricing.mdservesx-robots-tag: noindex;https://www.worldmonitor.app/pricing.mdserves no such header. 25/25 probed Mintlify/docs/*.mdtwins carry it, 0 of ours do.vercel.jsonsetsX-Robots-Tag: noindex, followon/stocks,/stocks/(.*),/stocks.mdand/blog/rss.xml. This issue applies the same header to the rest of the same class.Link: <…>; rel="canonical"is already present on these files (#4999) and does not help — a self-canonical tells Google which URL is authoritative, not that the document should be left out of the index.Why not robots.txt
robots.txtis the wrong tool here and #7660's reasoning applies directly./countries/<slug>/resilience.jsonis discovered two legitimate ways, both of which aDisallowwould break:A disallowed URL is never fetched, so the
DataDownloadclaim in our structured data would point at something Google is forbidden to verify.noindexkeeps the file fetchable and the claim honest.Scope
Data files — unambiguous, do these. Add
X-Robots-Tag: noindex, followinvercel.jsonheadersfor:/countries/(.*)\.json,/chokepoints/(.*)\.json,/crises/(.*)\.json,/sources/(.*)\.json,/research/(.*)\.json/docs/api/(.*)\.openapi\.yaml,/openapi.json,/openapi.yaml,/plugin.json/docs/snapshots/(.*)\.json/.well-known/(.*)\.jsonMatch the existing entries' shape and placement. Do not touch
robots.txt,robots.*.txt,sitemap*.xml,schemamap.xmlor/blog/rss.xml— the sitemaps must stay plain, and rss already has the header.Agent-facing markdown and llms surfaces — owner decision required, do not guess.
llms.txt,llms-full.txt,api/llms.txt, the 12 sitemap-declared.mdtwins,agents.md,developers.md, and.well-known/agent-skills/*/SKILL.mdare the GEO/AI-citation surface. Addingnoindexis correct for Google, but it is not established that OAI-SearchBot, PerplexityBot or Claude-SearchBot treatX-Robots-Tag: noindexas "do not index" rather than "do not cite". Until that is known, changing these risks trading a cosmetic Search Console count for AI-search citability, which is the opposite of the tradeoff we want.Deliver this half as a written recommendation on this issue, not a code change: name each surface, state what is known about each crawler's handling, and propose the call. The safe interim step that needs no crawler assumption is dropping the 12 non-HTML entries from
sitemap-main.xml(scripts/build-sitemap.mjs:95-106) — a Google sitemap is a request to index, and these are not pages we want ranked.llms.txtdiscovery does not depend on the XML sitemap;robots.txtand theLinkheaders carry it.Gate. Extend
tests/deploy-config.test.mjs(which already guards theX-Robots-TagandContent-Signalcontracts) so that everyvercel.jsonheaderssource matching a data-file extension assertsnoindex, and sositemap-main.xml's generator cannot re-add a non-HTML<loc>without the test failing. Grep that file for the existing robots-tag helpers before adding a new one.Out of scope
/docs/_next/static/chunks/*— 503 of the 1,000 sampled rows, all already servingnoindex. Mintlify's assets, nothing to do./docs/*.mdtwins — alreadynoindex.Acceptance criteria
curl -sI -A "<Googlebot UA>" https://www.worldmonitor.app/countries/iran/resilience.jsonshowsx-robots-tag: noindex, followand still200. Same forcii.json, achokepoints/*/reference.json, acrises/*/tracker.jsonand a/docs/api/*.openapi.yaml. Paste the output in the PR.contentUrlon/countries/iran/are unchanged, and the file still downloads. Structured data still validates.tests/deploy-config.test.mjspasses, and fails when thenoindexentry for/countries/(.*)\.jsonis removed. Prove it by deleting that entry locally and pasting the red output.npm run typecheckand the touchednpm run test:datasuites pass.Guardrails
vercel.jsonis also mirrored by the nginx confs for self-hosting. Check whetherdeploy/carries an equivalent robots-tag block before assumingvercel.jsonalone is the contract, andtests/deploy-config.test.mjsis the suite that catches the drift.vercel.jsonnear the CSP entries can change an inline-script hash. Runtests/deploy-config.test.mjslocally before pushing.Disallowline. fix(seo): unbounded map-state URL space is burning crawl budget — 1,271 redirects and rising #7660 settled that robots.txt is a crawl control, not a canonicalization or de-indexing tool, and theDataDownloadstructured data depends on these URLs staying fetchable.