Common mistakes and confusion points in this project. Add to this list when something surprises you.
-
There is no official Substack API for what this project does. Every endpoint here was found by observing the web app, and any of them can change without notice. Substack does run a separate, limited "Developer API" (support.substack.com article 45099095296916, terms at substack.com/api-tos), and it is not the endpoints used here.
substackapi.devis a different product with its own API keys; its per-minute limits don't apply to us. -
No published rate limits for the web app endpoints (checked 2026-09-25). The Developer API terms only say limits are "determined by Substack in its sole discretion", with no numbers.
-
https://<sub>.substack.com/api/...answers 301 to the custom domain when the publication has one. That's whySubstackHttpfollows redirects manually and decides per hop whether to attach the session cookie. Don't switch toredirect: "follow". Whether fetch strips a manually setCookieheader on cross-origin redirects is not guaranteed, so we don't rely on it. -
/api/v1/posts/<slug>returns the post object directly, but/api/v1/posts/by-id/<id>wraps it as{ post: ... }.SubstackClient.posthandles both. -
Paywalled posts return HTTP 200 with a truncated
body_html, not an error, and no field says whether you got the full post. The reliable signal: when the viewer has access,body_htmlcontains<div class="paywall-jump">where the paywall would be, and previews stop just before it. Word count alone is not enough: Lenny's Newsletter serves about 85% of the article as a free preview.isPreviewinformat.tschecks the marker first and uses word count only as a backup. Verified against real posts on 2026-09-25. -
audience: "only_paid"(thepaywalledflag in summaries) describes the post, not the viewer. It istrueeven for publications the user pays for. Paid podcasts often have only ~50 words of show notes as their full text; the paywall covers the audio. -
/api/v1/archive?search=...is the search endpoint. It is per publication; there is no cross-subscription search. -
Anonymous visitors also get a
substack.sid, so having the cookie doesn't mean being logged in. Always validate with/api/v1/user/profile/self(401 when not logged in). -
The subscription list comes from
/api/v1/user/profile/self, not/api/v1/subscriptions.profile/selfreturns every subscription, including ones hidden from the public profile, each with a nestedpublicationobject./api/v1/subscriptionslooks like the right endpoint, but it returns 400 unless you passtvOnly, and even then it gives an unrelated subset of publications with an emptysubscriptionsarray. Verified with a real account on 2026-09-25. -
Custom domains do not accept
substack.sid. For example,www.oneusefulthing.org/api/v1/subscriptionreturns 404 "Subscription not found" even though the user is subscribed. The browser's workaround, whichSubstackHttp.customDomainSessionreproduces:GET substack.com/sign-in?redirect=%2F&for_pub=<subdomain>with the session redirects (303) to<custom domain>/api/v1/sign-in/local/complete?token=…, and that response setsconnect.sidon the custom domain. The handoff needs the publication's subdomain, which is whytrustHosttakes it. Doing this does not rotate or invalidatesubstack.sid(verified 2026-09-25). -
Subscription endpoints (found in the web app's JS bundles and checked live):
GET substack.com/api/v1/subscription/<publicationId>returns{ subscription, publication }for any publication, custom domain or not; for no subscription it returnssubscription: nullor 404. Good for checking before and after a change.GET <pub>/api/v1/subscriptiongives the payment details:is_subscribed,stripe_subscription_id,expiry, and so on. Free subscriptions havemembership_state: "free_signup",is_subscribed: false, and no Stripe ID; paid ones have"subscribed",true, and a Stripe ID.POST <pub>/api/v1/free {email, source, first_url, ...}subscribes. It requires the account email, which only appears in page data:window._preloads.user.emailonsubstack.com/settings.- Free unsubscribe is
DELETE <pub>/api/v1/free {publication_id, source: "account"}, the mirror of subscribe. It lives only in the scripts for the publication's/accountpage, which load on demand, so a grep of the homepage bundles won't find it. DELETE <pub>/api/v1/subscriptionis a trap. It's the paid cancellation endpoint (the web app sends{force_now: true}). For a free subscription it returns200 {}and changes nothing, with or withoutforce_now,Origin, orReferer(verified 2026-09-25). Never use it inunsubscribe: on a paid subscription it would cancel the plan.unsubscribestill checkspaidSignal()before deleting and refuses anything that isn't unambiguously free. Don't loosen that. Afterwards it re-checks through the substack.com lookup, because a 200 on its own proves nothing.- To find endpoints, load the actual page with the saved browser profile (
playwright-core, headless) and collect every script it loads. Downloading only the bundles referenced by the page HTML misses code that loads on demand.
-
window._preloadson any publication homepage haspub(id, subdomain, custom_domain).publicationInfouses it to turn a URL or subdomain into canonical details. -
Some well-known newsletters have left Substack (Platformer moved to Ghost), so "doesn't look like a Substack publication" can be the correct answer.
-
Chat (publication community chats and DMs) all lives on
substack.com, so no custom-domain handoff is needed, and it always requires a login:GET /api/v1/messages/inbox?tab=alllists items oftype: "chat"(withpublicationandcommunityPost) and"direct-message"(withmessageThread; read it withmessageThread.id, not thedirect-message-…item id).GET /api/v1/community/publications/<pubId>/posts[?before=<created_at>]lists threads ({threads: [{communityPost, user}], moreBefore}). Pages hold 25. A publication without a chat returns 404.GET /api/v1/community/posts/<threadId>/comments?order=asc&initial=truereturns a window of replies withmoreBefore/moreAfter. Page withbefore_id=<first id>andafter_id=<last id>. Replies to a reply use the same scheme at/community/comments/<commentId>/comments.GET /api/v1/messages/dm/<messageThread.id>returns{replies: [{comment, user}], profile}.
-
Chat unread flags are useless. No inbox item had
communityPost.has_unread_commentsset in three digest runs (2026-09-30 to 10-02), andpubChatUnreadCountstays the same because nothing in this server marks chats as seen. UseSubstackChat.activity(since)instead. -
The inbox
timestamp(ourlastActivity) doesn't move when someone replies to an older thread. Seen 2026-10-02: Nate's chat had inbox time2026-10-01T22:00:42.514Z, equal to the newest thread'screated_at, while that thread'smost_recent_comment_created_atwas2026-10-02T00:49. Activity has to come from each chat's thread list:max(created_at, most_recent_comment_created_at). Threads are ordered by creation, so a reply to a thread beyond the first few pages is missed. -
Archive requests get rate-limited. At 6 concurrent archive requests with 500 ms / 1 s backoff, 7 of 35 publications failed with 429 on 2026-10-02, and 20 failed on each of two calls 17 s apart on 09-26. Now: concurrency 3, 429 backoff 2/5/10 s with jitter, then retry passes after 15 s and 30 s, one publication at a time. No numbers are published (see above), so these are guesses that worked.
-
Don't open chat pages in a browser to investigate. Loading
substack.com/chatmakes the pagePOST /api/v1/messages/inbox/seen, which clears the user's unread badges. The API GETs above don't do that (as far as observed). This happened once, on 2026-09-25. -
Reader shelves (found 2026-10-03 in the web app's public JS bundles, checked live):
GET substack.com/api/v1/reader/posts?inboxType=<seen|saved|archived|inbox|recommended>&limit=20[&cursor=…]returns{posts, publications, inboxItems, postReactions, savedPosts, more, cursor}. Thepostslack publication names, which come frompublicationsbypublication_id.inboxItems[].max_read_progress(0–1) says how far the user read in the web or app reader; email reads don't appear.postReactionsare the user's own hearts, andinboxType=archivedis posts the user dismissed. Likes:GET /api/v1/reader/feed/profile/<userId>?types=likereturns{items, nextCursor}, wheretype: "post"withcontext.type: "post_like"is a hearted post and"comment"/"note_like"is a liked note (skipped). -
Endpoints that look read-only but write:
/api/v1/posts/savedis only save (POST) and unsave (DELETE); list saved posts withinboxType=savedinstead./api/v1/posts/<id>/seen,/api/v1/reader/feed/<key>/seen,/api/v1/inbox/seenand/api/v1/reader/feed/<key>/dismisschange inbox state. Never call them. The GETs above don't mark anything.
- In server mode stdout is the MCP protocol channel. Never
console.logfrom server code paths; useconsole.error. CLI subcommands (status,logout,install) may use stdout. - Claude Code does not read MCP servers from
~/.claude/settings.json. The original Python version'smake configurewrote them there, so that registration never took effect. Useclaude mcp add(whatcli.ts installdoes) or~/.claude.json/.mcp.json.
-
Third-party docs are linked, never copied into the repo.
docs/jev/README.mdis our own summary with links to the TypeSafe and OpenRouter pages; keep local copies in the gitignoreddocs/jev/upstream/. -
The version is in two places:
package.jsonandVERSIONinsrc/server.ts(what MCP clients see inserverInfo). Bump both. -
docs/digest-tools.mdis checked bytest/docs-example.test.ts. Changing a digest tool's output or input fields fails that test until the doc is updated. Regenerate the output blocks withUPDATE_DOCS=1 npx vitest run test/docs-example.test.ts, review the diff, and update the argument tables and prose by hand. -
TypeScript is 7.x, the native compiler. Some older tsconfig options (
baseUrl,moduleResolution: node) no longer exist. -
If vitest fails with
Cannot find native binding(rolldown), it's the npm optional-dependency bug: deletenode_modulesandpackage-lock.json, then runnpm installagain. -
playwright-coreis an optional dependency and is imported dynamically only bylogin. Don't import it at the top level of anything the server loads.
digest_begin,digest_finish,digest_statusandmark_reportedare registered only whenSUBSTACK_DIGEST_DIRis set (absolute path), so the public tool list doesn't change. Paths come only from env, never from tool arguments.get_chat_activityis always registered.- Two server processes run at once in the Hermes container (seen 2026-10-02), so anything shared between calls lives in files in the digest dir, never in memory. Every write takes
state.lock(open(..., "wx"), stale after 60 s) and goes through write-then-rename. Symlinks in the digest dir are refused. digest_beginwrites onlycurrent_run.json;digest_finishcommitsstate.jsonand renames the run toprevious_run.jsonwith its rendered result, which is how a repeateddigest_finishcall returns the same text. Iflast_runchanged between begin and finish, finish refuses.reported_postsis deliberately the last key instate.json. The v2.x skill appended URLs by patching the end of that list, so keeping it last lets the old skill run against v2 state if the skill is rolled back.- Digest output is plain text, never JSON. Hermes wraps every MCP text result as
{"result": "<text>"}, so JSON comes out double-escaped and hard for a small model to copy. Don't addstructuredContentor anoutputSchema. Keep each result under ~40,000 characters (Hermes drops results over ~50k). - Hermes pauses an MCP server for 60 s after 3 consecutive error results, so
digest_finishis lenient (case-insensitive refs, duplicates resolved, unknown refs dropped) and reports warnings instead of errors wherever that's safe. - Hermes's MCP tool timeout is 300 s.
digest_begingives the post fetch about 140 s and leaves the rest for chats. Intl.DateTimeFormatcan't combinedateStylewithtimeZoneName(it throws), soformatLocalspells out the fields. Newer ICU puts U+202F before AM/PM;time.tsreplaces it with a space.
hermes/SKILL.mdis the generic, shareable copy of the digest skill; its settings are in the Settings block at the top. Since v3.0.0 it uses only the digest tools,get_chat_activity,read_postanddelegate_task, and must not use file tools. The copy running on the maintainer's Hermes host has those values filled in, and the two are kept in sync by hand. When you change one, change the other, and bumpversionin the frontmatter.- The skill is a prompt, so it can't be unit-tested. Check changes by running the cron job once (
hermes cron run <id>) and reading the tool calls in Hermes'sstate.dbmessagestable. That's how the result-size and missing-sinceproblems were found.
- Shared with medium-reader-mcp.
src/classifier/*andtools/classifier/{lib,collect,label,score,analyze}.mjsare identical in both repos; onlytools/classifier/source.mjsdiffers. Change both copies together, and keepsrc/classifier/free of Substack imports. docs/classifier.md explains the design; the evaluation behind it is in medium-reader-mcp'sexperiments/jev/. - Off by default here (
classifierFromEnv("SUBSTACK", …, "off")), unlike Medium, wheresamplingis the default. The Substack digest has always read every post, so turning a classifier on must be a choice. When it's off,digest_begin's output is exactly what it was before (no CLASSIFIER line, no rank column). - What the classifier changes: POSTS are sorted by rank with a rank column. Skipped posts (
skip ≥ SUBSTACK_DIGEST_SKIP_THRESHOLD) and below-floor posts (SUBSTACK_DIGEST_RANK_FLOOR) get their own lists.reconciledoesn't count them as missing: skipped ones go in the 🗑 line, and unread low ones under 📎 "Also new". All of them are still saved as reported. - The classifier's deadline is real time.
digestBeginhands itDate.now() + time left, notfetchStart + budget, because tests injectdeps.nowand the classifier uses the real clock. A timestamp on the injected clock made it give up at once ("no time left to classify"). runs/archive:digest_finishsaves each committed run asruns/<run_id>.jsonwithjudged: {picks, others}(newest 14,archiveRunin state.ts). Nothing in the digest reads it; it's history fortools/classifier/collect.mjs. A failed archive write never fails the run.interests.mdsections are parsed byparseProfile(## Interests,## Skip; a file without them is all interests). The model still sees the raw file in INTERESTS.OPENROUTER_API_KEYfor the tools can live in.env(gitignored); run them withnode --env-file=.env.- interests_evidence / save_interests_proposal:
src/classifier/{evidence,proposal}.tsare shared with Medium.src/digest/evidence.tsis Substack's gatherer. Labels and the dataset live in<digest dir>/classifier/, and runs inruns/; they're read throughsubdirPath, which refuses symlinks, sincesafePathonly takes flat names. Proposals are tested on held-out labels:isTrainLabelandanalyze.mjs --test-halfshare one split.