Skip to content

Commit 7315862

Browse files
committed
feat(claude): 开放受保护的批量导出
1 parent f74bf09 commit 7315862

22 files changed

Lines changed: 1067 additions & 165 deletions

‎PLAN.md‎

Lines changed: 11 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77
## 目标
88

99
- 在 chatgpt.com / claude.ai 页内一键导出对话为 Obsidian 等笔记软件友好的 Markdown
10-
(ChatGPT 支持批量与增量;Claude 首版只做当前对话,理由见 P5)
10+
(两站均支持当前、选择、批量与增量导出;Claude 使用更保守的独立风控)
1111
- 高保真:公式、引用链接、代码、图片/附件、思维链、Canvas 不丢不乱
1212
- 增量同步:重跑只导出有变化的对话
1313
- 全程本地处理,不经任何第三方服务
@@ -77,15 +77,16 @@ inkstone/
7777
ir.ts # 中间表示:IRConversation / IRTurn / IRBlock
7878
render.ts # IR → Markdown(轮次标题、callout、围栏、frontmatter)
7979
fetcher.ts # 限速 / 退避 / 并发池 / 取消 / 限流观测(每站点一个实例)
80+
batch-safety.ts # 站点级并发/重试策略 + 429/请求预算/失败率熔断
8081
sites/
8182
types.ts # SiteAdapter 契约(取数 + 转换 + 界面锚点 + 批量能力)
8283
index.ts # 按 location.host 分派
8384
chatgpt/
8485
index.ts # adapter 实装
8586
convert.ts # backend-api JSON → IR(content_type 分发、canmore 语义)
8687
claude/
87-
index.ts # adapter 实装(supportsBatch: false)
88-
api.ts # 内部 API 客户端 + 保守限流参数 + 分页器(未接界面)
88+
index.ts # adapter 实装 + Claude 专属批量风控策略
89+
api.ts # 内部 API 客户端 + 保守限流参数 + 有界分页器
8990
types.ts # 从宽的字段类型,[待测] 处已标注
9091
convert.ts # 内部 API JSON → IR(块级分发、主线回溯、附件两处来源)
9192
artifacts.ts # artifact create/update/rewrite 折叠成终稿
@@ -103,7 +104,7 @@ inkstone/
103104
fsaccess.ts # File System Access 直写 vault(句柄存 IndexedDB)
104105
test/
105106
fixtures/*.json # 对话 JSON(ChatGPT 真实脱敏 / Claude 合成)
106-
*.test.ts # bun test(132 个)
107+
*.test.ts # bun test 回归套件
107108
```
108109

109110
## 阶段
@@ -130,13 +131,13 @@ inkstone/
130131
- **Claude 侧的脏活更少**:Canvas 的正则 patch 重放(150 行)与私有区 Unicode 引用
131132
还原(178 行)在 Claude 都不需要——artifact 的 update 是字面量 `old_str`→`new_str`,
132133
引用是结构化数组。artifact 折叠约 40 行。
133-
- **⚠️ 首版刻意只做「导出当前对话」**:批量的地基(分页器、水位线、并发池、
134-
保护性中止)全部就位且已单测,但 `supportsBatch: false` 关着。理由是限流画像
135-
未知——ChatGPT 侧的参数是 344 + 432 对话实测调出来的,Claude 侧一条实测数据
136-
都没有。调研过的三个开源 claude.ai 导出器**没有一个实现了 429 退避**
137-
(最激进的是 3 并发 + 固定 200ms 间隔且不看 429),所以没有可借鉴的安全参数。
134+
- **批量分阶段开放**:首版因限流画像未知而只开放当前对话;2026-08-29 恢复选择、
135+
全部与增量导出,但不照搬 ChatGPT 参数。Claude 固定单并发、不做失败项整批二次重试;
136+
一次带 `Retry-After` 的 429、累计 3 次 429、单批 1000 次 HTTP 尝试,或至少 5 条失败
137+
且失败率超过 25%,都会保护性中止。成功条目才推进水位线,未尝试部分由下次增量补齐。
138+
列表分页另设 250 请求 / 10000 条硬上限,并检测重复页与缺失 uuid,防接口漂移后空转。
138139
- **Claude 限流起步参数**(保守,待实测调整):间距 1500ms(ChatGPT 侧的两倍慢)、
139-
上限 8000ms、每 40 请求歇 30s、最多重试 6 次。每个站点持有独立的 fetcher 实例,
140+
上限 8000ms、每 40 请求歇 30s、首次失败后最多重试 1 次。每个站点持有独立的 fetcher 实例,
140141
一边的限流不拖累另一边。吃到 429 时导出完成文案会报出次数、被推大的间距与
141142
服务端要求的最长等待——未知站点的节奏只能靠实测看清,先让它可见再谈调参。
142143
- **待实测**:`docs/claude-probe.js` 可直接粘进 claude.ai 控制台,打印字段骨架

‎README.md‎

Lines changed: 11 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -36,18 +36,18 @@ Inkstone runs inside the page and fetches conversations through the same backend
3636
| | ChatGPT | Claude |
3737
| --- | --- | --- |
3838
| Export current conversation | ✅ | ✅ |
39-
| Batch / export-all | ✅ | ⏳ not yet enabled |
40-
| Incremental sync | ✅ | ⏳ not yet enabled |
39+
| Batch / export-all | ✅ | ✅ conservative safeguards |
40+
| Incremental sync | ✅ | ✅ |
4141
| Rich documents | Canvas patch replay | Artifact fold-up to final version |
4242
| Thoughts / tool traces | ✅ opt-in | ✅ opt-in |
4343
| Attachments | images and files downloaded | images, documents, and generated files downloaded; text extractions inlined |
4444

45-
**Why no batch export on Claude yet?** It isn't missing, it's switched off. The pager,
46-
watermark, concurrency pool and protective abort are all in place and unit-tested — but
47-
there is no measured rate-limit profile for Claude yet. The ChatGPT numbers only became
48-
trustworthy after 344 + 432 real conversations. Until comparable evidence exists, the cost
49-
of a wrong guess lands on your account, and that isn't a call a default-on switch should
50-
make. See [`docs/claude-adapter-feasibility.md`](./docs/claude-adapter-feasibility.md).
45+
**Claude batch export uses a separate, deliberately conservative policy.** Its rate-limit
46+
profile is still not backed by a large real-world sample, so ChatGPT's settings are not
47+
reused: one worker, spacing from 1500 ms, a 30-second rest every 40 requests, and no second
48+
pass over failed items. One global `Retry-After` signal, three 429s, 1000 HTTP attempts, or
49+
an abnormal failure ratio stops the batch. Successful conversations are still written;
50+
unfinished ones do not advance the watermark and are picked up by the next incremental run.
5151

5252
## Screenshots
5353

@@ -117,10 +117,10 @@ bun run build # → dist/inkstone.user.js, drag it into Tampermonkey
117117

118118
Open chatgpt.com or claude.ai (logged in) → click the **⤓ button** in the top bar → pick **Markdown zip** or **raw JSON zip** → unzip into your Obsidian vault.
119119

120-
- On Claude only **current conversation** is offered; the batch options are hidden, not disabled
121120
- The button position is switchable (panel → advanced settings): next to Share in the top bar, or a glass button beside the input box
122-
- The UI follows the host page's appearance automatically (light/dark + accent color)
121+
- The UI follows the host page's light/dark appearance; ChatGPT follows its selected accent, while Claude uses its brand orange
123122
- Exports are cancelable; a single failed conversation never aborts the run — failures are summarized in `_failures.json`
123+
- Claude batches run with one worker; after a protective abort, wait as instructed instead of restarting immediately
124124

125125
## Offline CLI
126126

@@ -154,7 +154,7 @@ Adding a site means adding an adapter, not touching the orchestration. See `PLAN
154154

155155
## Roadmap
156156

157-
Batch export on Claude once its rate-limit profile has actually been measured, an MV3 browser extension (no Tampermonkey, store release), and Gemini support. Already done: the multi-site adapter architecture, Claude single-conversation export, incremental sync, direct-write to an Obsidian vault, settings panel, Canvas patch replay, Artifact fold-up, and the offline CLI. Details in [PLAN.md](./PLAN.md) (Chinese).
157+
An MV3 browser extension (no Tampermonkey, store release) and Gemini support. Already done: the multi-site adapter architecture, Claude batch and incremental export, direct-write to an Obsidian vault, settings panel, Canvas patch replay, Artifact fold-up, and the offline CLI. Details in [PLAN.md](./PLAN.md) (Chinese).
158158

159159
## License
160160

‎README.zh-CN.md‎

Lines changed: 11 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -36,17 +36,17 @@ Inkstone 直接运行在页内,通过应用自己使用的 backend API 抓取
3636
| | ChatGPT | Claude |
3737
| --- | --- | --- |
3838
| 导出当前对话 | ✅ | ✅ |
39-
| 批量 / 全部导出 | ✅ | ⏳ 暂不开放 |
40-
| 增量同步 | ✅ | ⏳ 暂不开放 |
39+
| 批量 / 全部导出 | ✅ | ✅ 保守风控 |
40+
| 增量同步 | ✅ | ✅ |
4141
| 富文档还原 | Canvas patch 重放 | Artifact 折叠还原终稿 |
4242
| 思维链 / 工具痕迹 | ✅ 可开关 | ✅ 可开关 |
4343
| 附件 | 图片与文件下载 | 图片下载、文档链接、文本抽取件内联 |
4444

45-
**Claude 端为什么先不做批量?** 不是没写,是没开。批量所需的分页器、水位线、并发池、
46-
保护性中止都已就位并通过单测,但 Claude 侧的限流画像还没有任何实测数据——
47-
ChatGPT 端那套参数是 344 + 432 条对话跑出来才敢用的。在拿到同等的实测证据之前,
48-
批量抓取整个历史的风险由用户账号承担,这个代价不该由一个默认开启的开关来决定。
49-
细节与实测计划见 [`docs/claude-adapter-feasibility.md`](./docs/claude-adapter-feasibility.md)。
45+
**Claude 批量导出采用更保守的独立风控。** Claude 的限流画像仍没有大样本实测,
46+
因此不照搬 ChatGPT 参数:固定单并发、请求间隔从 1500ms 起、每 40 次请求休息 30 秒,
47+
不对整批失败项做第二轮重试。一次带 `Retry-After` 的全局限流信号、累计 3 次 429、
48+
1000 次 HTTP 尝试或异常失败率过高都会保护性中止;已经成功的对话照常落盘,未完成项
49+
不推进水位线,下次增量导出会继续补齐。
5050

5151
## 截图
5252

@@ -116,8 +116,9 @@ bun run build # 产物 dist/inkstone.user.js,拖进 Tampermonkey 即可
116116
打开 chatgpt.com 或 claude.ai(已登录)→ 点**顶栏 Share 左侧的 ⤓ 按钮** → 选 **Markdown zip** 或**原始 JSON zip** → 解压到 Obsidian vault。
117117

118118
- 按钮位置可换(面板 → 高级设置):顶栏 Share 旁,或输入框旁的玻璃圆钮
119-
- UI 主题色自动跟随 ChatGPT 的外观设置(明暗 + accent color)
119+
- UI 明暗自动跟随所在站点;ChatGPT 跟随用户选择的重点色,Claude 使用品牌橙色
120120
- 可随时取消;单条对话失败不中断整体导出,失败汇总进 `_failures.json`
121+
- Claude 批量默认单并发慢速执行;触发保护性中止后不要立刻重跑,先按界面提示等待
121122

122123
## 离线 CLI
123124

@@ -151,12 +152,12 @@ claude.ai 的 CSP 可能拦掉 dev server 的脚本,Claude 端的改动请用
151152

152153
## 路线图
153154

154-
Claude 端的批量导出(等限流画像实测清楚再开)、MV3 浏览器扩展(脱离 Tampermonkey、上架商店)、Gemini 适配。已完成:多站点适配器架构、Claude 单对话导出、增量同步、直写 vault、设置面板、Canvas patch 重放、Artifact 折叠还原、离线 CLI。详见 [PLAN.md](./PLAN.md)。
155+
MV3 浏览器扩展(脱离 Tampermonkey、上架商店)、Gemini 适配。已完成:多站点适配器架构、Claude 批量与增量导出、直写 vault、设置面板、Canvas patch 重放、Artifact 折叠还原、离线 CLI。详见 [PLAN.md](./PLAN.md)。
155156

156157
## 许可证
157158

158159
[GPL-3.0](./LICENSE)
159160

160161
## 友情链接
161162

162-
[LINUX DO](https://linux.do)
163+
[LINUX DO](https://linux.do)

‎docs/claude-adapter-feasibility.md‎

Lines changed: 4 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,9 @@
11
# Inkstone → Claude 对话导出:可行性分析
22

3-
> **实施状态(2026-08-28)**:本文第四节的架构改造与第六节的 P1–P3 已完成,
4-
> Claude 单对话导出可用;批量(P4)按第五节的风险判断**刻意未开放**。
5-
> 落地记录见 `PLAN.md` § P5,待实测清单见本文第五节,探针脚本见 `docs/claude-probe.js`。
3+
> **实施状态(2026-08-29)**:本文第四节的架构改造与第六节的 P1–P4 已完成。
4+
> Claude 的当前、选择、全部与增量导出均已接入;因限流画像仍缺少大样本实测,P4
5+
> 使用单并发、独立慢速 Fetcher、429/Retry-After 熔断、1000 请求预算与低失败率阈值,
6+
> 不做整批二次重试。落地记录见 `PLAN.md` § P5,探针脚本见 `docs/claude-probe.js`。
67
> 本文其余部分保持评估当时的原貌,不随实施回填——它是决策依据的快照。
78
89
> 评估日期:2026-08-28 · 基准代码:`2121e9f`(v0.2.3,与上游 ZhenHuangLab/inkstone 同步)

0 commit comments

Comments
 (0)