Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 25 additions & 0 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
name: tests

on:
push:
branches: [main]
pull_request:

permissions:
contents: read

jobs:
offline:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.13"]
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
with:
python-version: ${{ matrix.python-version }}
- run: python -m pip install --upgrade pip
- run: python -m pip install -r scripts/requirements.lock pytest==8.4.2
- run: PYTHONPATH=. python -m pytest -q
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,9 @@ git clone https://github.com/O0000-code/paper-search-pro.git \
export PSP_HOME="$HOME/.claude/skills/paper-search-pro"

python3 -m pip install -r "$PSP_HOME/scripts/requirements.txt"

# Optional: install the exact dependency versions validated by CI.
python3 -m pip install -r "$PSP_HOME/scripts/requirements.lock"
```

Five free API keys (~15 min total) — see [`references/setup.md`](references/setup.md).
Expand Down Expand Up @@ -134,7 +137,7 @@ Three tabs · three hero layouts · two list densities · responsive at 860 px
</tr>
</table>

**Journal partitions.** Every paper is tagged with its 中科院 (CAS) / JCR / SJR tier — a quiet badge on each card, the full three-platform breakdown in the detail panel, and a zone filter (`Q1 / ≥Q2 / ≥Q3`, or `一区 / ≥二区`) in the toolbar. The tables are fetched at runtime from public mirrors — never bundled — and attributed; only JCR's figure is labelled an impact factor.
**Journal partitions.** Every paper can be tagged with its 中科院 (CAS) / JCR / SJR tier — a quiet badge on each card, the full three-platform breakdown in the detail panel, and a zone filter (`Q1 / ≥Q2 / ≥Q3`, or `一区 / ≥二区`) in the toolbar. Ranking datasets and download URLs are not bundled: provide compatible CSVs in the local cache, or explicitly configure source URLs you are authorised to use. Only JCR's figure is labelled an impact factor.

Outputs land in `$PWD/paper-search-results/<search_id>/`:

Expand Down
5 changes: 4 additions & 1 deletion README.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,9 @@ git clone https://github.com/O0000-code/paper-search-pro.git \
export PSP_HOME="$HOME/.claude/skills/paper-search-pro"

python3 -m pip install -r "$PSP_HOME/scripts/requirements.txt"

# 可选:安装 CI 验证过的精确依赖版本。
python3 -m pip install -r "$PSP_HOME/scripts/requirements.lock"
```

5 个免费 API key(共约 15 分钟)— 见 [`references/setup.md`](references/setup.md)。
Expand Down Expand Up @@ -134,7 +137,7 @@ python3 -m pip install -r "$PSP_HOME/scripts/requirements.txt"
</tr>
</table>

**期刊分区。** 每篇论文都标注其 **中科院 / JCR / SJR** 分区 —— 卡片上一个安静的徽章、详情面板里三家完整对照、工具栏里按分区筛选(`Q1 / ≥Q2 / ≥Q3`,中科院则是 `一区 / ≥二区`)。分区表在运行时从公开镜像拉取 —— 不随包分发 —— 并标注来源;只有 JCR 的数值才称为影响因子。
**期刊分区。** 每篇论文都可标注其 **中科院 / JCR / SJR** 分区 —— 卡片上一个安静的徽章、详情面板里三家完整对照、工具栏里按分区筛选(`Q1 / ≥Q2 / ≥Q3`,中科院则是 `一区 / ≥二区`)。仓库不附带分区数据或下载地址:请把兼容 CSV 放入本地缓存,或显式配置你有权使用的来源 URL。只有 JCR 的数值才称为影响因子。

输出落在 `$PWD/paper-search-results/<search_id>/`:

Expand Down
13 changes: 7 additions & 6 deletions SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: paper-search-pro
description: "Find academic papers across up to 7 sources (OpenAlex / Semantic Scholar / CrossRef / PubMed / arXiv for English, plus native-Chinese retrieval via NSSD 国家哲社文献中心 + yiigle 中华医学期刊) with adjustable depth — Quick scan (5 min) to Audit prep (3 hr). Use when the user wants to find papers, run a literature search, gather references, scope a research topic, search Chinese-language / 中文原生 literature (中文文献/中文核心/CSSCI/C刊/国内研究/国内文献/中华××期刊/心理学报/经济研究), or filter results by journal tier (中科院分区/一区/几区, Q1, JCR/SJR quartile, 影响因子/impact factor, 期刊分区, 顶刊/top journal, '按分区筛'). Triggers on search verbs ('find papers', 'literature search', 'papers about X'), review types ('scoping review', 'systematic review', 'SR prep', 'literature review', 'lit review', 'help me write a lit review'), Chinese ('找文献', '找论文', '论文搜索', '学术检索', '文献检索', '文献综述', '综述前期', '求文献', '中文文献', '中文核心', 'CSSCI', 'C刊', '国内研究', '找中文的'). Outputs Shadcn HTML report + BibTeX/RIS/CSV + PRISMA-S log. Do NOT use for: concept explanations ('what is X' / 'X 是什么', e.g. '影响因子怎么算'), writing ('帮我写' / 'help me write a paragraph'), single-paper interpretation or PDF download with metadata (use paper-downloader-portable), or when the user already has a literature set (use literature-set-review)."
description: "Find academic papers across OpenAlex, Semantic Scholar, CrossRef, PubMed, arXiv, NSSD 国家哲社文献中心, and yiigle 中华医学期刊, with Quick-to-Audit depth. Use for literature search, gathering references, scoping a topic, review preparation (literature/scoping/systematic/meta-analysis), Chinese literature (找文献/找论文/文献检索/中文文献/中文核心/CSSCI/C刊/国内研究), and journal-tier filters (中科院分区/一区/Q1/JCR/SJR/影响因子/顶刊). Produces an HTML report plus BibTeX, RIS, CSV, and PRISMA-S log. Do not use for concept explanations, prose drafting, interpreting a known paper, PDF downloads (use paper-downloader-portable), or reviewing an existing literature set (use literature-set-review)."
license: Apache-2.0
allowed-tools: Bash, Read, Write, Edit, Glob, Grep, Task
metadata:
Expand Down Expand Up @@ -449,11 +449,11 @@ everything else in this step it is **opt-in and off by default — skip it and t
report is byte-for-byte unchanged** (R-19). 📖 Read `references/journal_metrics.md`
first (it is the SSOT for sources, the ISSN join, attribution, and R-04 naming).

- **First use needs a one-time fetch** (init-once; data is pulled at runtime into
`~/.paper-search-pro/ranks/` and **never bundled in the repo**). If you have not
fetched before, run it once (and tell the user it is a one-time step):
- **First use needs user-provided data.** The repo bundles neither ranking data
nor download URLs. Put compatible CSVs in `~/.paper-search-pro/ranks/`, or
configure URLs the user is authorised to use under `rank.sources` before fetch:
```bash
PYTHONPATH=$PSP_HOME python3 -m scripts.journal_rank fetch # all three
PYTHONPATH=$PSP_HOME python3 -m scripts.journal_rank fetch # configured sources
# or a single platform: ... journal_rank fetch --platform cas
PYTHONPATH=$PSP_HOME python3 -m scripts.journal_rank info # what's cached
```
Expand All @@ -469,7 +469,8 @@ first (it is the SSOT for sources, the ISSN join, attribution, and R-04 naming).
`journal_rank.load()` returns **None** when nothing is cached — then this layer
silently degrades (no partitions; the OpenAlex open-impact figure from the block
above is still the influence placeholder) and you tell the user they can
`journal_rank fetch` to enable partitions.
configure `rank.sources` before `journal_rank fetch`, or populate the cache, to
enable partitions.
- **Filter (only when a tier was requested — see STEP 11 for the full flow):** call
`rank_filter.filter_by_rank(papers, platform, tiers=…, quartiles=…, top=…)`. It
returns `(kept, filtered_out, no_platform_data)` — the third bucket (journals not
Expand Down
4 changes: 4 additions & 0 deletions THIRD_PARTY.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,10 @@ homepage links below carry the canonical license file.

## Python runtime (`scripts/requirements.txt`)

`scripts/requirements.lock` is a generated, exact-version snapshot of these
same runtime packages and their transitive dependencies. It adds no runtime
component; `requirements.txt` remains the default portable install path.

| Package | License | Homepage |
|---|:---:|---|
| **pyalex** — OpenAlex client | MIT | <https://github.com/J535D165/pyalex> |
Expand Down
11 changes: 4 additions & 7 deletions assets/default_config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -88,7 +88,7 @@ quota_fallback_threshold_usd: 0.05
# Journal rank (v2.2 Feature A — additive; omit entirely for v2.1 behavior)
# ============================================================================
# Multi-platform journal partition / quartile (CAS 中科院 / JCR / SJR). Data is
# fetched at runtime from public mirrors into cache_dir NEVER bundled in-repo.
# read from cache_dir and is NEVER bundled in-repo. Downloads are opt-in.
# CAS 区 & SJR quartile are PARTITIONS; only JCR IF(2024) is a real Impact Factor.
rank:
# Factory default reporting standard. The skill labels every paper with all
Expand All @@ -98,12 +98,9 @@ rank:
default_platform: jcr
# Where fetched ranking CSVs are cached (never the repo).
cache_dir: ~/.paper-search-pro/ranks
# Mirror URLs (swap to another mirror if one moves). All GitHub-raw -> plain
# HTTP works (no Cloudflare). Leave unset to use the built-in defaults.
sources:
cas: "https://raw.githubusercontent.com/hitfyd/ShowJCR/master/%E4%B8%AD%E7%A7%91%E9%99%A2%E5%88%86%E5%8C%BA%E8%A1%A8%E5%8F%8AJCR%E5%8E%9F%E5%A7%8B%E6%95%B0%E6%8D%AE%E6%96%87%E4%BB%B6/FQBJCR2025-UTF8.csv"
jcr: "https://raw.githubusercontent.com/hitfyd/ShowJCR/master/%E4%B8%AD%E7%A7%91%E9%99%A2%E5%88%86%E5%8C%BA%E8%A1%A8%E5%8F%8AJCR%E5%8E%9F%E5%A7%8B%E6%95%B0%E6%8D%AE%E6%96%87%E4%BB%B6/JCR2024-UTF8.csv"
sjr: "https://raw.githubusercontent.com/saramabrouk173/zotero-sjr-ranker/main/scimagojr%202024.csv"
# Downloads are opt-in. Add only URLs you are authorised to use.
# With no sources, cached CSV lookup still works and fetch performs no request.
sources: {}
# NOTE: no persistent tier default — tier filtering is per-request / per-explicit
# default only (never auto-persisted from a single request).

Expand Down
5 changes: 5 additions & 0 deletions pytest.ini
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
[pytest]
addopts = -m "not live"
markers =
live: requires external network access and may require user credentials
testpaths = tests
13 changes: 7 additions & 6 deletions references/agent_mode.md
Original file line number Diff line number Diff line change
Expand Up @@ -370,10 +370,10 @@ Process **exit code mirrors `error.code`** so a non-LLM caller can branch on `$?
| `--with-nssd` | flag / off | **Supplemental (zh) axis.** Also query NSSD (国家哲社文献中心, Chinese social sciences) and MERGE its results into the candidate pool. Requires zh in scope (an `en` scope is upgraded to `both`). Explicit-only — no discipline auto-routing. Reported in `meta.language.chinese_sources_used`. |
| `--with-yiigle` | flag / off | **Supplemental (zh) axis.** Also query yiigle (中华医学期刊全文数据库, Chinese medicine) and MERGE its results. Same zh-scope requirement + reporting as `--with-nssd`. |
| `--verify` | flag / off | Attach per-paper existence + abstract + cross-source consistency markers (see below). |
| `--quartile Q1,Q2` | csv / none | OPT-IN filter on `journal_rank.sjr.best_quartile` (the single layer); keep only these SJR quartiles. The partition is always attached. Needs cached journal-rank data (`journal_rank fetch` → `ranks/`). SJR分区, **NOT** JCR (R-04). For 区/category control use `--rank-platform`. |
| `--quartile Q1,Q2` | csv / none | OPT-IN filter on `journal_rank.sjr.best_quartile` (the single layer); keep only these SJR quartiles. The partition is always attached. Needs compatible data in `rank.cache_dir`; configure authorised `rank.sources` before fetch. SJR分区, **NOT** JCR (R-04). For 区/category control use `--rank-platform`. |
| `--min-impact X` | float / none | Drop papers whose OPEN journal impact (`journal_rank.openalex.mean_citedness_2yr`, OpenAlex 2yr mean citedness) is below `X`. **NOT** the JCR Impact Factor (R-09); relative use only. |
| `--journal-category NAME` | str / none | **DEPRECATED** (legacy SJR-only layer). Use `--rank-platform sjr --rank-category` to pin a per-category quartile in the single layer. |
| `--sjr-csv PATH` | str / none | **DEPRECATED / ignored** — the legacy sjr_helper SJR-only path is retired. Cache SJR data with `journal_rank fetch --platform sjr` (→ `ranks/`) instead. |
| `--sjr-csv PATH` | str / none | **DEPRECATED / ignored** — the legacy sjr_helper SJR-only path is retired. Place a compatible SJR CSV in `rank.cache_dir`, or configure an authorised source before `journal_rank fetch --platform sjr`. |
| `--no-journal-metric` | flag / off | Skip journal-rank enrichment entirely (`journal_rank` stays `null`: no partition labels, no OpenAlex open impact). |
| `--rank-platform {cas,jcr,sjr}` | choice / none | **Multi-platform layer** (the A-line partition feature). The platform to FILTER on this run: `cas` (中科院 区) / `jcr` / `sjr`. Omit to take the platform from the query's NL intent (`中科院一区` → cas), else config `rank.default_platform` (which only LABELS, never filters). CAS 区 & SJR quartile are PARTITIONS; only JCR is a real IF (R-04). |
| `--keep-tiers 1,2` | csv / none | **Multi-platform layer.** Tiers/quartiles to KEEP, mapped onto the chosen platform: CAS uses 区 numbers (`1,2`); JCR/SJR use quartiles (`Q1,Q2`). OPT-IN: with none given every paper is labelled with all three platforms but nothing is filtered. With a tier filter the search ADAPTIVELY DEEPENS to meet the target count. |
Expand Down Expand Up @@ -573,10 +573,11 @@ with **all three** platforms (plus an `openalex` open-impact sub-slot) and, when
tier was requested, filters on **one**. The opt-in `--quartile` / `--min-impact`
filters also read this record (`journal_rank.sjr.best_quartile` /
`journal_rank.openalex.mean_citedness_2yr`); there is no separate per-paper
`journal_metric` any more. The data is fetched at runtime from public mirrors into the user's local
cache (`~/.paper-search-pro/ranks/`) by `scripts/journal_rank.py` — **never bundled
in the repo**. First use needs a one-time `journal_rank fetch` (see
`references/journal_metrics.md`); until then this layer degrades gracefully
`journal_metric` any more. Data is read from the user's local cache
(`~/.paper-search-pro/ranks/`) by `scripts/journal_rank.py` — **never bundled in
the repo**. The user can place compatible CSVs there, or explicitly configure
authorised source URLs before `journal_rank fetch` (see
`references/journal_metrics.md`); until data exists this layer degrades gracefully
(`meta.rank.data_loaded: false`, slots stay `null`, the run still succeeds — it
**never** fetches inline).

Expand Down
Loading