Skip to content

fix: license per-claim grounding refusal in the default report prompt - #1961

Closed
MrSampson wants to merge 140 commits into
assafelovic:masterfrom
MrSampson:fix/report-grounding-instruction
Closed

MrSampson wants to merge 140 commits into
assafelovic:masterfrom
MrSampson:fix/report-grounding-instruction

Conversation

@MrSampson

Copy link
Copy Markdown

The default report prompt has no instruction for what to do when a cited source doesn't support a specific claim -- see linked issue for a worked example where an ad hoc custom prompt produced correct refusal behavior the default prompt never licensed.

3mk4yl and others added 30 commits March 23, 2026 20:51
Re-signed branch history.

Previous subjects:
- Add Codex CLI plugin manifest
- Fix: add skills field to plugin manifest
- Fix: address code review feedback
- Add Codex plugin quality gate CI
- Remove CI workflow from plugin PR
- Remove CI workflow from plugin manifest
…ixes assafelovic#1676)

When a WebSocket request enabled MCP, run_agent() permanently mutated
os.environ["RETRIEVER"] and os.environ["MCP_STRATEGY"] for the lifetime
of the server process. Subsequent requests that did not enable MCP would
inherit the modified RETRIEVER and end up using MCPRetriever unexpectedly,
causing No MCP server configurations found errors.

The MCP retriever and strategy are already passed directly to GPTResearcher
via the mcp_configs and mcp_strategy constructor parameters, so the env
mutations were redundant and only caused side-effects.

Changes:
- Remove os.environ["RETRIEVER"] and os.environ["MCP_STRATEGY"] writes
  from websocket_manager.run_agent().
- Rewrite GPTResearcher._process_mcp_configs() to inject mcp directly
  into self.cfg.retrievers (a per-instance list) instead of os.environ.
  This is safe under concurrent async requests because each researcher
  instance owns its own cfg object.
OpenAlex (https://openalex.org) is an open catalog of ~300M
scholarly works. No API key is required. It complements the
existing semantic_scholar and pubmed_central retrievers for
academic queries.

Key behaviors:
- Reconstructs abstracts from OpenAlex's inverted-index format.
- Prefers the open-access PDF URL, falls back to the landing
  page, then the OpenAlex work URL, so non-OA works are not
  silently dropped.
- Reads optional OPENALEX_EMAIL (polite pool) and
  OPENALEX_API_KEY (authenticated pool) from env; works without
  either.

Closes assafelovic#1264
Large scraped documents (especially PDFs) can exceed embedding API token
limits when all chunks are embedded in a single batch request. For example,
OpenAI's text-embedding models have a 300k token-per-request cap, which
is easily breached when multiple large pages are split into many 1000-char
chunks by the compression pipeline.

Truncate raw_content to 50000 chars per document before creating LangChain
Document objects in SearchAPIRetriever. This keeps the total embedding batch
well within provider limits while preserving enough content for accurate
similarity filtering. The limit is configurable via the MAX_CONTENT_CHARS
environment variable.

Fixes assafelovic#1524
…ndpoints (fixes assafelovic#1525)

The mistralai provider did not honor MISTRAL_BASE_URL, causing users with
self-hosted or vLLM-based Mistral deployments to always hit the public
Mistral API and receive 401 errors. Add the same env-var pattern already
used by the openai provider (OPENAI_BASE_URL) so that setting
MISTRAL_BASE_URL overrides the endpoint passed to ChatMistralAI.

Co-Authored-By: Octopus <liyuan851277048@icloud.com>
Adds a CI workflow that runs tokentoll on pull requests to detect
changes in LLM API calls (model swaps, new call sites, removed
endpoints) and posts the cost impact as a PR comment.

The action is pinned to SHA 753ca4d1150c74169b52a843439049b65e256d2b (v0.6.1) and installs
tokentoll==v0.6.1 from PyPI. Zero runtime dependencies.
Signed-off-by: box4wangjing <box4wangjing@outlook.com>
Replaces literal-prefix string parsing (startswith('Query:'),
startswith('Question:'), startswith('Learning')) with JSON-strict
prompts and json_repair parsing, matching the approach used in
actions/query_processing.py. A regex fallback is kept for the
legacy output format so existing users on custom models are not
broken.

Validated end-to-end with Claude Haiku 4.5, which previously
triggered the silent failure mode. Tests cover clean JSON,
markdown-fenced JSON, malformed-but-repairable JSON, empty input,
and legacy prefix formats.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…t models

Modern models support 64k-128k output tokens (Claude Haiku 4.5 /
Sonnet 4.6 = 64k, Opus 4.7 / GPT-5 family = 128k). The hardcoded 32k
cap predates these models and silently corrupts long-form report
generation: the ValueError is swallowed by report_generation.py and an
empty report is returned without surfacing the cause.

The cap is kept as a sanity guard against absurd typos
(e.g., SMART_TOKEN_LIMIT=1000000), since provider-side error reporting
is inconsistent across backends (cloud APIs surface clear errors,
Ollama / litellm providers may not). Raised to 200k, leaving ~50%
headroom above the largest currently available cloud output limit.

The error message now points at the env vars users should check.

Refs: assafelovic#1335

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adds commented examples for FAST_TOKEN_LIMIT / SMART_TOKEN_LIMIT /
STRATEGIC_TOKEN_LIMIT with recommended values per model class.
Default behavior unchanged (variables stay unset -> code defaults apply).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ecommendations

- Update gptr/config.md defaults to match runtime values
  (FAST=3000, SMART=6000, STRATEGIC=4000)
- Add a recommended-values table for modern LLMs with large output
- Cross-link from llms.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… as full content

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…imports-and-tests

fix: properly resolve report_type imports in websocket_manager and re…
Fix: resolve unhashable dict error in detailed report context deduplication
…v-pollution-mcp-retriever

fix: eliminate process-level env pollution from MCP retriever setup
feat(cli): generate filename from LLM and add YAML frontmatter for re…
…very-route

feat(server): add Agent Discovery Protocol manifest endpoint
Bartok9 and others added 23 commits July 14, 2026 12:27
PyMuPDFScraper downloads a remote PDF to NamedTemporaryFile(delete=False)
and only called os.remove() on the success path. When PyMuPDFLoader.load()
raises on a malformed/partial PDF, the broad except swallowed the error and
left the temp .pdf on disk every time. Wrap the load in try/finally so the
temp file is always cleaned up.

(cherry picked from commit 124ec5c)
The loader_dict in _load_document keys formats in lower-case (pdf, docx,
xlsx, ...), but _get_extension returned the raw suffix. A URL ending in
an upper-case extension (e.g. report.PDF, doc.DOCX — common on direct
links and signed CDN URLs) produced an upper-case format string that
missed every loader_dict key, so the document was silently skipped.

Lower-case the extracted extension so the loader lookup matches.

(cherry picked from commit c67a179)
…yntax

The generate_search_queries_prompt() function said 'Write N google search
queries'. Instruction-following LLMs (e.g. GPT-5.4) interpret this literally
and inject Google-specific operator syntax such as site:arxiv.org,
filetype:pdf, inurl:, intitle:, OR, AND into the generated queries.

Most search backends (Tavily, DuckDuckGo, Semantic Scholar, Bing, Exa,
SerpAPI, etc.) do not support these operators and return zero results when
they appear. This silently kills sub-query result sets mid-research, causing
the final report to be based on far fewer sources than expected.

Changes:
- Remove the word 'google' from the prompt (was the trigger for operator injection)
- Add an explicit instruction not to use operator syntax and explain why
- Keep the fix backend-neutral so it is correct regardless of which
  retriever(s) are configured

(cherry picked from commit c15a863)
When CURATE_SOURCES=True, ResearchConductor.conduct_research() returns
result['context'] as a List[dict] (each dict has Title/Content/Source keys)
rather than a plain str. Two separate crash paths:

1. all_context.append(result['context']) nests the list as a single item,
   so context_with_citations ends up as a list containing List[dict] items.
   The subsequent '\n'.join(final_context) then raises:
     TypeError: sequence item N: expected str instance, dict found
   Fix: extend instead of append when context is a list.

2. As a safety net, the final join is made type-safe so any residual
   dict items are extracted via their 'Content' key rather than crashing.

Fixes assafelovic#1279, fixes assafelovic#1332

(cherry picked from commit 216899e)
…rcher.context

curate_sources() returns List[dict] (each dict contains Title, Content, Source
keys), but researcher.context is expected to be a str throughout the codebase.
Assigning the List[dict] directly causes any downstream str operation to crash:

  - len(researcher.context) returns the list length (not character count)
  - researcher.context.split() raises AttributeError: 'list' has no attribute 'split'
  - '\n'.join(...) or other string ops on the value fail

Fix: capture the return value, detect if it is a list, and format each dict
entry into a readable 'Title / Content / Source' block joined by double newlines.
Non-dict items fall back to str(). Non-list returns are passed through unchanged.

(cherry picked from commit f71b95b)
- Add MiniMax-M3 to model list and set as default
- Keep MiniMax-M2.7 and MiniMax-M2.7-highspeed
- Remove older models (M2.5/M2.5-highspeed)
- Update related documentation

(cherry picked from commit 18893c0)
Anthropic deprecates the temperature parameter on Claude Sonnet 4.5+,
Opus 4.5+, and Haiku 4.5+. Passing temperature to these models now
returns a 400 error: "'temperature' is deprecated for this model".

This breaks the report-writer step in gpt-researcher when configured
with anthropic:claude-sonnet-4-5 or anthropic:claude-opus-4-7 (the
default-recommended modern models).

Adding these to NO_SUPPORT_TEMPERATURE_MODELS routes them through the
existing temperature=None code path that already handles GPT-5 and
the o1/o3/o4 reasoning models.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
(cherry picked from commit e199a5b)
Resolve post-merge conflicts: 4 retriever/context fixes (supersedes assafelovic#1895-1898)
…guards

Retriever hardening: 20 guard fixes against malformed results (supersedes assafelovic#1837…assafelovic#1890)
…bustness

Scraper robustness: 6 fixes (title/PDF-detection/temp-file/dimensions) (supersedes assafelovic#1824…assafelovic#1842)
…s-bounds

Multi-agent robustness: bound revision loops + exact sentinels (supersedes assafelovic#1883/assafelovic#1885/assafelovic#1886)
Core/misc hardening: 11 fixes (costs/query/llm/agent/mcp/config) (supersedes assafelovic#1822…assafelovic#1902)
Reconcile master→main: GetXAPI retriever, Claude 4.x temp fix, MiniMax M3, curate_sources fixes
…actory-provider

feat: add Nebius Token Factory as LLM and embedding provider
Move the standalone multi_agents_ag2/ directory to multi_agents/ag2/ so the
AG2 (AutoGen) variant lives under the multi_agents package instead of a
separate top-level directory. The AG2 code already depended on multi_agents
(reusing its agents + utils), so this reflects the real relationship.

Updated all references: module path (multi_agents_ag2 -> multi_agents.ag2)
in backend/server/multi_agent_runner.py and ag2/main.py, plus doc/README paths.
History preserved via git mv; primary multi_agents package and the runner
fallback both verified importable, tests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…nder-multi-agents

chore(multi_agents): nest AG2 variant under multi_agents/ag2
Audited the Claude skill references against current code and fixed verified
drift (skill itself + other references were accurate):

- retrievers.md: list all 21 retrievers (was 13) incl. Brave/SearX/GroundRoute/
  fastCRW/BoCha/Xquik/OpenAlex/GetXAPI; fix get_retrievers() signature to the
  real (headers, cfg) form
- multi-agents.md: ChiefEditorAgent is in agents/orchestrator.py (not editor.py);
  Reviser/agents/reviser.py (not 'Revisor'/revisor.py); import from
  multi_agents.main (not multi_agents)
- api-reference.md: /report response nests research_information.research_costs
  (no top-level 'costs'); DELETE /files/{filename} (not /delete/); add GET /files/;
  remove non-existent /getConfig & /setConfig endpoints

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ence-refresh

docs(skill): fix drift in .claude reference docs (retrievers, multi-agents, api)
Ports the fix from assafelovic#1943
onto this production branch. gpt_researcher/actions/query_processing.py
failed to import at all on the currently-published 0.16.0 release: a
typing import sat after its first use, evaluated eagerly at module load
time. Moves all imports to the top in conventional order.
… get scraped (assafelovic#17)

Native port of a long-standing internal monkeypatch for this fix
(previously applied externally as a patch, now maintained as
fork-native code here).

gpt_researcher.skills.researcher._search_relevant_source_urls() treats
any search result whose raw_content/body exceeds 100 characters as
already-fetched full text -- a heuristic meant for retrievers (e.g.
PubMed Central) that genuinely return full article text inline.
SearxNG and ddgs both populate "body" with an ordinary search-result
snippet, which routinely exceeds 100 characters, so every result from
either retriever was wrongly treated as already-fetched and the real
page was never actually scraped -- get_source_urls() always returned
[], and report citations weren't grounded in anything fetched.

Capped at the retriever source rather than patched in
_search_relevant_source_urls() itself, so any future upstream change
to that function's classification logic is inherited automatically
when this fork is rebased. Confirmed working in production since the
original patch shipped: retrieval quality improved markedly.
… PDFs (assafelovic#20, assafelovic#24)

Consolidates two separate upstream PRs onto this production branch,
combined into extract_data_from_url with a deliberate check order:

1. assafelovic#1951 (assafelovic#24) --
   PDFs on extensionless URLs (institutional repository download
   endpoints) get routed to a text/HTML backend and ingested as raw
   binary. Detected via PDF structural tokens and retried with
   PyMuPDFScraper. Runs FIRST, before the content-quality checks below,
   so a misdispatched PDF gets a chance to be recovered as real content
   before anything downstream might otherwise discard it.
2. assafelovic#1944 (assafelovic#20) --
   anti-bot/challenge pages (Anubis, Cloudflare, ResearchGate) and
   word-list/vocab dumps are ingested as real content. Detected via
   anchored markers (block pages) and sentence-ending-punctuation
   density including CJK terminators (word lists), and rejected in the
   same shape the function already returns for too-short content.

Both PRs' full rationale, false-positive analysis, and real-world
verification (including reproducing the exact reported PDF URL and
confirming clean text recovery) are in their respective PR
descriptions. Also extracts the shared reject-and-log shape into a
_reject() helper and removes a pre-existing unreachable duplicate
too-short check while touching this method.
generate_report_prompt (used by deep_research's default write_report()
call, no custom_prompt) repeatedly demands maximal comprehensiveness
and has exactly one anti-fabrication guard -- 'do not cite sources
absent from context' -- which says nothing about a present source
that doesn't actually support the specific claim next to its
citation. Real-world evidence demonstrated the model is capable of the
correct behavior (an ad hoc custom_prompt asking it to flag
unreadable/unsupported content produced a grounded, partially-
refusing report on the same context that the default prompt
fabricated a confidently cited section from) -- the default path
just never told it refusal/hedging was an acceptable output.

Adds an explicit per-claim grounding guard, modeled on
generate_quick_summary_prompt's existing 'if the results are
insufficient to answer the query, state that clearly' -- the same
pattern already proven out for a different tool (quick_search),
now applied at the granularity report synthesis actually needs
(per-claim, not whole-report).
@assafelovic

Copy link
Copy Markdown
Owner

Reviewed — the actual change here is small and worth having, but it's buried.

GitHub reports 27,040 additions across 311 files. Nearly all of that is a branch that picked up an unrelated working tree: .claude/, .codex-plugin/, .mcp.json, .vscode/, workflow files, and a .claude/worktrees/crazy-curie-e43899 gitlink — that last one was itself a bug on main (it broke recursive clones and git-URL Docker builds) and was removed in #2074.

The fix described in the title looks like roughly twenty lines. Rebasing onto main and keeping only that would make this reviewable and, most likely, mergeable.

Worth saying plainly: the content concerns are real and I'd like these to land. Your #1944, #1951 and #1952 all merged in #2074 and were among the highest-value fixes in the batch — the scraper quality work in particular. This is a packaging problem, not a judgement on the work.

Also: this targets master, which is no longer the trunk. main is the default branch and is now 160 commits ahead; master last moved in April. Retargeting to main is needed regardless of anything else.

CI now runs on every PR — imports and unit across Python 3.11–3.14, plus collect — so once it's rebased you'll get a direct answer.

@assafelovic assafelovic added the needs-rebase Sound work that no longer applies to main label Aug 23, 2026
@assafelovic

Copy link
Copy Markdown
Owner

Thanks for this! This PR targets master, which is retired — main is the canonical branch, and the diff here now carries hundreds of unrelated files from the branch divergence. We asked for a rebase on Aug 23 and there's been no update since, so I'm closing it.

If you'd like to continue, please open a fresh PR against main containing only your change and we'll review it promptly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-rebase Sound work that no longer applies to main

Projects

None yet

Development

Successfully merging this pull request may close these issues.