This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Search Agent — a FastAPI service that implements a 3-stage web search pipeline using Pydantic AI agents and SearXNG as the search backend. Also exposes search as an MCP tool.
All Python commands run via docker compose (never directly on host). The Taskfile uses docker compose exec (requires running services), so start services first.
task up # Start all services (required before other task commands)
task test # Run tests (uses exec)
task test -- tests/test_pipeline.py::TestSearchPipeline::test_pipeline_runs_all_stages # Single test
task lint:check # Lint (src/ only)
task lint:format # Format (src/ only)
task lint # Lint + format check
task coding-standards:apply # Lint fix + format
task build:image # Build and push prod image to ghcr.io/aarhusai/search-agent
task build:image TAG=v1.0.0 # With custom tagAlternative: direct docker compose (does not require running services):
docker compose run --rm --no-deps agent uv run pytest # Run all tests
docker compose run --rm --no-deps agent uv run ruff check src tests # Lint
docker compose run --rm --no-deps agent uv run ruff format src tests # Format- Query Planner — Decomposes complex questions into up to
search_max_queriestargeted search queries via LLM. Skipped for "simple" queries (word/?thresholds in settings, plus a hardcoded complexity regex) controlled by_is_simple_query(). - Search Executor — Calls the configured search provider (
search_provider:searxngdefault, orstaan) via HTTP, runs multiple queries concurrently, deduplicates by URL, caps atsearch_max_results. The Staan provider can return full page content / scored chunks per result directly intoRawSearchResult.content. - (Optional) Page Fetch — If
search_fetch_page_content=true,fetch.pyfetches the topsearch_fetch_max_pagesresult URLs and extracts main text viatrafilatura; the extracted text lands inRawSearchResult.contentfor the synthesizer. Guarded by content-type check, byte cap, and an SSRF filter that rejects private/loopback/link-local hosts. MCP path is snippet-only regardless. - Analyze + Synthesize — Single combined agent (
analyze_synthesizer) that filters, ranks, extracts passages, and generates a cited summary with[1],[2]style inline citations.
main.py— FastAPI app with/health,/api/v1/searchendpoints. Mounts the MCP ASGI app at/, which serves the Streamable HTTP transport at/mcp. The mount must stay at module scope:mcp.session_manager(used by the lifespan) raisesRuntimeErroruntilstreamable_http_app()has been called. Wraps the handler incache.bypass()whenSearchRequest.no_cache=True.pipeline.py— Orchestrates the 3 stages with timeout handling (default 90s). Planner output is cached per (normalized query, context,YYYY-MM-DD) — the date bucket is in the key because the prompt includesdatetime.now(), so a TTL crossing midnight would otherwise leak a stale date.config.py—Settingsclass using pydantic-settings. All env vars useSEARCH_AGENT_prefix.deps.py— Shared httpx client and Pydantic AI model, initialized at app startup via lifespan.cache.py— Pluggable cache with aCacheBackendprotocol and three implementations:RedisBackend(production, shared across pods),InMemoryBackend(tests/dev only — per-process, not safe for multi-pod prod),DisabledBackend. Fail-open on Redis errors/timeouts (300ms socket timeouts). Versioned namespace keys (fetch:v1:…).bypass()context manager +no_cacheContextVaravoid threading a flag through every call.providers/— Pluggable search backends behind aSearchProviderprotocol (base.py, mirrors theCacheBackendpattern).providers/__init__.pyis the registry:init_provider()(lifespan, fails fast on missing Staan key),get_provider()(lazy fallback for tests),set_provider_for_testing(). Shared helpers inbase.py:search_multiple(concurrency + URL dedup, then enforces the provider'scontent_result_capglobally across all queries so multi-query content can't overflow the LLM context window),is_valid_url,read_capped_json(rejects non-object JSON bodies),normalize_query.providers/searxng.py—SearxngProvider(default). Warns onunresponsive_engines(e.g. Brave rate-limited) and, at DEBUG, logs which engines contributed results. Results cached per(normalize(query), searxng_url); empty result lists are not cached so transient outages can retry immediately.providers/staan.py—StaanProviderfor the Staan "Web for AI" API (Bearer auth,GET /v2/search/web, results underweb.results). Enrichment viastaan_enrichment:full_content(markdown page body) orextra_snippets(scored chunks) →RawSearchResult.content, capped per result (staan_content_max_chars). The count cap (staan_content_max_resultstop reranked results keep content) is applied globally insearch_multipleacross all queries — not insidesearch()— so the cached payload is cap-independent and the synthesizer prompt stays bounded regardless of query count. Response read is capped atstaan_max_response_bytes(larger than SearXNG's sincefull_contentreturns whole page bodies). Never log its request headers — they carry the API key.
fetch.py— Optional per-result page fetch with trafilatura extraction. Concurrent, with content-type / byte-size / SSRF guards. Skips results whosecontentis already populated (e.g. by the Staan provider) — only content-less results consumesearch_fetch_max_pagesslots.trafilatura.extractis run viaasyncio.to_threadbecause it's sync._fetch_onewraps_fetch_one_uncachedwith a cache (positive TTL for extracted text, shorter negative TTL forNone/failed fetches).models.py—SearchRequest(with optionalno_cacheflag),RawSearchResult(with optionalcontentfield populated by fetch),SearchResult,Source.mcp_server.py— MCP server (MCPServerfrom the mcp 2.x SDK) exposingsearch_webtool. Usesrun_search_pipeline_raw(steps 1+2 only, no LLM synthesis) so callers get raw results for their own citation handling (e.g. Open WebUI). Does not exposeno_cache; MCP callers always hit the shared cache.streamable_http_app()builds the ASGI app thatmain.pymounts — mcp 2.x takestransport_security(themcp_allowed_hostsDNS-rebinding allowlist) on the transport factory rather than the server constructor, so it is applied there.agents/— Pydantic AI agent definitions.analyze_synthesizer.pyis the combined analyze+synthesize agent.
SearchRequest(query, context, no_cache) → query planner (Redis-cached) → [str] queries → search provider (Redis-cached; SearXNG or Staan) → [RawSearchResult] → fetch_pages (Redis-cached per URL, optional, skips results that already have content) → analyze_synthesizer → SearchResult(summary, sources)
All env vars use SEARCH_AGENT_ prefix (via pydantic-settings). Key settings:
SEARCH_AGENT_DEBUG(default:false) — setssearch_agentandpydantic_ailoggers to DEBUG and logs full agent prompts + outputs, plus the SearXNG engine list per query. Does not enablehttpx/httpcoreDEBUG (those leakAuthorizationheaders and full LLM response bodies); enable those manually if you need them.SEARCH_AGENT_LLM_BASE_URL(default:http://localhost:11434/v1) — OpenAI-compatible endpoint (Ollama, etc.)SEARCH_AGENT_LLM_API_KEY,SEARCH_AGENT_LLM_MODEL(default:llama3)SEARCH_AGENT_LLM_STRICT_TOOLS(default:true) — OpenAI strict tool definitionsSEARCH_AGENT_SEARCH_PROVIDER(default:searxng) — search backend:searxngorstaan. Deep health check (/health?deep=true) reportsprovider+search_backendfields.SEARCH_AGENT_SEARXNG_URL(default:http://searxng:8080)SEARCH_AGENT_STAAN_API_KEY— required when provider isstaan(startup fails without it). PlusSEARCH_AGENT_STAAN_URL(https://api.staan.ai),SEARCH_AGENT_STAAN_MARKET(en-us),SEARCH_AGENT_STAAN_TIMEOUT(10s),SEARCH_AGENT_STAAN_ENRICHMENT(full_content|extra_snippets|none),SEARCH_AGENT_STAAN_MAX_SNIPPETS(3),SEARCH_AGENT_STAAN_MIN_SCORE(0.1),SEARCH_AGENT_STAAN_CONTENT_MAX_CHARS(5000),SEARCH_AGENT_STAAN_CONTENT_MAX_RESULTS(5, top reranked results that keep content, enforced globally across all planner queries — with the char cap this bounds the synthesizer prompt for 32k-context models),SEARCH_AGENT_STAAN_MAX_RESPONSE_BYTES(10_000_000)SEARCH_AGENT_SEARXNG_TIMEOUT(15s),SEARCH_AGENT_SEARCH_PIPELINE_TIMEOUT(90s),SEARCH_AGENT_LLM_TIMEOUT(60s)SEARCH_AGENT_DATETIME_TIMEZONE(default:UTC),SEARCH_AGENT_DATETIME_FORMAT— used in query planner promptsSEARCH_AGENT_MCP_ALLOWED_HOSTS(default:["search-agent:8001","localhost:8001"]) — Host header allowlist for the MCP transport; a mismatchedHostgets a 421SEARCH_AGENT_SEARCH_SKIP_PLANNER_FOR_SIMPLE_QUERIES(default:true)SEARCH_AGENT_SEARCH_MAX_QUERIES(default:3) — cap planner outputSEARCH_AGENT_SEARCH_MAX_RESULTS(default:15) — cap results reaching the synthesizerSEARCH_AGENT_SEARCH_SIMPLE_QUERY_MAX_WORDS(default:15),SEARCH_AGENT_SEARCH_SIMPLE_QUERY_MAX_QUESTIONS(default:1) —_is_simple_querythresholdsSEARCH_AGENT_SEARCH_FETCH_PAGE_CONTENT(default:false) — fetch and extract main text from result pages via trafilatura; when enabled the synthesizer sees acontentfield per result in addition to the SearXNGsnippet. MCP path (search_web) is snippet-only regardless.SEARCH_AGENT_SEARCH_FETCH_MAX_PAGES(5),SEARCH_AGENT_SEARCH_FETCH_TIMEOUT(10s),SEARCH_AGENT_SEARCH_FETCH_MAX_CHARS(5000),SEARCH_AGENT_SEARCH_FETCH_MAX_BYTES(2_000_000)SEARCH_AGENT_CACHE_BACKEND(default:redis) —redisfor production (shared across pods),memoryfor tests/dev only (per-process), ordisabled.SEARCH_AGENT_CACHE_REDIS_URL(default:redis://redis:6379/0) — used when backend isredis.SEARCH_AGENT_CACHE_FETCH_TTL(3600s),SEARCH_AGENT_CACHE_FETCH_NEGATIVE_TTL(300s),SEARCH_AGENT_CACHE_SEARXNG_TTL(300s),SEARCH_AGENT_CACHE_STAAN_TTL(300s),SEARCH_AGENT_CACHE_PLANNER_TTL(21600s) — TTLs per cache namespace. Planner cache key is date-bucketed so it rolls daily even inside TTL.- Agent prompts overridable via
SEARCH_AGENT_SEARCH_QUERY_PLANNER_PROMPT,SEARCH_AGENT_SEARCH_ANALYZE_SYNTHESIZE_PROMPT
Tests mock all external services (LLM and SearXNG). conftest.py sets SEARCH_AGENT_* env vars before any module imports — this ordering matters because config.py reads env vars at module level via settings = Settings(). conftest.py also pins SEARCH_AGENT_CACHE_BACKEND=disabled so existing tests assert fresh behaviour; cache-specific tests opt in by swapping the module-level backend via cache.set_backend_for_testing(InMemoryBackend()) (see the in_memory_backend fixture in tests/test_cache.py). RedisBackend is exercised with fakeredis.
pytest-asyncio is configured with asyncio_mode = "auto" so async tests don't need the @pytest.mark.asyncio decorator.
- Python 3.12+, ruff with rules: E, F, I, N, W, UP, B, RUF
- Line length: 100
- Async-first — all I/O uses async httpx and FastAPI async handlers
Multi-stage Dockerfile with dev and prod targets. docker-compose defines three services: agent (container port 8001), searxng (container port 8080), and redis (container port 6379, redis:7-alpine with 256MB maxmemory + allkeys-lru, data persisted at .docker/data/redis/). All three have health checks — ports are not host-mapped (random host ports unless overridden). Services connect via a bridge app network; agent is also on an external frontend network. Source is volume-mounted for live reload in dev. Build target is controlled by ENV variable (defaults to dev). Note: Taskfile.yml currently sets SERVICE: search-agent, which no longer matches the compose service name agent — task commands that use docker compose exec {{.SERVICE}} will fail until that var is updated. Use docker compose exec agent … directly in the meantime.
- Never read
.envfiles — they contain secrets. Usedocker-compose.ymlorconfig.pyto understand env vars. - Always use docker compose to run Python commands — never run
uvor Python directly on the host.