Llestrade is a PySide6 (Qt) desktop application for analyzing and summarizing complex and variable documents using multiple LLM providers (Anthropic Claude, Anthropic Claude via AWS Bedrock, Google Gemini, Azure OpenAI).
- Multiple LLM Providers: Support for Anthropic Claude (cloud & AWS Bedrock), Google Gemini, and Azure OpenAI GPT-4
- Document Processing: Convert PDFs to markdown and analyze forensic psychological reports
- Smart Chunking: Markdown-aware document chunking for large files
- Batch Processing: Process multiple documents with progress tracking
- Integrated Analysis: Combine multiple reports into comprehensive bulk analysis outputs
- Extended Thinking: Support for Claude and Gemini's advanced reasoning capabilities
- Debug Dashboard: Real-time monitoring and debugging tools (in debug mode)
- Error Recovery: Robust error handling with retry logic and crash recovery
- Python 3.12 or higher
- uv package manager (recommended) or pip
-
Clone the repository
git clone https://github.com/phren0logy/llestrade.git cd llestrade -
Install dependencies using uv
uv sync
-
Configure API keys
Use the Settings panel to add Azure DI, Azure OpenAI, and Anthropic credentials.
-
Run the application
uv run main.py # Or with debug instrumentation enabled ./run_debug.sh
For first-time users, run the interactive setup:
uv run scripts/setup_env.pyThis will:
- Create your
.envfile - Guide you through API key configuration
- Test your LLM connections
- Run a sample analysis
-
Create or open a project
- Use the welcome screen buttons to create a new project or open an existing
.frpdworkspace. - The project dialog prompts for a project name, project folder, source folder, and conversion helper.
- Pick a placeholder set (or start empty) so the workspace can substitute client-specific values in prompts.
- Use the welcome screen buttons to create a new project or open an existing
-
Documents tab
- The Documents tab shows the selected source folders and FileTracker counts (
Converted X / Y). - Use "Re-scan for new files" after dropping additional PDFs/markdown into the source tree.
- Start conversions from the banner; the conversion worker preserves folder structure inside
converted_documents/.
- The Documents tab shows the selected source folders and FileTracker counts (
-
Highlights tab
- Extract PDF highlights once conversions are complete. Only PDF-derived markdown contributes to the highlight denominator (
Highlights: X of Y). - Placeholder
.highlights.mdfiles are written when a PDF has no highlights so counts remain accurate.
- Extract PDF highlights once conversions are complete. Only PDF-derived markdown contributes to the highlight denominator (
-
Bulk Analysis and Reports tabs
- Create bulk analysis groups to run prompts per folder subset. Each group stores its config in
bulk_analysis/<group>/config.jsonand writes outputs underbulk_analysis/<group>/outputs/. - Use the inline log feed to monitor worker progress and rerun pending or all documents as needed.
- The Reports tab orchestrates draft/refinement prompts with the same placeholder map used by bulk jobs.
- Create bulk analysis groups to run prompts per folder subset. Each group stores its config in
llestrade/
├── main.py # Application entry point (launches src.app)
├── src/
│ ├── app/
│ │ ├── __init__.py # Re-exports project entry points (run, ProjectManager, etc.)
│ │ ├── core/ # Dashboard domain logic (project manager, file tracker, metrics)
│ │ ├── ui/
│ │ │ ├── dialogs/ # Qt dialogs
│ │ │ ├── stages/ # Top-level Qt widgets (main window, workspace shell)
│ │ │ ├── workspace/ # Decomposed workspace tabs
│ │ │ │ ├── controllers/
│ │ │ │ ├── services/
│ │ │ │ ├── bulk_tab.py / highlights_tab.py / reports_tab.py
│ │ │ └── widgets/ # Shared UI components
│ │ ├── workers/ # QRunnable-based background jobs
│ │ └── resources/
│ │ ├── prompts/
│ │ └── templates/
│ ├── config/ # Application configuration modules
│ │ ├── app_config.py
│ │ ├── config.py
│ │ ├── logging_config.py
│ │ └── startup_config.py
│ ├── core/ # Shared utilities reused by the dashboard
│ │ ├── exception_handler.py
│ │ ├── file_utils.py
│ │ ├── ingest_markdown.py
│ │ ├── pdf_utils.py
│ │ └── prompt_manager.py
│ └── common/llm/ # LLM provider abstractions and helpers
│ ├── base.py
│ ├── providers/
│ ├── chunking.py
│ ├── tokens.py
│ └── factory.py
├── tests/ # Test suite
├── scripts/ # Utility scripts
├── var/ # Runtime artefacts (gitignored contents)
│ ├── logs/
│ └── test_output/
When you create a project, the application maintains a self-contained workspace with derived outputs. The key folders are:
<project>/
├── project.frpd # Project metadata (source config, helper, UI state)
├── sources.json # Included folders (+ warnings for root files)
├── .llestrade/
│ └── citations.db # Canonical citation/evidence store (SQLite)
├── converted_documents/ # Markdown outputs mirroring selected folder structure
│ ├── medical_records/
│ │ ├── report1.md
│ │ └── report2.md
│ └── legal_docs/
│ └── case_summary.md
├── highlights/ # Highlight outputs for PDFs (mirrors converted_documents)
│ ├── medical_records/
│ │ ├── report1.highlights.md
│ │ └── report2.highlights.md
│ └── legal_docs/
│ └── case_summary.highlights.md
├── bulk_analysis/
│ ├── clinical_records/
│ │ ├── config.json # Prompts, model, folder subset
│ │ └── outputs/
│ │ ├── medical_records/
│ │ │ ├── report1.md
│ │ │ └── report2.md
│ │ └── legal_docs/
│ │ └── case_summary.md
│ └── legal_documents/
│ ├── config.json
│ └── outputs/
│ └── legal_docs/
│ └── case_summary.md
└── backups/
└── 2025-01-01T120000Z/ # Snapshot copies created by the app
Notes:
- Highlights are extracted only for PDFs. If a PDF has no highlights, a placeholder
.highlights.mdfile is created with a processed timestamp. - Dashboard highlight counts use a PDF-only denominator (e.g.,
Highlights: X of YwhereYis the number of PDF-converted documents), so DOCX and other non-PDF sources are excluded from the “pending highlights” count. - Bulk analysis and report prompts now substitute project placeholders (client, case, project name) along with per-document metadata. Ensure required placeholders are filled in Project Settings before running jobs.
- Citation-aware outputs can include inline markers like
[CIT:ev_<id>]. These IDs resolve through.llestrade/citations.db. - Use
uv run scripts/export_citations.py <project_dir>to export citation tables to JSON for debugging.
All markdown artefacts generated by Llestrade include a YAML front-matter block that captures provenance and runtime metadata. This is handled centrally by src/common/markdown/frontmatter_utils.py using the python-frontmatter library, so every worker shares the same structure.
Each document records:
project_path: Absolute path to the project workspace that produced the file.created_at: ISO 8601 timestamp (UTC) for when the markdown was written.generator: Identifier for the component that generated the file (conversion_worker,highlight_extraction,bulk_analysis_worker,bulk_reduce_worker,report_worker, etc.).sources: List of inputs (absolute path, project-relative path, file kind, role, and checksum).prompts: Prompt files or template IDs that influenced the output (role-labelled).- Additional keys specific to the workflow (for example
converter,pages_detected,highlight_count,prompt_hash,document_type, orrefinement_tokens).
Example front matter from a converted PDF:
---
project_path: /Users/me/Documents/cases/case-a
created_at: 2025-01-04T19:22:18.304218+00:00
generator: conversion_worker
sources:
- path: /Users/me/Documents/cases/case-a/sources/report.pdf
relative: sources/report.pdf
kind: pdf
role: primary
checksum: 2d5df4…
converter: pdf-local
pages_detected: 14
pages_pdf: 14
---This metadata is consumed by downstream tooling (dashboards, manifests, audits) and provides a consistent way to trace where every markdown document came from. When extending the app, prefer augmenting the front matter via the helper rather than writing YAML by hand.
Prompt templates use {placeholder} tokens that the workers populate at runtime. Placeholder requirements are defined in src/app/core/prompt_placeholders.py and validated whenever a prompt is loaded, so missing required tokens fail fast instead of producing malformed requests. The same registry powers the UI tooltips on prompt selectors.
| Prompt | Template file | Required placeholders | Optional placeholders |
|---|---|---|---|
| Document analysis (system) | prompts/document_analysis_system_prompt.md |
None | {subject_name}, {subject_dob}, {case_info} |
| Document bulk analysis (user) | prompts/document_bulk_analysis_prompt.md |
{document_content} |
{subject_name}, {subject_dob}, {case_info}, {document_name}, {chunk_index}, {chunk_total} |
| Integrated analysis | prompts/integrated_analysis_prompt.md |
{document_content} |
{subject_name}, {subject_dob}, {case_info} |
| Report generation (user) | prompts/report_generation_user_prompt.md |
{template_section}, {transcript}, {additional_documents} |
{section_title}, {document_content} |
| Report refinement (user) | prompts/refinement_prompt.md |
{draft_report}, {template} |
{transcript} |
| Report instructions | prompts/report_generation_instructions.md |
{template_section}, {transcript} |
None |
| Report generation (system) | prompts/report_generation_system_prompt.md |
None | None |
| Report refinement (system) | prompts/report_refinement_system_prompt.md |
None | None |
Bulk and report workers automatically inject additional runtime placeholders:
{source_pdf_filename},{source_pdf_relative_path},{source_pdf_absolute_path},{source_pdf_absolute_url}for each document derived from a PDF- Combined runs expose
{reduce_source_list},{reduce_source_table},{reduce_source_count}summarising aggregated inputs {project_name}and{timestamp}resolve at execution time
When building custom prompts, include every required placeholder shown above. Optional placeholders are always supplied (with an empty string if the value is unavailable), so they can be added or removed without breaking validation. If you introduce a new prompt template, add its specification to the registry so documentation and tooltips stay aligned.
The application stores settings in var/app_settings.json (created on first run):
{
"selected_llm_provider_id": "anthropic",
"llm_provider_configs": {
"anthropic": {
"enabled": true,
"default_model": "claude-sonnet-4-5-20250929"
},
"anthropic_bedrock": {
"enabled": true,
"default_model": "anthropic.claude-sonnet-4-5-v1"
},
"gemini": {
"enabled": true,
"default_model": "gemini-1.5-pro"
},
"azure_openai": {
"enabled": true,
"default_deployment_name": "gpt-4"
}
}
}Report draft/refinement and bulk map/reduce use a Gateway-backed execution backend by default.
- Required auth for Gateway mode:
PYDANTIC_AI_GATEWAY_API_KEY=<your key>(orPAIG_API_KEY)
- Managed Gateway:
- no additional endpoint configuration required
- Optional custom endpoint (self-hosted Gateway):
PYDANTIC_AI_GATEWAY_BASE_URL=<gateway base url>(orPAIG_BASE_URL)PYDANTIC_AI_GATEWAY_ROUTE=<route override>
- Optional shared concurrency cap for Gateway-backed worker requests:
PYDANTIC_AI_GATEWAY_MAX_CONCURRENCY=4- set
0to disable model-level concurrency limiting
- Fallback switch (force legacy native provider clients):
FRD_ENABLE_PYDANTIC_AI_GATEWAY=false
Gateway backend fallback behavior:
- Unsupported provider IDs automatically use the legacy backend implementation.
- Runtime Gateway errors are surfaced to workers with existing stage-level failure handling.
Self-hosting notes:
- Use a custom domain such as
https://gateway.example.comforPYDANTIC_AI_GATEWAY_BASE_URL. - Keep the runtime gateway on API-key auth; the current desktop app does not participate in Cloudflare Access login flows.
- For AWS Bedrock through the self-hosted gateway, store a Worker secret named
AWS_BEARER_TOKEN_BEDROCKand set the rendered Worker varAWS_DEFAULT_REGION. - The full Cloudflare/1Password operator workflow lives in
gateway/README.md.
Claude models delivered through AWS Bedrock rely on the AWS CLI credential chain. Run aws configure (for long-term access keys) or aws configure sso (for IAM Identity Center) so credentials are written to ~/.aws/credentials and ~/.aws/config. Llestrade reads those settings automatically; no AWS secrets are stored in the application. Optional overrides for profile, region, and the default Bedrock Claude model can be set under Settings → Configure API Keys → AWS Bedrock (Claude).
Enable debug mode for enhanced logging and monitoring:
# Via command line
uv run main.py --debug
# Via environment variable
DEBUG=true uv run main.pyDebug mode features:
- Debug Dashboard with real-time monitoring
- Detailed logging to
~/Documents/llestrade/logs/ - System resource tracking
- Operation timing and performance metrics
For complex analysis requiring step-by-step reasoning:
- Anthropic Claude: Automatically uses thinking mode for integrated analysis
- Google Gemini: Uses extended thinking API when available
The application handles large documents through:
- Smart chunking with configurable overlap
- Token counting with caching
- Progress tracking for long operations
- Memory-efficient processing
Built-in resilience features:
- Automatic retry with exponential backoff
- Crash recovery on startup
- Detailed error logging
- Transaction-safe file operations
-
"Module not found" errors
# Ensure dependencies are installed uv sync -
API Connection Issues
# Test your API connections uv run scripts/setup_env.py -
Large Document Timeouts
- Increase timeout in settings
- Check token limits for your model
- Enable debug mode to see detailed progress
-
Qt Plugin Issues (macOS)
- The application automatically handles Qt plugin paths
- If issues persist, check
QT_PLUGIN_PATHenvironment variable
scripts/setup_env.py: Interactive environment setup and testingtests/test_api_keys.py: Verify API key configurationtests/test_large_document_processing.py: Test large document handling
Logs are stored in:
- macOS/Linux:
~/Documents/llestrade/logs/ - Windows:
%USERPROFILE%\\Documents\\llestrade\\logs\\
Crash reports are saved to:
~/Documents/llestrade/crashes/
# Run the default suite (includes all markers unless filtered)
scripts/run_pytest.sh tests/
# Run deterministic PR-like suite (recommended for daily development)
scripts/run_pytest_pr.sh
# Run optional live-provider suite (requires provider API keys)
scripts/run_pytest_live.sh
# Override the env file used for live-provider runs
LLESTRADE_LIVE_ENV_FILE=.env.live scripts/run_pytest_live.sh
# Run a specific test file
scripts/run_pytest.sh tests/app/ui/test_file_tracker.py -vIf .env.live exists, scripts/run_pytest_live.sh runs pytest through
op run --env-file=.env.live. That file may contain plain values or op://...
1Password references.
The project uses:
- Type hints throughout
- Qt signal/slot patterns
- Async operations in worker threads
- Comprehensive error handling
- Fork the repository
- Create a feature branch
- Make your changes
- Run tests
- Submit a pull request
Key dependencies:
- PySide6 (Qt for Python)
- anthropic (Claude API)
- google-genai (Gemini API)
- openai (Azure OpenAI)
- pypdf (PDF processing)
- pdfplumber (PDF text extraction)
- psutil (System monitoring)
- python-dotenv (Environment management)
See pyproject.toml for complete dependency list.
The MIT License (MIT)
Copyright © 2025 Andrew Nanton
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
For issues and feature requests, please use the GitHub issue tracker.
For detailed documentation, see the docs/ directory and CLAUDE.md for AI assistant guidance.
Llestrade currently ships a single dashboard workflow launched by uv run main.py (or uv run -m src.app).
The active workspace flow is:
- Welcome / project open-create
- Documents (source selection + conversion)
- Highlights (PDF-only extraction and tracking)
- Bulk Analysis (grouped per-document and combined runs)
- Reports (draft + refinement runs)
Note: there is no standalone Progress tab in the dashboard workflow. Bulk Analysis inline logs are the canonical activity feed.
Legacy transition notes and historical plans live under docs/archive/ and are not the source of truth for current behavior.
CLAUDE.md- AI assistant guidance and technical detailsdocs/current_behavior.md- concise current behavior baselinedocs/work_plan.md- active implementation prioritiesdocs/progress.md- dashboard-era milestone summarydocs/testing_strategy.md- marker taxonomy and CI test lanesdocs/placeholder_reference.md- placeholder behavior and usage