This repository contains the planning and implementation blueprint for a standalone web-facing AI tagging and research system.
- Frontend: Next.js, TypeScript, Tailwind CSS
- Backend: FastAPI, Python, Pydantic
- Database: PostgreSQL + pgvector
- Queue: Redis
- Worker: Python worker service
- Storage: S3-compatible object storage, MinIO locally
- LLM: Centralized LLM gateway
- Local runtime: Docker Compose
- Planning and execution: Jira tickets + GitHub branches/PRs
- Upload individual files.
- Upload folders recursively from browser.
- Preserve relative paths during folder upload.
- Upload ontology CSV files.
- Upload multiple entity CSV files.
- Parse, chunk, and embed documents.
- Detect entities from uploaded entity files.
- Use entities and ontology terms for candidate tag selection.
- Run LLM-assisted structured tagging.
- Review, approve, reject, or edit tag candidates.
- Query tagged documents through a persistent research assistant.
- Maintain persistent project memory so agents can resume after chat resets.
- Docker Desktop (with Docker Compose)
# 1. Clone the repository
git clone https://github.com/eogbadu/ai-tagging-system.git
cd ai-tagging-system
# 2. Copy the environment example
cp .env.example .env
# 3. Build and start all services
docker compose up --build| Service | URL | Notes |
|---|---|---|
| Frontend | http://localhost:3000 | Next.js app |
| Backend API | http://localhost:8000/api/health | FastAPI health check |
| API Docs | http://localhost:8000/docs | Swagger UI |
| PostgreSQL | localhost:5432 | pgvector enabled (pg15) |
| Redis | localhost:6379 | |
| MinIO API | http://localhost:9000 | S3-compatible object storage |
| MinIO Console | http://localhost:9001 | Web UI (admin/minioadmin) |
curl http://localhost:8000/api/health
# Expected: {"status":"ok","service":"backend","version":"0.1.0"}docker compose down
# To remove volumes (database data):
docker compose down -vCLAUDE.mdAGENTS.mddocs/implementation_blueprint.mddocs/workflows/jira_github_workflow.mddocs/workflows/memory_compaction_policy.mddocs/project_memory/context_summary.mdspecs/database_schema.mdspecs/api_contract.mdspecs/frontend_spec.mdspecs/tagging_pipeline_spec.mdspecs/review_workflow_spec.mdspecs/research_assistant_spec.mdtickets/implementation_sequence.mdtickets/jira_ticket_template.md
The repo includes generated architecture assets under architecture/images/:
architecture/images/high_level_architecture.pngarchitecture/images/tagging_pipeline_workflow.pngarchitecture/images/deployment_container_architecture.png
See also:
architecture/diagrams.md
The architecture and specs are written first. Claude/Codex should implement against these files. Human review should constrain, test, and correct implementation decisions.
Development should happen in small Jira-ticket-sized increments:
- Select a Jira ticket.
- Create a GitHub branch with the ticket key.
- Implement only the ticket scope.
- Run tests/verification.
- Update project-memory files.
- Open a PR using the repo template.
The repo contains persistent memory files under:
docs/project_memory/
These files must be updated after meaningful changes. context_summary.md must be compacted and rewritten whenever project memory changes so future agents can resume without long chat history.