Skip to content

feat: tool capping, skills system, conversation summarization, config centralization - #8

Merged
AhmadHammad21 merged 7 commits into
mainfrom
feat/tool-response-capping
May 7, 2026
Merged

AhmadHammad21 merged 7 commits into
mainfrom
feat/tool-response-capping

Conversation

@AhmadHammad21

Copy link
Copy Markdown
Owner

Summary

  • Tool response capping — all tools wrapped with with_cap() at startup; responses over TOOL_RESPONSE_MAX_CHARS (default 40K chars) are truncated with a notice so the agent knows to narrow its query
  • Config centralization — all settings moved from src/agent/config.py to src/config/appsettings.py, re-exported via src/config/__init__.py; all 21 import sites updated
  • Skills system — auto-discovers src/skills/*/SKILL.md files; skill names injected into system prompt at startup; full content loaded on demand via use_skill(name); ships with lambda-throttling skill
  • Conversation summarization — before each agent call, if session exceeds SUMMARIZATION_THRESHOLD_CHARS, compacts old messages via LLM call, preserving last SUMMARIZATION_KEEP_CHARS of context; tracks events in usage_events.metadata
  • Dashboard summarization stats — /stats now returns total_summarizations and total_chars_compacted; dashboard UI shows a "Context management" row when non-zero
  • Docs — docs/tool_capping.md, docs/conversation_summarization.md, docs/skills.md

Migration required (Postgres only)

psql $DATABASE_URL -f migrations/003_usage_events_metadata.sql

Adds metadata jsonb column to usage_events. SQLite auto-migrates on startup.

New env vars

TOOL_RESPONSE_MAX_CHARS=40000
SUMMARIZATION_ENABLED=true
SUMMARIZATION_THRESHOLD_CHARS=60000
SUMMARIZATION_KEEP_CHARS=20000

Test plan

  • uv run pytest tests/ -q — 172 passed, 4 skipped
  • Set TOOL_RESPONSE_MAX_CHARS=2000, ask "list all Lambda functions", confirm agent sees _capped response
  • Set SUMMARIZATION_THRESHOLD_CHARS=500, send 3+ messages, confirm log line summarizer: session=... chars, removing N msgs
  • GET /stats — confirm total_summarizations and total_chars_compacted are non-zero after summarization fires
  • Dashboard UI — confirm "Context management" row appears with Sessions compacted + Context saved cards
  • python -c "from tools.skills import list_skills; print(list_skills())" — confirm lambda-throttling appears

Tool response capping:
- src/tools/_cap.py: cap_tool_result() truncates dicts that exceed
  TOOL_RESPONSE_MAX_CHARS (default 40 000 chars ≈ 10 K tokens) and
  appends a notice telling the agent to narrow its query
- src/agent/core.py: wrap ALL_TOOLS with with_cap() at agent init
- 9 unit tests in tests/test_tools/test_cap.py

Config centralization:
- src/config/appsettings.py: Settings class (moved from agent/config.py)
- src/config/__init__.py: re-exports Settings and settings
- Deleted src/agent/config.py; all 21 import sites updated to
  `from config import settings`
- src/tools/skills.py: auto-discovers src/skills/*/SKILL.md files,
  exposes list_skills() and use_skill(name) tools to the agent
- src/skills/lambda-throttling/SKILL.md: step-by-step throttling
  investigation (concurrency checks, CloudTrail changes, account limits,
  traffic spikes, mitigation table)
- src/agent/prompts.py: injects available skill names + descriptions into
  system prompt at startup; full content only loaded on use_skill() call
- src/config/appsettings.py: checkpoint_backend uses Literal type
- README: skills system marked done in mid-term roadmap
Automatically compacts long sessions before the LLM context window overflows.
Fires before each agent call when total message chars exceed
SUMMARIZATION_THRESHOLD_CHARS (default 60 000 chars ≈ 15 K tokens).

- src/agent/summarizer.py: maybe_summarize() reads LangGraph state,
  splits at HumanMessage boundaries (no orphaned ToolMessages), calls LLM
  to produce a structured summary, applies RemoveMessage + injects summary
  HumanMessage via aupdate_state()
- src/api/routers/chat.py: single await maybe_summarize() call before astream()
- src/config/appsettings.py: summarization_enabled, threshold, keep_chars settings
- save_usage_event() gains metadata: dict param across all backends so
  summarization events are recorded with {"summarization": True,
  "messages_removed": N, "chars_removed": N} for dashboard queries
- migrations/003: ADD COLUMN metadata jsonb to usage_events (postgres)
- SQLite DDL + inline ALTER TABLE migration for existing databases
- dashboard /stats now includes total_summarizations and total_chars_compacted
  (queried from usage_events.metadata across all three backends)
- add docs/tool_capping.md and docs/conversation_summarization.md
- update _SUMMARY_KEYS in test_backends to reflect new fields
Add context management stat row (hidden until first summarization fires):
- Sessions compacted: total summarization runs
- Context saved: estimated tokens (~chars/4) compacted across all sessions
@AhmadHammad21
AhmadHammad21 merged commit b9f8045 into main May 7, 2026
2 checks passed
@AhmadHammad21
AhmadHammad21 deleted the feat/tool-response-capping branch May 7, 2026 18:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant