Repository navigation
feat: tool capping, skills system, conversation summarization, config centralization - #8
Merged
Merged
Conversation
Tool response capping: - src/tools/_cap.py: cap_tool_result() truncates dicts that exceed TOOL_RESPONSE_MAX_CHARS (default 40 000 chars ≈ 10 K tokens) and appends a notice telling the agent to narrow its query - src/agent/core.py: wrap ALL_TOOLS with with_cap() at agent init - 9 unit tests in tests/test_tools/test_cap.py Config centralization: - src/config/appsettings.py: Settings class (moved from agent/config.py) - src/config/__init__.py: re-exports Settings and settings - Deleted src/agent/config.py; all 21 import sites updated to `from config import settings`
- src/tools/skills.py: auto-discovers src/skills/*/SKILL.md files, exposes list_skills() and use_skill(name) tools to the agent - src/skills/lambda-throttling/SKILL.md: step-by-step throttling investigation (concurrency checks, CloudTrail changes, account limits, traffic spikes, mitigation table) - src/agent/prompts.py: injects available skill names + descriptions into system prompt at startup; full content only loaded on use_skill() call - src/config/appsettings.py: checkpoint_backend uses Literal type - README: skills system marked done in mid-term roadmap
Automatically compacts long sessions before the LLM context window overflows.
Fires before each agent call when total message chars exceed
SUMMARIZATION_THRESHOLD_CHARS (default 60 000 chars ≈ 15 K tokens).
- src/agent/summarizer.py: maybe_summarize() reads LangGraph state,
splits at HumanMessage boundaries (no orphaned ToolMessages), calls LLM
to produce a structured summary, applies RemoveMessage + injects summary
HumanMessage via aupdate_state()
- src/api/routers/chat.py: single await maybe_summarize() call before astream()
- src/config/appsettings.py: summarization_enabled, threshold, keep_chars settings
- save_usage_event() gains metadata: dict param across all backends so
summarization events are recorded with {"summarization": True,
"messages_removed": N, "chars_removed": N} for dashboard queries
- migrations/003: ADD COLUMN metadata jsonb to usage_events (postgres)
- SQLite DDL + inline ALTER TABLE migration for existing databases
- dashboard /stats now includes total_summarizations and total_chars_compacted (queried from usage_events.metadata across all three backends) - add docs/tool_capping.md and docs/conversation_summarization.md - update _SUMMARY_KEYS in test_backends to reflect new fields
Add context management stat row (hidden until first summarization fires): - Sessions compacted: total summarization runs - Context saved: estimated tokens (~chars/4) compacted across all sessions
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
with_cap()at startup; responses overTOOL_RESPONSE_MAX_CHARS(default 40K chars) are truncated with a notice so the agent knows to narrow its querysrc/agent/config.pytosrc/config/appsettings.py, re-exported viasrc/config/__init__.py; all 21 import sites updatedsrc/skills/*/SKILL.mdfiles; skill names injected into system prompt at startup; full content loaded on demand viause_skill(name); ships withlambda-throttlingskillSUMMARIZATION_THRESHOLD_CHARS, compacts old messages via LLM call, preserving lastSUMMARIZATION_KEEP_CHARSof context; tracks events inusage_events.metadata/statsnow returnstotal_summarizationsandtotal_chars_compacted; dashboard UI shows a "Context management" row when non-zerodocs/tool_capping.md,docs/conversation_summarization.md,docs/skills.mdMigration required (Postgres only)
psql $DATABASE_URL -f migrations/003_usage_events_metadata.sqlAdds
metadata jsonbcolumn tousage_events. SQLite auto-migrates on startup.New env vars
Test plan
uv run pytest tests/ -q— 172 passed, 4 skippedTOOL_RESPONSE_MAX_CHARS=2000, ask "list all Lambda functions", confirm agent sees_cappedresponseSUMMARIZATION_THRESHOLD_CHARS=500, send 3+ messages, confirm log linesummarizer: session=... chars, removing N msgsGET /stats— confirmtotal_summarizationsandtotal_chars_compactedare non-zero after summarization firespython -c "from tools.skills import list_skills; print(list_skills())"— confirm lambda-throttling appears