Fix bugs in error handling, URL filtering, and response validation - #6
Merged
Conversation
- Fix cli_toolbox.py error flag: was checking is_spam truthiness instead of checking for missing response keys - Fix process_email.py: reuse BeautifulSoup object instead of parsing HTML twice; strip whitespace from URLs; filter out data:, javascript:, vbscript:, and mailto: schemes from extracted links/images - Fix email_ai_interface.py: add errors="replace" to file open for consistent encoding handling with rest of codebase - Fix external_ai_test.lua: defensive nil handling for missing confidence field in AI response - Add response validation (email_utils/validation.py): normalize is_spam to yes/no, clamp confidence to 0-100, ensure reason is string - Add 15 new tests (57 total): validation logic, URL filtering, data URI filtering, whitespace stripping
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
is_spamtruthiness instead of detecting missing response keysdata:,javascript:,vbscript:, andmailto:schemes from extracted links/imageserrors="replace"to file open for consistent encoding handling with rest of codebaseconfidencefield in AI responseis_spamto yes/no, clamps confidence to 0-100, ensures reason is a string. Used by both Anthropic and OpenAI providers.Test plan
ruff check . && ruff format --check .— cleanpytest -v— 57 tests pass