This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
This is an RSS feed generator that creates RSS feeds for blogs that don't provide them. The project uses Python scripts to scrape blog websites and convert them to RSS XML feeds, with automated updates via GitHub Actions.
# Create virtual environment
make env_create
# Activate virtual environment (run the output of this command)
$(make env_source)
# Install dependencies
make uvx_install
# Clean generated files
make clean# Generate all RSS feeds
make generate_all_feeds
# Generate specific feeds
make generate_anthropic_news_feed
make generate_anthropic_engineering_feed
make generate_anthropic_research_feed
make generate_openai_research_feed
make generate_ollama_feed
make generate_paulgraham_feed# Format Python code
make py_format# Test the feed generation workflow locally (requires 'act' tool)
make test_feed_workflow
# Run test feed generator
make test_feed_generate-
Feed Generators (
feed_generators/): Individual Python scripts that scrape specific blogs and generate RSS feeds- Each script follows the pattern: fetch HTML → parse content → generate RSS XML
- Common utilities: BeautifulSoup for HTML parsing, feedgen for RSS generation, requests for HTTP
- Output location:
feeds/feed_*.xml
-
Shared module (
feed_generators/_common.py): All generators must use this for content cleaning and feed output. It providesclean_article_html(sanitizes scraped article HTML to the reader-safe subset: headings, lists, figures/captions, tables, code blocks; resolves lazy-loaded images; strips nav/CTA/share chrome),extract_summary(og:description-based item summaries),build_feed/save_feed(CDATA content:encoded, permalink guids, media:content lead image, correct channel link order, 50-item cap, skip-write-when-unchanged), andload_cached_entries(reads the previous feed XML back as a content cache). Generators keep only site-specific fetching/parsing; never add per-generator HTML cleaning. -
Orchestration (
run_all_feeds.py): Main script that executes all feed generators automatically -
Automation (
.github/workflows/run_feeds.yml): GitHub Action that runs twice daily (midnight and noon UTC) to update all feeds
Each feed generator script follows this structure:
- Fetch HTML content from target blog using requests
- Parse HTML with BeautifulSoup to extract article data
- Use feedgen library to create RSS XML
- Save to
feeds/directory with patternfeed_*.xml
Key Python packages:
beautifulsoup4&bs4: HTML parsingfeedgen: RSS feed generationrequests: HTTP requestsselenium&undetected-chromedriver: Browser automation for dynamic contentpython-dateutil&pytz: Date/time handling
- Create new Python script in
feed_generators/following existing patterns - Add corresponding make target to Makefile
- Script will be automatically included in
run_all_feeds.pyexecution - GitHub Actions will run the new script twice daily
feed_generators/: Python scripts for individual blog scrapersfeeds/: Generated RSS XML files and cachesrequirements.txt: Python dependenciesMakefile: Development commands and shortcuts.github/workflows/: Automated feed generation and testing