Skip to content

Latest commit

 

History

430 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

FetchSmith

Pay-per-result web data extraction tools. Each Actor in actors/ targets one site or one kind of page, reads only what is publicly accessible (no logins, no personal-data harvesting), and returns tidy JSON.

Every Actor is HTTP-only — no headless browser anywhere in the stack. That is why they start in well under a second and never pay Chromium's launch time or memory footprint. We benchmarked the difference on a 1 vCPU / 2 GB box: 0.2 s and 50 structured records vs. 7 s, 146 MB and 83 unstructured links for the same page.

They run on Apify Store with pay-per-event pricing (you pay per result returned, no per-run start fee) and are catalogued with docs and sample output at fetchsmith.com.

Actors

Actor What it returns Links
Google News Scraper Articles by keyword, topic or publisher with the resolved publisher URL (not Google's redirect token), plus optional full article text, author, image and keywords. Any language or country. Apify Store · Docs · Source
App Store Reviews Scraper Apple App Store reviews for any iOS app and country storefront: rating, title, text, version, author, date, plus app metadata. Apify Store · Docs · Source
Google Play Reviews Scraper Google Play reviews and app details by app ID or search term: rating, text, date, developer replies, installs, score. Apify Store · Docs · Source
Shopify Products Scraper Full product catalog of any Shopify store or collection: prices, compare-at prices, currency, variants, SKUs, stock, images, tags. Apify Store · Docs · Source
Hacker News Scraper Stories, comments, Ask HN, Show HN and Who's Hiring threads via the official Algolia API, with author, points, comment-count and date filters. Apify Store · Docs · Source
Substack Scraper Any Substack publication as structured data: posts with full cleaned article text, engagement stats and comment threads. Custom domains supported. Apify Store · Docs · Source
Apple Podcasts Scraper Four modes in one Actor: every episode of a show with its direct audio file URL, listener reviews, podcast search, and country top charts. Apify Store · Docs · Source
Steam Reviews Scraper Steam player reviews with full text, playtime at review and total, recommended/not, helpfulness votes and verified-purchase flag — plus a game-details mode (price, genres, review score). Apify Store · Docs · Source
Scholarship Scraper (bold.org) Every bold.org scholarship as a row: award amount and number of awards, deadline (rolling flagged), the actual essay prompt with word limits, eligibility and a competitiveness ratio. Apify Store · Docs · Source
EU TED Tenders Scraper The EU's official TED public-procurement journal by country, CPV code and date: buyer, value, deadlines and notice links, deduplicated and language-flattened. Apify Store · Docs · Source
UK Public Contracts Both official UK portals in one deduplicated feed — Find a Tender (above threshold) and Contracts Finder (sub threshold): buyer email, phone and address, contract value, CPV codes, lots and deadlines. Apify Store · Docs · Source
USAspending Scraper Every US federal contract, IDV, grant, loan and direct payment from USAspending.gov: recipient UEI and address, awarding/funding agency, NAICS/PSC, CFDA program, place of performance. Apify Store · Docs · Source
FDA Recall Scraper Every US FDA product recall from the official openFDA enforcement API — food, drug and device in one schema: Class I/II/III severity, recalling firm, reason, ISO dates, plus NDC/UPC/brand/substance on drug recalls. Apify Store · Docs · Source
Federal Register Scraper US Federal Register rules, proposed rules, notices and presidential documents from the official government API: comment-close deadline, EO 12866 significance, RIN, docket IDs, CFR references — cursor paging past the API's own 10,000-row wall. Apify Store · Docs · Source
ClinicalTrials.gov Scraper US clinical trials from the official NIH API: condition, intervention, sponsor, location, status, type and phase filters, plus an optional one-row-per-trial-site mode for site-selection. No contact people, phones or emails shipped, ever. Apify Store · Docs · Source
Grants.gov Scraper US federal grant opportunities from the official Grants.gov API: keyword, agency, status, eligibility and funding-category filters, plus optional detail enrichment for award ceiling/floor, applicant-eligibility text and the full synopsis. Apify Store · Docs · Source
NIH RePORTER Scraper NIH-funded research projects from the official RePORTER API: keyword, fiscal year, institute, activity code, organization, state and PI filters, award amounts, study sections and an optional join to the PubMed papers each project produced — auto-chunks past the API's 15,000-row offset wall. Apify Store · Docs · Source
ATS Jobs Scraper Live job postings from any company's Greenhouse, Ashby, Lever, Recruitee, Workable or SmartRecruiters career board, normalized into one schema: location, remote status, salary where the ATS exposes it, department, team, employment type. Companies that moved off an ATS are skipped, not failed. Apify Store · Docs · Source
FEC Campaign Finance Scraper US federal candidates (House, Senate, President) from the official FEC open.fec.gov API, by name, state, office, party or election cycle, with each candidate's campaign financial totals — receipts, disbursements, cash on hand, individual contributions. Apify Store · Docs · Source

Guides

Write-ups of things we hit while building these — each one is a real, reproduced finding, not a tutorial rehash.

Layout

  • actors/<slug>/ – one Apify Actor (Node 20, apify SDK, HTTP-only, pay-per-event charging)
  • actors/_template/ – starting point for new Actors
  • site/ – FastAPI app behind fetchsmith.com (catalog, docs, guides, credit API)
  • bin/ – operations helpers (publishing, health checks, revenue snapshot)
  • notes/, tasks/, state/ – the autonomous operator's playbook, task queue and status

How these are built

FetchSmith is run end-to-end by an AI agent: it picks the niches, writes and tests the Actors, publishes them, writes the guides, and answers support mail. The playbook it follows is in notes/PLAYBOOK.md and its running notes are in notes/LEARNINGS.md — both are worth reading if you are curious what that actually looks like in practice. Every Actor is re-tested nightly against the live source.

Issues and requests: support@fetchsmith.com

About

Pay-per-result web scrapers: Google News, App Store & Google Play reviews, Shopify catalogs, Hacker News, Substack, Apple Podcasts, Steam reviews, scholarships, EU/UK/US public tenders & awards, and FDA recalls. HTTP-only, no headless browser. Runs on Apify Store.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages