Skip to content

Latest commit

 

History

132 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

WebGrep

WebGrep is a high-performance CLI keyword search tool with three modes: web crawl (recursively search a website and its linked documents), local file (search a single file offline), and local folder (recursively search all files in a directory). All modes use Apache Tika for text extraction from PDF, DOCX, TXT, and 100+ other formats, and report exact page and line numbers for every match.

Tech stack: Java 17+, Maven, Jsoup (HTML fetching/parsing), Apache Tika (document text extraction), Playwright for Java (JavaScript rendering for SPA pages), JUnit 4 (tests).

WebGrep demo


Getting Started

Option 1 - Native JAR (requires Java 17+)

mvn package
java -jar target/WebGrep-1.2.0.jar -u https://example.com -k "your keyword"

Option 2 - Docker (no Java required)

docker build -t webgrep .
docker run --rm webgrep -u https://example.com -k "your keyword"

For local files and folders, mount the host path into the container:

# Single file
docker run --rm -v /path/to/file.pdf:/data/file.pdf webgrep -f /data/file.pdf -k revenue

# Entire folder
docker run --rm -v /path/to/documents:/data webgrep -F /data -k confidential

The Docker image uses a multi-stage build - Maven compiles the JAR in the build stage, only the lightweight Alpine JRE is included in the final image (~90 MB).

SPA rendering in Docker: The Docker image does not include a headless browser. JavaScript-rendered SPA pages cannot be crawled from within the container. For SPA support, use the JAR directly and run --install-browser once to set up a browser on the host.


Usage

WebGrep has three input modes. Exactly one must be specified per run:

# Web crawl mode
java -jar WebGrep.jar -u <URL> -k <keyword> [options]

# Local file mode
java -jar WebGrep.jar -f <path> -k <keyword> [options]

# Local folder mode
java -jar WebGrep.jar -F <path> -k <keyword> [options]

Options

Input (exactly one required):

  • -u, --url <URL>: The starting URL for a web crawl. The scheme is optional and defaults to http://, so example.com/docs and localhost:8080 both work.
  • -f, --file <path>: Search a single local file. Supports all formats Apache Tika understands (PDF, DOCX, TXT, and more).
  • -F, --folder <path>: Recursively search all files in a local directory. Results are grouped by file.

Matching:

  • -k, --keyword <word>: The keyword to search for (required).
  • -m, --mode <mode>: Match strategy: default, exact, or fuzzy.

Web crawl controls:

  • -d, --depth <n>: Maximum crawl depth (default: 1). See Depth Definition below.
  • -p, --max-pages <n>: Stop after visiting N pages (default: 5000).
  • -t, --timeout-ms <n>: Network timeout per request in milliseconds (default: 20000).
  • -r, --delay-ms <n>: Delay between requests in milliseconds (default: 100).
  • -a, --all-urls: Disable smart URL deduplication. By default, page?id=1&sort=asc is treated as a variant of page?id=1 and skipped. Use this flag to visit every URL regardless.
  • -s, --dfs: Use depth-first search instead of the default breadth-first. BFS covers the site level by level; DFS follows each link chain as deep as possible before backtracking.
  • -e, --allow-external: Follow links outside the starting domain.
  • -i, --insecure: Disable SSL certificate verification. Use with caution.

SPA rendering:

  • --browser <type>: Browser to use for JavaScript rendering: auto (default), firefox, or chromium. See JavaScript-Rendered Pages below.
  • --install-browser: Pre-install a browser for SPA rendering and exit without crawling. Useful for one-time setup before running WebGrep in a non-interactive environment. Respects --browser preference.

General:

  • -n, --max-hits <n>: Stop after N total matches are found (default: 0 = no limit). Applies to both web crawl and folder scan. The check fires after each page or file completes - a single page/file may contribute more matches than the limit before the search stops. Pairs well with --dfs for web crawl to surface deeply buried results quickly.
  • -b, --max-bytes <n>: Maximum response or file size in bytes (default: 10MB). Responses or files exceeding this limit are skipped and counted in the summary.
  • -o, --output <format>: Output format: text (default) or json.
  • -h, --help: Show help message.

Matching Modes

  • Default: Case-insensitive with Unicode and diacritic support. cafe matches Café, CAFE, café. Both a case-insensitive regex pass and a diacritic-stripped, punctuation-removed pass run; the higher match count is returned. This ensures mixed text like "cafe Café" counts both variants rather than only one. (The diacritic-stripped pass is skipped for keywords that contain ASCII special characters such as node.js, .NET, or C++, since stripping punctuation would change the intended search term.)
  • Exact: Strict case-sensitive literal match. hello does not match Hello.
  • Fuzzy: First tries a normalised substring match. If that fails, splits the text into words and accepts any word within Levenshtein edit distance 1 (for keywords ≤ 4 characters) or 2 (for longer keywords). Catches common typos and spelling variants.

Depth Definition

Applies to web crawl mode only.

  • Depth 0: Fetches and searches only the seed URL. No links are followed.
  • Depth 1: Fetches the seed URL, then follows all links found on that page.
  • Depth N: Continues recursively, following links up to N hops from the seed.

Each URL is visited at most once per run. By default, URLs that look like navigation variants of an already-visited page are skipped - list?id=1&sort=asc is treated as a variant of list?id=1, but list?id=2 is treated as new content. Use --all-urls to disable this.


JavaScript-Rendered Pages (SPAs)

Many modern websites - government portals, dashboards, intranets - are built as single-page applications (Angular, React, Vue). Their content is injected into the page by JavaScript at runtime, which means a plain HTTP fetch returns an almost-empty HTML skeleton with no searchable text.

WebGrep detects these pages automatically and, when one is encountered, asks whether you want to enable headless browser rendering for the session:

  The page at https://example.gov/portal
  is a JavaScript-rendered single-page application (Angular, React, or Vue).
  Its content is not visible in a plain HTML fetch and will return no results
  without a headless browser.

  WebGrep can render it automatically. A compatible browser will be used
  if one is already installed; otherwise a one-time download (~105 MB) is
  required.

  Enable JavaScript rendering for this session? [Y/n]:

Answer Y (or just press Enter) to enable rendering. The choice applies for the rest of the crawl - you won't be asked again.

How the Browser Is Selected

WebGrep works through the following tiers in order, using the first option that succeeds:

  1. System Chromium or Chrome - if already installed on your machine, WebGrep uses it silently. No download required.
  2. System Firefox - used if found. Note: Firefox Developer Edition and Nightly may be incompatible with Playwright's protocol; WebGrep falls through to the next tier silently if they fail.
  3. Previously downloaded Playwright Firefox - if you've run WebGrep with SPA rendering before, the cached Firefox build is reused automatically.
  4. Previously downloaded Playwright Chromium - same as above for Chromium.
  5. First-time download - if no browser is available, WebGrep prompts you to choose Firefox (Mozilla Foundation, ~105 MB, default) or Chromium (Google, ~120 MB). The browser is downloaded once to ~/.cache/ms-playwright and reused on all future runs.

Use --browser firefox or --browser chromium to skip to a specific tier. --browser auto (the default) always tries the above order.

Pre-installing a Browser

To avoid the interactive prompt - for example, before running WebGrep in a script or CI pipeline - install a browser upfront with --install-browser:

# Let WebGrep pick (checks for system browsers first, prompts if none found)
java -jar WebGrep.jar --install-browser

# Force a specific browser
java -jar WebGrep.jar --install-browser --browser firefox
java -jar WebGrep.jar --install-browser --browser chromium

If a compatible system browser is already installed, this command reports it and exits without downloading anything.

Additional Details

  • Intercepted API links: Beyond rendering page text, WebGrep intercepts the JSON API calls a SPA makes during page load and extracts document download URLs that are not exposed as plain <a href> links (e.g. click-handler-driven PDF downloads). These are queued and searched like any other link.
  • Non-interactive mode: When stdin is not a terminal (piped input, scripts, CI), the SPA consent prompt is skipped and rendering proceeds automatically if a compatible browser is already installed. If no browser is found in a non-interactive context, WebGrep will not download one silently - it skips SPA rendering and logs a warning. Run --install-browser ahead of time to pre-install a browser for CI use.
  • Opt out: Answer n at the prompt to disable SPA rendering for the entire session. Detected SPA pages will be processed as plain HTML and will likely return no content.

Examples

Web Crawl

Basic search:

java -jar WebGrep.jar -u https://example.com -k domain

Deep crawl including linked PDFs and documents:

java -jar WebGrep.jar -u https://example.com -k "annual report" -d 2

Search directly inside a remote PDF:

java -jar WebGrep.jar -u https://example.com/report.pdf -k "revenue" -d 0

Search an Angular/React/Vue portal (SPA):

java -jar WebGrep.jar -u https://portal.example.gov -k "contract" -d 1
# WebGrep detects the SPA, prompts once, then renders all pages automatically

JSON output:

java -jar WebGrep.jar -u https://example.com -k domain -o json

Fast crawl with no delay:

java -jar WebGrep.jar -u https://example.com -k topic -d 2 -r 0

Crawl all sort/filter variants of a listing page:

java -jar WebGrep.jar -u https://example.com/listings -k topic -d 1 --all-urls

Stop after the first 5 matches (BFS - most prominent results first):

java -jar WebGrep.jar -u https://example.com -k "annual report" -d 2 -n 5

Stop after the first 5 matches (DFS - dives deep before backtracking):

java -jar WebGrep.jar -u https://example.com -k "annual report" -d 3 -s -n 5

Local File

Search a PDF or plain text file:

java -jar WebGrep.jar -f /path/to/report.pdf -k "revenue"

Case-sensitive search in a DOCX file:

java -jar WebGrep.jar -f /path/to/contract.docx -k "Clause 4.2" -m exact

JSON output:

java -jar WebGrep.jar -f /path/to/report.pdf -k "revenue" -o json

Local Folder

Search all files in a folder recursively:

java -jar WebGrep.jar -F /path/to/documents -k "confidential"

Skip files larger than 5 MB:

java -jar WebGrep.jar -F /path/to/documents -k "invoice" -b 5242880

JSON output:

java -jar WebGrep.jar -F /path/to/documents -k "confidential" -o json

Document Support

All three modes use Apache Tika for text extraction. Supported formats include:

  • PDF (.pdf)
  • Word (.doc, .docx)
  • Plain text (.txt)
  • And many more - Tika supports 100+ formats including ODT, RTF, EPUB, XLS, XLSX, PPT, PPTX, and more.

Each file is parsed with a 30-second timeout. If parsing takes longer (e.g. a corrupt or malformed file), it is skipped and the raw bytes are used as a UTF-8 fallback.

Web crawl mode: static assets (images, video, CSS, JS, fonts, archives, social share links) are filtered by URL before any request is made.

File and folder modes: every file is passed to Tika regardless of extension. Files exceeding --max-bytes are skipped and counted in the summary.


Sample Output (JSON)

Web Crawl

{
  "query": {
    "url": "https://example.com",
    "depth": 1,
    "keyword": "domain",
    "mode": "default"
  },
  "stats": {
    "duration_ms": 1243,
    "total_matches": 13,
    "pages_visited": 1,
    "pages_parsed": 1,
    "docs_parsed": 0,
    "pages_blocked": 0,
    "errors": {
      "network_error": 0,
      "blocked": 0,
      "parse_error": 0,
      "skipped_size": 0,
      "skipped_type": 0
    }
  },
  "results": [
    { "url": "https://example.com/", "count": 13, "snippets": ["...the domain example.com is used to illustrate...", "...illustrative examples in documents without prior coordination..."] }
  ],
  "blocked": []
}

Local File

{
  "query": {
    "file": "/path/to/report.pdf",
    "keyword": "revenue",
    "mode": "default"
  },
  "stats": {
    "duration_ms": 84,
    "total_matches": 3
  },
  "matches": [
    { "page": 1, "line": 4,  "count": 1, "snippet": "Total revenue for the year was $4.2M" },
    { "page": 2, "line": 11, "count": 2, "snippet": "Revenue growth and revenue targets exceeded" }
  ]
}

For plain text files (no page structure), the page field is omitted from each match object.

Local Folder

{
  "query": {
    "folder": "/path/to/documents",
    "keyword": "confidential",
    "mode": "default"
  },
  "stats": {
    "duration_ms": 312,
    "files_scanned": 5,
    "files_skipped": 1,
    "files_failed": 0,
    "files_with_matches": 2,
    "total_matches": 4
  },
  "results": [
    {
      "file": "/path/to/documents/contract.docx",
      "total_matches": 3,
      "matches": [
        { "line": 2, "count": 1, "snippet": "CONFIDENTIAL - Do not distribute" },
        { "line": 17, "count": 2, "snippet": "This document is confidential and confidential use only" }
      ]
    },
    {
      "file": "/path/to/documents/notes.txt",
      "total_matches": 1,
      "matches": [
        { "line": 5, "count": 1, "snippet": "Mark as confidential before sending" }
      ]
    }
  ]
}

JSON Fields Reference

Not all fields are listed - self-explanatory fields such as url, count, snippet, and query.* are visible in the sample output above. The table below covers fields whose meaning may not be immediately obvious.

Field Meaning
stats.duration_ms Wall-clock time for the entire run in milliseconds
stats.docs_parsed (web crawl) Binary documents (PDF, DOCX, etc.) parsed during the crawl
stopped_early (web crawl) Present only when --max-hits triggered an early stop
stats.errors.network_error (web crawl) Requests that failed (DNS, timeout, connection refused, non-403/429 HTTP errors)
stats.errors.network_error_reasons (web crawl) Breakdown by cause (e.g. Timeout: 30, HTTP 404: 5)
stats.errors.blocked (web crawl) Pages returning 403/429 or detected bot-protection challenges
stats.errors.skipped_size (web crawl) Files that exceeded --max-bytes and were not parsed
stats.files_skipped (folder) Files that exceeded --max-bytes and were not scanned
stats.files_failed (folder) Files that could not be read or parsed (permission error, corrupt file, etc.)
matches[].page (file/folder) Page number (1-based); omitted when the file has no page structure
matches[].line (file/folder) Line number within the page (1-based)
matches[].count (file/folder) Keyword occurrences on that line
matches[].snippet (file/folder) Trimmed line content, truncated to 120 characters

Architecture

WebGrep is designed with a modular architecture for performance and maintainability:

  • CliOptions: Parses and validates all command-line arguments. Enforces mutual exclusion between --url, --file, and --folder.
  • Crawler: Manages the multi-level crawl queue, domain scoping, body size limits, and configurable politeness delays. Maintains a session cookie jar across requests, and retries automatically on HTTP 429 with exponential backoff. Displays a live progress indicator during the crawl. Separates links into navigation links (respect depth limit) and document links (always enqueued - documents are leaves that cannot expand the crawl frontier). On first SPA encounter, prompts the user and delegates to PlaywrightRenderer.
  • PlaywrightRenderer: Headless browser renderer for JavaScript-rendered SPAs. Manages browser lifecycle and selection across five tiers (system Chromium → system Firefox → cached Playwright Firefox → cached Playwright Chromium → first-time download). Uses a.href (the DOM property) for link resolution so that Angular/React apps with <base href> resolve relative URLs correctly. Also intercepts JSON API responses to capture document download URLs that SPAs serve via click handlers rather than plain <a href> links.
  • BrowserFinder: Locates system-installed browser binaries (Chromium/Chrome and Firefox) across Linux, macOS, and Windows via known install paths, falling back to which/where for non-standard package manager locations.
  • ContentExtractor: Extracts searchable text from HTML pages via Jsoup, and from binary documents via Apache Tika with a 30-second timeout per file to prevent hangs on corrupt or oversized content.
  • MatchEngine: Pluggable matching strategies (case-insensitive with diacritic support, exact, fuzzy/Levenshtein). Compiled regex patterns are cached per keyword/mode pair for performance. Returns both match counts and context snippets.
  • UrlDeduplicator: Smart URL deduplication. Treats ?id=1&sort=asc as a variant of ?id=1 (same content, extra parameter) but allows ?id=2 through (different content). Use --all-urls to visit every URL regardless.
  • ReportWriter: Renders results as human-readable text or structured JSON. Covers all three input modes.

Limitations

  • JavaScript (inline): WebGrep renders full SPA frameworks (Angular, React, Vue) automatically. Inline JavaScript that is not part of an SPA - custom onclick handlers, lazy loaders triggered by user interaction - is not evaluated.
  • Bot Protection: JavaScript-based challenges (e.g. Cloudflare Managed Challenges) cannot be bypassed. They are detected and reported under blocked.
  • Robots.txt: Not parsed. Use --delay-ms to be polite to servers.
  • Authentication: No support for login forms or HTTP Basic Auth. Session cookies set by the server are maintained across requests automatically.

Written by and belongs to Simon D.
Free to use for personal and educational purposes.
For commercial use please contact me at simon.d.dev@proton.me.

About

CLI keyword search across websites, documents, and local folders. Crawls multi-level sites, renders JavaScript SPAs with a headless browser, and extracts text from PDF, DOCX and 100+ formats via Apache Tika — reporting exact page and line numbers. Fuzzy matching, JSON output, Docker image.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages