Graven (grype + maven) is a recursive and optimized crawler for scraping the Maven Central Repository.
Graven 2.0 introduces three additional workers into the pipeline, while existing workers have been specialized:
-
Crawler: Recursive Maven Central crawler that parses file trees looking for jars to download
-
Downloader: Mass downloader for jars from Maven Central
-
Generator: Use syft to scan the downloaded jars to generate SBOMs
-
Scanner: Use grype to scan the generated SBOMs for CVEs
-
Analyzer: Parse the
syftandgrypereports and save vulnerability information into the database
Adjacent to the main pipeline, there is an additional NVD & MITRE module that add additional vulnerability information in realtime.
- Project Overview
- Starting the Database
- Launching Graven (Docker: RECOMMENDED)
- Local Deployment
- Usage
Graven is a high-performance crawler and analysis dataset generation tool that generates Software Bill of Materials (SBOMs) and scan for Common Vulnerabilities and Exposures (CVEs) for JAR binaries from Maven Central Repository. It automates the end-to-end process of discovering and processing and storing CVE and dependency relationships between Maven JARs in a single easy to use pipeline. Graven is powerful solution for security researchers, developers, and organizations to quickly generate datasets at scale for analysis and greater insight in their Maven supply chains.
The compose file will create a volume so the database contents will persist between launches
- Create an
.envfile Creating an.envfile in this directory and copy the values below. Set any<...-here>with your credentials:
# connection details
MYSQL_HOST=localhost
MYSQL_DATABASE=<db-name-here>
EXTERNAL_PORT=3306
# user details
MYSQL_ROOT_PASSWORD=<root-pw-here>
MYSQL_USER=<username-here>
MYSQL_PASSWORD=<user-pw-here>
# optional NVD API Key
NVD_API_KEY=<your-key-here>
- Launch the database
docker compose up -dRemove -d if you want to see the logs
To access the database, run
docker exec -it graven_database mysql -u <name set for MYSQL_USER> -p <db set for MYSQL_DATABASE>and enter the MYSQL_PASSWORD set in the .env file when prompted.
For external connections, the database will be hosted at localhost:3306
- Build the graven image
docker build -t graven:2.2.1 .- Run the container attached to the database network
docker run --rm -it --env-file .env \
-e MYSQL_HOST=mysql \
-v grype_db:/home/graven/.cache/grype \
--network=graven_database_graven graven:2.2.1 run --root-url <start-url>--rm: remove container when finished-it: open an interactive terminal (so can see logs)--env-file <path/to/.env>: .env file to use, same one as for the database-e MYSQL_HOST=mysql: Set the MYSQL_HOST tomysql(name of database service) since database not running on localhost inside the container-v grype_db:/home/graven/.cache/grype: Create a cached volume for the grype security database, so it doesn't need to download everytime a new container is launched--network=graven_database_graven: Attach to the same network the database is running ongraven: Name of image--root-url <start-url>: Root url to start parsing from
If using the --seed-urls-csv flag, you will also need to mount the directory containing the csv file like so:
docker run --rm -it --env-file .env \
-e MYSQL_HOST=mysql \
-v grype_db:/home/graven/.cache/grype \
-v "<path-to-csv-dir>:/csv" \
--network=graven_database_graven graven:2.2.1 run --seed-urls-csv /csv/<your csv file>On first run, the grype database will take 1-3 minutes to initialize. If you want to setup the grype database before running graven, you can do so with the following docker command:
docker run --rm -it -v grype_db:/home/graven/.cache/grype --entrypoint ash graven -c "grype db update"Graven uses syft and grype to generate SBOMs and scan for CVEs.
For Linux:
curl -sSfL https://raw.githubusercontent.com/anchore/grype/main/install.sh | sh -s -- -b /usr/local/bincurl -sSfL https://raw.githubusercontent.com/anchore/syft/main/install.sh | sh -s -- -b /usr/local/binFor Windows: Go to the releases page and download the latest version.
Either place the path to the binary on the PATH or use the --syft-path and --grype-path arguments respectively to
use an absolute path.
- Create venv
python3 -m venv venv - Activate venv
. venv/bin/activate - Install dependencies
pip install -r requirements.txtusage: graven [-h] [-l <log level>] [-s]
{run,crawl,process,update-vuln,export} ...
Recursive and optimized crawler for scraping the Maven Central Repository
positional arguments:
{run,crawl,process,update-vuln,export}
run Run the entire graven pipeline
crawl Crawl Maven Central for jars
process Process jars stored in the database
update-vuln Update CVE and CWE data
export Export SBOMs in the database to file
options:
-h, --help show this help message and exit
-l <log level>, --log-level <log level>
Set log level (Default: INFO) (['INFO', 'DEBUG',
'ERROR'])
-s, --silent Run in silent mode
usage: graven run [-h]
(--root-url <starting url> | --seed-urls-csv <path to csv>)
[--update-domain] [--update-jar] [-u]
[--download-cache-size <cache size in MB>]
[--jar-limit <max jar count>]
[--syft-path <absolute path to syft binary>]
[--syft-cache-size <cache size in MB>] [--disable-syft]
[--grype-path <absolute path to grype binary>]
[--grype-db-source <url of grype database to use>]
[--grype-cache-size <cache size in MB>]
[--max-concurrent-maven-requests <number of requests>]
[--max-cpu-threads <number of the threads>]
[--disable-update-vuln]
Run the entire graven pipeline
options:
-h, --help show this help message and exit
Input Options:
Only one flag is permitted
--root-url <starting url>
Root URL to start crawler at
--seed-urls-csv <path to csv>
CSV file of root urls to restart the crawler at once
the current root url is exhausted
Crawler Options:
--update-domain Update domains that have already been crawled. Useful
for ensuring no jars were missed in a domain
--update-jar Update jars that have already been crawled
-u, --update Update domains AND jars that have already been
crawled. Supersedes --update-* flags
Downloader Options:
--download-cache-size <cache size in MB>
Limit of the number of jars to be saved at one time.
(Default: 5120.0 MB)
--jar-limit <max jar count>
Limit the number of jars downloaded at once
Generator Options:
--syft-path <absolute path to syft binary>
Path to syft binary to use. By default, assumes syft
is already on the PATH
--syft-cache-size <cache size in MB>
Limit of the number of grype files to be saved at one
time. (Default: 5120.0 MB)
--disable-syft Disable SBOM generation and scan jars directly
Scanner Options:
--grype-path <absolute path to grype binary>
Path to Grype binary to use. By default, assumes grype
is already on the PATH
--grype-db-source <url of grype database to use>
URL of specific grype database to use. To see the full
list, run 'grype db list'
--grype-cache-size <cache size in MB>
Limit of the number of grype files to be saved at one
time. (Default: 5120.0 MB)
Miscellaneous Options:
--max-concurrent-maven-requests <number of requests>
Max number of requests can make at once to Maven
Central. (Default: 100)
--max-cpu-threads <number of the threads>
Max number of threads allowed to be used to generate
anchore results. Increase with caution (Default: 8)
--disable-update-vuln
Disable real-time queries for CVE and CWE details
usage: graven crawl [-h]
(--root-url <starting url> | --seed-urls-csv <path to csv>)
[--update-domain] [--update-jar] [-u]
[--max-concurrent-maven-requests <number of requests>]
Crawl Maven Central for jars
options:
-h, --help show this help message and exit
Input Options:
Only one flag is permitted
--root-url <starting url>
Root URL to start crawler at
--seed-urls-csv <path to csv>
CSV file of root urls to restart the crawler at once
the current root url is exhausted
Crawler Options:
--update-domain Update domains that have already been crawled. Useful
for ensuring no jars were missed in a domain
--update-jar Update jars that have already been crawled
-u, --update Update domains AND jars that have already been
crawled. Supersedes --update-* flags
Miscellaneous Options:
--max-concurrent-maven-requests <number of requests>
Max number of requests can make at once to Maven
Central. (Default: 100)
usage: graven process [-h] [--download-cache-size <cache size in MB>]
[--jar-limit <max jar count>]
[--syft-path <absolute path to syft binary>]
[--syft-cache-size <cache size in MB>] [--disable-syft]
[--grype-path <absolute path to grype binary>]
[--grype-db-source <url of grype database to use>]
[--grype-cache-size <cache size in MB>]
[--max-concurrent-maven-requests <number of requests>]
[--max-cpu-threads <number of the threads>]
[--enable-update-vuln]
Process jars stored in the database
options:
-h, --help show this help message and exit
Downloader Options:
--download-cache-size <cache size in MB>
Limit of the number of jars to be saved at one time.
(Default: 5120.0 MB)
--jar-limit <max jar count>
Limit the number of jars downloaded at once
Generator Options:
--syft-path <absolute path to syft binary>
Path to syft binary to use. By default, assumes syft
is already on the PATH
--syft-cache-size <cache size in MB>
Limit of the number of grype files to be saved at one
time. (Default: 5120.0 MB)
--disable-syft Disable SBOM generation and scan jars directly
Scanner Options:
--grype-path <absolute path to grype binary>
Path to Grype binary to use. By default, assumes grype
is already on the PATH
--grype-db-source <url of grype database to use>
URL of specific grype database to use. To see the full
list, run 'grype db list'
--grype-cache-size <cache size in MB>
Limit of the number of grype files to be saved at one
time. (Default: 5120.0 MB)
Miscellaneous Options:
--max-concurrent-maven-requests <number of requests>
Max number of requests can make at once to Maven
Central. (Default: 100)
--max-cpu-threads <number of the threads>
Max number of threads allowed to be used to generate
anchore results. Increase with caution (Default: cpu corse)
--enable-update-vuln Enable real-time queries for CVE and CWE details
usage: graven update-vuln [-h]
Update CVE and CWE data. Will use 'NVD_API_KEY' env variable if available
options:
-h, --help show this help message and exit
usage: graven export [-h] -d DIRECTORY -c {zip,tar.gz}
Export SBOMs in the database to file
options:
-h, --help show this help message and exit
-d DIRECTORY, --directory DIRECTORY
Directory to save dump to
-c {zip,tar.gz}, --compression-method {zip,tar.gz}
Compression mode to export data to
If you encounter a bug, have a question, or want to suggest a feature, feel free to open a GitHub issue! For contributing, see CONTRIBUTING.md.