Skip to content

About

LAPKT extensions and benchmark utilities for DNTP planning, including domain-driven axiom evaluation, BFS/GBFS experiments, and validation scripts.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

41 Commits

Folders and files

Repository files navigation

DNTP Planning Experiments for LAPKT

This repository contains the DNTP planning extension developed on top of LAPKT, together with source overlays, benchmarks, validation scripts and experiment reports.

The intended workflow is:

  1. clone a clean LAPKT checkout;
  2. copy exactly one source overlay;
  3. rebuild LAPKT;
  4. run and validate the DNTP benchmarks;
  5. use the provided scripts for ablation, tuning, tables, and figures.

Detailed setup commands are in steps.md.

Repository Layout

  • benchmarks/dntp/: DNTP domain, paper, PDDL instances, and JSON metadata.
  • possible_solutions/axioms/: generic lifted/compiled axiom evaluator with BFS support.
  • possible_solutions/axioms_gbfs/: the same generic evaluator plus the GBFS search layer used in the controlled evaluator comparison.
  • possible_solutions/dntp/: standalone domain-driven DNTP evaluator.
  • possible_solutions/dntp_gbfs/: latest DNTP evaluator with incremental graph updates, successor pruning, GBFS, and BFS fallback.
  • scripts/: execution, validation, ablation, tuning, analysis, and plotting.
  • results/: generated report-ready tables and figures.
  • pip_requirements.txt: Python build and analysis dependencies.
  • steps.md: complete installation and experiment procedure.

The overlays are intentionally independent. Do not combine files from multiple overlays unless you are deliberately merging those approaches.

Recommended Overlay

Use possible_solutions/dntp_gbfs for current experiments. It provides:

  • DNTP-specific graph evaluation instead of exhaustive generic axiom grounding;
  • ordered axiom checks with early termination;
  • incremental graph updates with optional differential verification;
  • successor pruning;
  • BFS(f) and GBFS search;
  • configurable GBFS time and plateau budgets;
  • BFS fallback when GBFS reaches a budget.

Use possible_solutions/axioms when the goal is to test the domain-independent axiom evaluator with BFS only. Use possible_solutions/axioms_gbfs when the same evaluator must also expose GBFS. Use possible_solutions/dntp when only the simpler domain-specific evaluator is needed.

Quick Start

After copying possible_solutions/dntp_gbfs/src/ over a clean LAPKT src/ and rebuilding:

uv run python Release/lapkt_package/lapkt.py BFS_f_Planner \
  -d benchmarks/dntp/domain.pddl \
  -p benchmarks/dntp/instances/P22.pddl \
  --grounder FD

Run and validate all instances with both planners:

uv run python scripts/run_dntp_instances.py \
  --planners BFS_f_Planner,GBFS_H_Add_Rp_Planner \
  --python-hash-seed 0 \
  --timeout 120

Ablation Study

The ablation study separates three effects:

  • full: incremental graph updates and successor pruning;
  • rebuild: complete graph reconstruction and successor pruning;
  • no_pruning: incremental updates without successor pruning.
uv run python scripts/run_dntp_ablation.py \
  --python-hash-seed 0 \
  --gbfs-budget 1.5 \
  --gbfs-plateau-limit 300 \
  --timeout 120 \
  --repetitions 5 \
  --output test_runs/dntp_ablation_runtime

Repeated runs are reduced to one median observation per instance before statistical inference. This prevents repetitions from being treated as independent benchmark instances.

Generate robust tables:

uv run python scripts/analyze_dntp_ablation.py \
  test_runs/dntp_ablation_runtime/combined_summary.csv \
  --out-dir results/dntp_ablation

Generate PDF and PNG boxplots:

uv run python scripts/plot_dntp_ablation.py \
  test_runs/dntp_ablation_runtime/combined_summary.csv \
  --out-dir results/dntp_ablation/figures

The reports include quartiles, IQR, median absolute deviation, Tukey outliers, outlier sensitivity, paired speedups, deterministic bootstrap confidence intervals, sign tests, and Holm-adjusted p-values.

Staged GBFS Tuning

The tuning process avoids selecting and evaluating parameters on the same instances. It screens 25 budget/plateau combinations on a stratified development subset, then evaluates the best three plus the baseline on an untouched holdout set.

Inspect the complete command matrix without launching planners:

uv run python scripts/run_dntp_tuning.py \
  --baseline-summary test_runs/dntp_ablation_runtime/combined_summary.csv \
  --output test_runs/dntp_tuning \
  --dry-run

Start or resume the real experiment by removing --dry-run. Completed configuration summaries are reused automatically.

After tuning:

uv run python scripts/analyze_dntp_ablation.py \
  test_runs/dntp_ablation_runtime/combined_summary.csv \
  --screening test_runs/dntp_tuning/screening_combined.csv \
  --holdout test_runs/dntp_tuning/holdout_combined.csv \
  --out-dir results/dntp_tuning

uv run python scripts/plot_dntp_ablation.py \
  test_runs/dntp_ablation_runtime/combined_summary.csv \
  --screening test_runs/dntp_tuning/screening_combined.csv \
  --holdout test_runs/dntp_tuning/holdout_combined.csv \
  --out-dir results/dntp_tuning/figures

Validation Notes

  • Matching JSON files are used for lightweight DNTP-specific validation.
  • P30 and P40 are known invalid inputs: they mark primary nodes as full, while the domain declares full only for secondary nodes.
  • PYTHONHASHSEED=0 is required for reproducible grounding/action ordering.
  • Rebuild LAPKT whenever any C++ source file changes.
  • The external VAL validator remains useful for small instances, but it may be memory-heavy for large derived-predicate problems.

General Evaluator Comparison

The repository also includes a controlled comparison between the general axiom evaluator and DNTP. It uses two permanent builds generated from the same clean LAPKT revision, then runs BFS and GBFS with identical settings, a 120-second timeout, an 8 GiB memory limit and VAL validation.

The experiment reports coverage (PASS, INVALID, NO_PLAN, TIMEOUT, OOM, CRASH) separately from paired runtime and peak-memory measurements. This avoids favouring an evaluator by silently excluding its failed runs.

The final general evaluator keeps normalized axioms lifted, compiles the proof support graph reachable from action and goal queries when feasible, evaluates compact supports by stratum, and falls back to demand-driven tabling when a compilation limit is reached. It is domain-independent: unlike the DNTP evaluator, it does not rely on graph-specific predicates such as radiality or restoration.

See steps.md for build, execution, analysis, plotting and tooling verification. The included benchmarks/dntp/smoke_instances.txt provides a four-instance pre-flight set before committing resources to the complete matrix.

About

LAPKT extensions and benchmark utilities for DNTP planning, including domain-driven axiom evaluation, BFS/GBFS experiments, and validation scripts.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages