Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Passive Wi-Fi Probe Request Fingerprinting Localization Benchmark

Python License: CC BY 4.0 Status Reproducibility

This repository accompanies the manuscript “An End-to-End Pipeline for Passive Wi-Fi Probe Request Fingerprinting-Based Localization with Comparative Evaluation from Classical Retrieval to Attention-Based Models”, submitted to the Journal of Network and Computer Applications.

The repository provides an anonymized passive Wi-Fi probe-request dataset, derived fingerprint matrices, benchmark scripts, reference result tables, paper-output figures, and reproducibility instructions for comparing passive PR fingerprinting localization methods in one controlled indoor testbed.

At a Glance

Component Included
Passive PR localization dataset Yes
Raw anonymized probe-request frames Yes
Processed fingerprint matrices Yes
Reproducible benchmark scripts Yes
Retrieval, ensemble, neural, and attention-based models Yes
Paper figures and result tables Yes
Chart reproduction materials Yes
Identifier sanitization audit Yes
Public release validation checks Yes

Visual Overview

This repository is designed to help readers understand the full benchmark story quickly. The figures below move from the passive sensing principle, to the end-to-end pipeline, to the data-acquisition setup, model families, and key benchmark findings.

1. Passive Wi-Fi Probe-Request Fingerprinting

Passive probe-request fingerprinting differs from conventional active Wi-Fi RSS fingerprinting. In the active setting, the user device scans surrounding APs and sends an RSS vector to a localization server. In the passive setting, monitor-mode APs or passive sniffers capture probe-request frames emitted by the device, and localization is performed from the infrastructure side.

Active Wi-Fi RSS fingerprinting versus passive Wi-Fi probe-request fingerprinting

2. End-to-End Benchmark Pipeline

The benchmark follows a unified end-to-end workflow: passive PR data acquisition, centralized packet logging, preprocessing, fingerprint construction, controlled cross-family localization benchmarking, and deployment-sensitivity analysis.

End-to-end pipeline for passive Wi-Fi probe-request fingerprinting localization

3. Application Context

Passive Wi-Fi probe-request sensing can support infrastructure-side spatial awareness in healthcare, retail, venue management, occupancy analytics, and emergency response. This benchmark focuses specifically on the localization-oriented part of this broader sensing landscape.

Passive Wi-Fi probe-request sensing applications

4. Centralized Passive PR Collection

The data-acquisition setup uses multiple monitor-mode AP streams, centralized ncat listeners, and separate log files for each AP. This structure enables synchronized multi-AP passive packet capture and reproducible fingerprint construction.

Centralized passive probe-request data collection architecture

Benchmark Design

The repository evaluates representative localization paradigms under the same passive PR fingerprint matrix, spatially disjoint train-test split, and coordinate-based evaluation protocol.

Stage Purpose
Data acquisition Capture passive probe-request frames using monitor-mode APs
Centralized logging Store AP-specific packet streams in synchronized logs
Fingerprint construction Aggregate RSS observations into AP-wise fingerprint vectors
Benchmark formation Use 36 reference locations and 9 held-out test locations
Model comparison Evaluate retrieval, ensemble, neural, and attention-based families
Evaluation Analyze accuracy, robustness, worst-case error, latency, and receiver-layout sensitivity

Benchmarked Model Families

Retrieval-Based and Scalable Fingerprint Matching

The retrieval family provides transparent and strong fingerprinting baselines. The HNSW pipeline accelerates nearest-neighbor retrieval while preserving the matching logic of classical KWNN localization.

HNSW-based approximate nearest-neighbor localization pipeline

Tree-Based Ensemble Models

Tree-based regressors provide nonlinear supervised baselines between classical retrieval and higher-capacity neural models.

Random Forest localization pipeline
Random Forest
XGBoost localization pipeline
XGBoost
CatBoost localization pipeline
CatBoost
Baseline MLP architecture for RSS fingerprint localization
Baseline MLP

Neural and Attention-Based Models

The benchmark also includes MLP-based neural architectures and Transformer-style attention-based models to assess whether higher-capacity learned representations improve passive PR localization under common benchmark conditions.

MLP architecture comparison for passive PR fingerprint localization

ANN sensitivity analysis for RMSE and latency across training epochs

Spatial prediction comparison between Transformer architectures 3 and 4

Key Benchmark Findings

The benchmark does not support a simple “deeper is always better” conclusion. Instead, the evaluated methods occupy different regions of the accuracy, robustness, latency, and deployment-sensitivity trade-off space.

Cross-Family Accuracy-Latency Trade-Off

The strongest retrieval methods remain highly competitive for typical-case accuracy and latency. The most favorable attention-based configuration provides strong large-error control while retaining low inference time.

Comprehensive cross-family comparison of localization error, latency, accuracy-latency trade-off, and robustness

Error Distribution Across Method Families

The cross-family error distributions show that retrieval methods remain strong typical-case baselines, while attention-based models become useful when upper-tail error control is emphasized.

Cross-family error distribution across statistical, tree-based, ANN, and Transformer methods

Transformer-Family Error Behavior

The Transformer-family results show that attention-based modeling is not automatically superior. The most favorable configuration is the regularized AE+Transformer variant, which reduces the upper-tail error in this benchmark.

Transformer architecture error distributions

Monitor-Node Density and Receiver-Geometry Sensitivity

The deployment-sensitivity analysis shows that localization performance is affected not only by the number of monitor nodes, but also by receiver geometry. These results should be interpreted as within-environment sensitivity findings for the evaluated testbed.

Monitor-node density and receiver-geometry sensitivity analysis

Main Takeaways

  • Passive PR localization requires an end-to-end benchmark workflow, not only a localization estimator.
  • Classical retrieval methods remain highly competitive under the evaluated passive PR testbed.
  • HNSW indexing preserves KWNN-level accuracy while reducing online retrieval cost.
  • Tree-based ensembles provide useful intermediate operating points.
  • Standard MLP models add computational cost without clear benchmark-level advantage in this dataset.
  • The regularized attention-based configuration improves upper-tail error behavior under the evaluated setting.
  • Monitor-node density and receiver geometry materially affect localization performance within the same environment.
  • The results are benchmark-level findings for one controlled passive PR testbed, not broad cross-building generalization claims.

Scope of the Benchmark

The benchmark is intentionally single-environment and reproducibility-oriented. It supports controlled comparison under a fixed data-acquisition and evaluation protocol, but it should not be interpreted as evidence of broad cross-building, cross-device, or cross-deployment generalization.

Item Value
Indoor environments 1
Training/reference locations 36
Held-out test locations 9
Monitor nodes / AP columns 6
Raw training frames 9,373
Raw test frames 2,322
Coordinate unit millimeters in raw/processing, meters in reported errors

Repository Layout

.
├── configs/                    # Shared experimental constants and hyperparameters
├── data/
│   ├── raw/                    # Anonymized raw passive PR captures
│   ├── processed/              # Aggregated fingerprint matrices for quick reuse
│   └── README.md               # Dataset schema and privacy notes
├── experiments/                # Reproducible benchmark entry points
├── plots/                      # Plotting scripts and selected reference plots
├── paper_outputs/              # Manuscript figures, result tables, and run logs
├── reproducibility/
│   └── paper_charts/           # Chart provenance and reproduction materials
├── results/                    # Executable benchmark outputs
├── scripts/                    # Repository validation and reproduction checks
├── src/                        # Data loading, metrics, models, and utilities
├── tests/                      # Data-integrity and smoke tests
├── CITATION.cff                # Citation metadata
├── DATASET_CARD.md             # Dataset scope, limitations, and ethical notes
├── IDENTIFIER_AUDIT.md         # Identifier-sanitization and legacy-file audit
├── REPRODUCIBILITY.md          # Exact steps for reproducing benchmark outputs
├── RELEASE_VALIDATION.md       # Commands used to check this release
├── requirements.txt            # Core dependencies
├── requirements-full.txt       # Optional full-stack dependencies
└── pyproject.toml              # Project metadata and tooling configuration

Quick Start

# 1. Create and activate an environment
python -m venv .venv
source .venv/bin/activate      # Linux/macOS
# .venv\Scripts\activate       # Windows PowerShell

# 2. Install the core dependencies
python -m pip install --upgrade pip
pip install -r requirements.txt

# 3. Run the fast release-verification path
python scripts/verify_release.py

Alternatively, run the main checks step by step:

python scripts/check_repository.py
python scripts/check_paper_outputs.py
python scripts/check_chart_provenance.py
python scripts/check_chart_reproduction.py
python -m pytest -q -p no:cacheprovider
python experiments/run_kwnn_benchmark.py

The release-verification path is designed to be reviewer-friendly and CPU-safe. It checks the repository structure, data integrity, core benchmark execution, and paper-output availability.

Full Benchmark

The extended benchmark includes optional packages such as PyTorch, XGBoost, and CatBoost. Install them only when the corresponding model families are needed:

pip install -r requirements-full.txt

python experiments/run_all.py --full

# Or run selected components manually
python experiments/run_ensemble_benchmark.py
python experiments/run_mlp_benchmark.py
python experiments/run_transformer_benchmark.py
python experiments/run_latency_benchmark.py
python experiments/run_bootstrap_analysis.py
python experiments/run_ap_density_analysis.py
python plots/generate_all_plots.py
python plots/generate_paper_figures.py

Some extended experiments may take longer and can be sensitive to package versions and hardware. TensorFlow/Keras paths are optional and can be installed separately with requirements-tensorflow.txt on supported Python/platform combinations. See REPRODUCIBILITY.md for details.

Data Files

The repository includes both raw and processed forms of the anonymized dataset.

File Description
data/raw/df_all.csv Raw anonymized passive PR frames for the 36 reference locations
data/raw/test_df.csv Raw anonymized passive PR frames for the 9 held-out test locations
data/processed/train_fingerprints.csv Mean RSSI fingerprint matrix with coordinates for training/reference locations
data/processed/test_fingerprints.csv Mean RSSI fingerprint matrix with coordinates for held-out test locations

The raw files use semicolon separators. The processed files use comma separators and are included to make the benchmark easier to inspect and reuse.

Data Anonymization

The public release does not include the original device MAC addresses or real AP/BSSID identifiers. Device-level identifiers were replaced by deterministic pseudonymous IDs, and AP/BSSID-like identifiers were replaced by generic AP labels such as ap1 to ap6.

Internal legacy notebooks and raw chart-export files that contained original identifiers were excluded from the public release because they are not required for executing the benchmark or regenerating the paper figures. See IDENTIFIER_AUDIT.md for the identifier-sanitization audit.

Implemented Model Families

Family Representative scripts
Weighted nearest-neighbor retrieval experiments/run_kwnn_benchmark.py
Tree ensembles experiments/run_ensemble_benchmark.py
Neural MLP models experiments/run_mlp_benchmark.py
Attention-based models experiments/run_transformer_benchmark.py
Latency evaluation experiments/run_latency_benchmark.py
Bootstrap uncertainty summaries experiments/run_bootstrap_analysis.py
Monitor-node density sensitivity experiments/run_ap_density_analysis.py

Paper-Output Archive

The manuscript-facing outputs are included under paper_outputs/:

Folder/file Purpose
paper_outputs/figures/ Figure files using the same filenames referenced in the manuscript
paper_outputs/results/ Generated per-point outputs and paper-reference aggregate tables
paper_outputs/logs/ Command logs from release validation and figure generation
paper_outputs/results/paper_figure_manifest.csv Machine-readable check that all manuscript figure files are present

The repository separates executable release verification from the paper-output archive. The executable path is compact and CPU-friendly so that reviewers can check the repository quickly. The paper-output archive preserves the manuscript-facing result tables and all figure files, including reference aggregate values for neural and attention-based experiments that can be hardware- and package-version-sensitive.

For GitHub README rendering, selected PDF figures are also provided as PNG files in paper_outputs/figures/.

Paper-Chart Reproduction

The repository includes a chart-level reproduction layer that maps submitted paper figures to their plotting scripts and regenerates the manuscript-facing chart files. From the repository root, run:

python reproducibility/paper_charts/standalone_scripts/generate_all_paper_charts.py
python scripts/check_chart_reproduction.py

The regenerated figures are saved in:

reproducibility/paper_charts/generated_figures/

and copied to:

paper_outputs/figures/

The provenance table is available at:

reproducibility/paper_charts/CHART_SOURCE_MAP.csv

Reference Results

Executable benchmark outputs are provided under results/. Manuscript-facing reference tables are archived under paper_outputs/results/. New runs may overwrite files in results/; use paper_outputs/ for the stable paper-output archive.

Limitations

This benchmark uses one indoor environment, six monitor nodes, 36 reference locations, and 9 held-out test locations. It is designed for controlled within-environment comparison and reproducibility. Model rankings, confidence intervals, latency measurements, and deployment-sensitivity results should therefore be interpreted within the evaluated testbed rather than as universal conclusions across buildings, devices, floors, or receiver layouts.

Privacy and Ethics

The raw data distributed here have been anonymized before release. Original device identifiers and AP/BSSID-like identifiers are not included in the public files. The benchmark uses only the signal, frequency, timestamp, AP-label, location, and coordinate fields required for reproducibility. See DATASET_CARD.md and IDENTIFIER_AUDIT.md for scope, limitations, and identifier-sanitization details.

Citation

If you use this dataset, code, or paper-output archive, please cite the associated manuscript. Citation metadata is provided in CITATION.cff.

@article{elrifaee2026passivepr,
  title   = {An End-to-End Pipeline for Passive Wi-Fi Probe Request Fingerprinting-Based Localization with Comparative Evaluation from Classical Retrieval to Attention-Based Models},
  author  = {Elrifaee, Mohamed and Mansour, Ahmed and Zayed, Tarek},
  journal = {Journal of Network and Computer Applications},
  year    = {2026},
  note    = {Under review}
}

License

Unless otherwise noted, this repository is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). This applies to the code, anonymized dataset, result tables, and figure-reproduction materials included in the repository.

About

Anonymized passive Wi-Fi probe-request fingerprinting localization benchmark dataset, code, and reproducibility materials.

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages