Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DGSF: A DeepSets-Guided Scheduling Framework for Single-Machine Total Tardiness

Code and reproducibility materials for the DeepSets-Guided Scheduling Framework (DGSF) experiments on the single-machine total tardiness problem.

Release

This is the production-ready v1.0.0 reproducibility release. The release tag identifies the complete code and documentation snapshot used for the public artifact package.

The model predicts a static job-priority vector for a non-preemptive single-machine scheduling instance. The predicted sequence can then be improved with local swap refinement and continuous-time MIP post-processing.

Contents

  • Code/1_data_processing/: instance generation, feature construction, and time-indexed MIP reference-solution scripts.
  • Code/2_dgsf_main/: DeepSets training, ML inference with swap refinement, and continuous-time MIP post-processing.
  • Code/3_evaluation/: dispatching-rule baselines, schedule evaluation, and feature-importance analysis.
  • Data/: product tables, trained checkpoints, and external benchmark data.
  • Results/: supplied or regenerated experiment outputs (Tables/, F1/, F2/, F3/, F4/, and F5/).
  • environment.yml: conda environment used for the reproducibility workflow.
  • REPRODUCIBILITY.md: step-by-step setup and command-line workflow.
  • docs/artifact_manifest.md: expected external artifacts and where to place them.

Large .xlsx, .csv, and .pth artifacts are not committed. Download the project data archive, preserve its folder structure, and merge it into this repository. The exact layout and filenames are listed in docs/artifact_manifest.md.

Artifact archive: https://drive.google.com/drive/folders/1Lo8WRZabBUxGNA0nOD50TMavkKyhDwHP

Setup

git clone https://github.com/dz5430/dgsf-single-machine.git
cd dgsf-single-machine
conda env create -f environment.yml
conda activate scheduling_env

Alternatively, in an existing Python 3.10 environment:

pip install -r requirements.txt

The MIP scripts use Pyomo with Gurobi. Results were produced with Gurobi 10.0.1. Install Gurobi separately, activate a valid license, and verify access:

python -c "import pyomo.environ as pyo; print(pyo.SolverFactory('gurobi').available(False))"

The command should print True.

Instance Types, Facilities, and Filename Tokens

The manuscript studies two instance types and five facility configurations.

Token Meaning
theta_max_6Itau Type A instances, with all jobs released at time zero.
theta_0max_40tau_avg Type B instances, with staggered release times.
F1 Base facility with integer processing times (time resolution 1.0).
F2, F3 Alternative facilities used in the generalization study.
F4, F5 F1 evaluated at time resolutions 0.5 and 0.1, respectively.

Artifact filenames retain the identifiers used to generate the reported results. Dev3 identifies the data-generation configuration, _u4 identifies a processed workbook containing model features and reference-solution columns, 50k denotes 50,000 training instances, and dev9_lean identifies the DeepSets model architecture. The _dgsf and _mip suffixes distinguish DGSF outputs and their MIP-postprocessed counterparts. These filenames are retained so that the repository paths correspond directly to the accompanying data archive.

Data Layout

Place external artifacts as follows:

Data/
    Facility Products/
    Trained Models/
Results/
    Tables/
    F1/
        input/
        output/
            F1_DGSF/
            F1_Recursive/
            Time resolution/
        Max Tardiness Evaluation/
    F2/
        input/
        output/
    F3/
        input/
        output/
    F4/
        input/
        output/
    F5/
        input/
        output/

The expected filenames are listed in docs/artifact_manifest.md.

Run the Pipeline

The scripts can be run from the repository root.

Run the supplied Dev9-Lean model with local-swap refinement:

python Code/2_dgsf_main/evaluate_sms_model.py \
  --input Results/F1/input/Dev3_singlemachine_instances_100_theta_max_6Itau_u4.xlsx \
  --model "Data/Trained Models/30_theta_max_6Itau_Dev3_50k_dev9_lean.pth" \
  --architecture dev9_lean \
  --device cpu \
  --output Results/F1/output/F1_DGSF/Dev3_singlemachine_instances_100_theta_max_6Itau_u4_dgsf.xlsx

Run continuous-time MIP post-processing:

python Code/2_dgsf_main/Solver_MIP_ct_post.py \
  --input Results/F1/output/F1_DGSF/Dev3_singlemachine_instances_100_theta_max_6Itau_u4_dgsf.xlsx \
  --output Results/F1/output/F1_DGSF/Dev3_singlemachine_instances_100_theta_max_6Itau_u4_dgsf_mip.xlsx \
  --time-limit 60

Evaluate a schedule against the reference solution:

python Code/3_evaluation/Schedule_evaluation.py \
  --input Results/F1/output/F1_DGSF/Dev3_singlemachine_instances_100_theta_max_6Itau_u4_dgsf_mip.xlsx \
  --method-obj-col tardiness_dgsf_mip \
  --ref-obj-col tardiness_dtime \
  --pred-rank-col rank_dgsf_mip \
  --ref-rank-col rank_vector_dtime

Evaluate static dispatching-rule baselines:

python Code/3_evaluation/Dispatching_heuristics.py \
  --input Results/F1/input/Dev3_singlemachine_instances_100_theta_max_6Itau_u4.xlsx \
  --output Results/F1/output/F1_Recursive/Dev3_singlemachine_instances_100_theta_max_6Itau_u4_heuristics.xlsx

Metric

Reported normalized tardiness gaps use the conventional optimality-gap definition:

gap (%) = 100 * (TT_method - TT_opt) / TT_opt

All reported benchmark summaries use instances with positive TT_opt.

Full Reproducibility Notes

See REPRODUCIBILITY.md for the complete workflow, including feature generation, reference-solution generation, optional retraining, and post-processing.

About

This work presents a hybrid machine learning and optimization framework for single machine non-preemptive total tardiness scheudling problems.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages