Authors: Elena Ryumina, Alexandr Axyonov, Dmitry Sysoev, Timur Abdulkadirov, Kirill Almetov, Yulia Morozova, Dmitry Ryumin
This repository contains the code for Team LEYA's submission to the 10th ABAW Competition for the ambivalence / hesitation recognition task.
The project implements a multimodal classification pipeline that combines four modalities:
- face
- audio
- text
- scene
The current training and inference pipeline is built around pre-extracted unimodal features and a multimodal fusion model with an optional prototype-based auxiliary head.
Related paper:
- Feature exporters for face, audio, text, and scene modalities
- A multimodal dataset loader over serialized feature artifacts
- Fusion training with:
- single-run training
- grid search
- exhaustive search
- Optuna-based hyperparameter search
- Challenge feature preparation and challenge/evaluation inference scripts
- Support for single-checkpoint inference and ensemble inference
.
+-- assets/ # checkpoints and auxiliary model files
+-- data/ # local CSVs and challenge metadata
+-- features/ # extracted modality artifacts
+-- results/ # training and inference outputs
+-- scripts/
| +-- prepare_challenge_csv.py
| +-- run_challenge_feature_prepare.py
| '-- run_challenge_inference.py
+-- src/
| +-- data_loading/
| +-- exporters/
| +-- models/
| +-- utils/
| '-- train.py
+-- config.toml
+-- config.challenge.toml
+-- search_params.toml
'-- main.py
Python 3.12 was used in the current setup.
Install dependencies:
pip install -r requirements.txtNotes:
requirements.txtis pinned to CUDA 12.4 PyTorch wheels.- The repository assumes local checkpoints and local datasets are already available.
- Some modality checkpoints and feature assets are expected to be taken from the related modality-specific branches of the same project. In practice, face, audio, text, and scene assets should be searched for in their corresponding branches and prepared locally before running the full pipeline.
There are two main configs:
- config.toml: training / local evaluation pipeline
- config.challenge.toml: challenge feature preparation and challenge inference
The main sections are:
datasets.*: dataset roots, CSV templates, split-specific pathsdataloader: batch size, workers, shuffling, prepare-only modesearch: training mode (none,greedy,exhaustive,optuna)model: multimodal fusion architecturetraining: optimizer, scheduler, early stopping, prototype-loss weightsmultimodal: active modalities and artifact rootface_export,audio_export,text_export,scene_export: modality-specific export settings
The pipeline operates on serialized feature artifacts stored under features/.
Expected artifact layout:
features/
+-- face/<artifact_tag>/<split>.pkl
+-- audio/<artifact_tag>/<split>.pkl
+-- text/<artifact_tag>/<split>.pkl
'-- scene/<artifact_tag>/<split>.pkl
Each artifact is a single pickle file per split. The multimodal loader reads these artifacts and joins modalities by sample_id.
If a required artifact is missing, main.py will attempt to run the corresponding exporter automatically.
Run the main training pipeline:
python main.pyBehavior depends on search.type in config.toml:
none: one training rungreedy: greedy hyperparameter searchexhaustive: full grid searchoptuna: Optuna-based search
Outputs are written to:
results/results_multimodal_pipeline_<timestamp>/
Typical contents:
session_log.txtconfig_copy.tomloverrides.txtfusion_metrics.jsoncheckpoints/iftraining.save_checkpoints = true
Challenge preparation is handled by:
Run:
python scripts/run_challenge_feature_prepare.pyWhat it does:
- Builds the challenge CSV from the official split text file
- Validates or mirrors precomputed audio features
- Runs face, text, and scene exporters for the challenge split
Important:
- This script is configured through module-level constants inside the file, not command-line arguments.
- Before running it, check:
CONFIG_PATHSPLITRUN_FACERUN_TEXTRUN_SCENEAUDIO_PRECOMPUTED_SOURCE
Inference is handled by scripts/run_challenge_inference.py.
Run:
python scripts/run_challenge_inference.pyThis script supports two modes:
challenge_submiteval_metrics
The mode is selected by editing RUN_MODE inside the script.
Set:
RUN_MODE = "challenge_submit"The script writes submission files under:
results/challenge_submissions/<tag>_<suffix>_<timestamp>/
Generated files:
no_probabilities/trial-0.txtno_probabilities/trial-0.csvwith_probabilities/trial-0.txtwith_probabilities/trial-0.csvwith_probabilities_hard/trial-0.txtwith_probabilities_hard/trial-0.csvsubmission_meta.json
Set:
RUN_MODE = "eval_metrics"The script evaluates one checkpoint or an ensemble on dev / test and writes outputs to:
results/eval_inference/<tag>_<suffix>_<timestamp>/
Generated files:
dev_predictions.csvtest_predictions.csveval_metrics.json
The inference script supports:
- one checkpoint via
CHECKPOINT_PATH - multiple checkpoints via
CHECKPOINT_PATHS
If CHECKPOINT_PATHS is non-empty, the script performs probability averaging across models.
Important:
- Like the challenge feature preparation script, this script is configured through module-level constants.
- Before running it, check:
CONFIG_PATHEVAL_CONFIG_PATHCHECKPOINT_PATHCHECKPOINT_PATHSRUN_MODERUN_TAGEVAL_TAG
The current multimodal pipeline works in two stages:
- Train or load unimodal feature extractors for face, audio, text, and scene
- Train a multimodal fusion model over the extracted modality embeddings
Supported fusion backbones currently include:
exchange_transformervideoformerconcat_mlpattnclass_weighted
The main working configuration in this repository is the transformer-based multimodal fusion model defined in src/models/fusion_model.py.
Prototype support is optional and controlled by:
[model]
use_prototypes = trueWhen enabled, the prototype branch contributes an auxiliary training loss. Final predictions still come from the main classifier logits.
Hyperparameter search is configured in search_params.toml.
The repository supports:
- standard grid search
- exhaustive search
- Optuna
Optuna can be configured to:
- persist trials to SQLite
- continue an existing study
- evaluate multiple random seeds for the same hyperparameter configuration
- disable checkpoint saving during search
- Paths in the provided configs are local and should be adapted to your machine.
config.tomlandconfig.challenge.tomlare the source of truth for dataset roots and artifact locations.- Some local folders such as
features/andresults/are intentionally not versioned. - The repository is currently optimized for local experimentation rather than packaging as a reusable library.
python main.pypython scripts/run_challenge_feature_prepare.py- Open scripts/run_challenge_inference.py
- Set
RUN_MODE = "challenge_submit" - Set checkpoint path(s)
- Run:
python scripts/run_challenge_inference.py- Open scripts/run_challenge_inference.py
- Set
RUN_MODE = "eval_metrics" - Set checkpoint path(s)
- Run:
python scripts/run_challenge_inference.py