This project aims to evaluate the performance of various LLMs models agents in a colaborative rescue task
Paper: arXiv:2508.14635
For installing the dependencies of the project, run:
pip install -r requirements.txt
After that, make sure you have Ollama set up. If not, download and run the program from:
If you wish to run Ollama with docker, follow this article:
https://ollama.com/blog/ollama-is-now-available-as-an-official-docker-image
Finally, after having Ollama set up, download the following models:
- cogito:14b
- qwen3:14b
- qwen2.5:14b
- mistral-small:24b
- cogito:32b
- qwen3:32b
- qwen2.5:32b
- qwen2.5-coder:32b
You can download them using:
ollama pull cogito:14b
ollama pull qwen2.5:14b
# ... and so onTo run experiments, use the following command:
python main.py --models <model_codes> --maps <map_codes>Use the following short codes for the --models argument:
c14- Cogito:14bq14- Qwen2.5:14bm24- Mistral-small:24bc32- Cogito:32bq32- Qwen2.5:32bq32c- Qwen2.5-coder:32bq3-14- Qwen3:14bq3-32- Qwen3:32bheuristic- Baseline heuristic agent (non-LLM)
Use the following map codes for the --maps argument:
map1throughmap8- Different problem configurations with varying room layouts and victim distributions
Run a single model on a single map:
python main.py --models q32 --maps map1Run multiple models on multiple maps:
python main.py --models c14 q14 m24 --maps map1 map2 map3Run LLM models and compare with heuristic baseline:
python main.py --models heuristic q32 c32 --maps map1 map2agentRescue/
├── src/
│ ├── agents/ # Agent implementations
│ │ ├── conversational_agent.py # LLM-based cooperative agent
│ │ └── heuristic_agent.py # Baseline heuristic agent
│ ├── graphEnv/ # Environment and graph representation
│ │ ├── environment.py # Room graph and simulation
│ │ └── env_elements.py # Victim and Room classes
│ ├── loggers/ # Experiment logging utilities
│ │ ├── experiment_logger.py # CSV experiment logger
│ │ └── agent_logger.py # Individual run logger
│ └── main_*.py # Main execution scripts
├── problem_instances_data/ # Test scenarios (maps, agents, victims)
├── experiments/ # Output folder for experiment results
├── plotted_maps/ # Generated visualizations
└── main.py # Entry point
After running experiments, results are saved in the experiments/ directory:
-
CSV files: Contain quantitative metrics such as:
- Number of steps to completion
- Victims saved (urgent vs non-urgent)
- Agent coordination efficiency
- Temperature variations (0.0 and 0.5)
-
Log folders: Detailed execution logs including:
- Agent communication history
- Step-by-step environment visualizations
- Decisions made by agents
Each experiment run creates a timestamped folder with all relevant data for analysis.
Agents must navigate through connected rooms to rescue victims who need water, food, and/or medicine. Key challenges include:
- Coordination: Multiple agents must coordinate to avoid redundant work
- Planning: Efficient path planning to minimize rescue time
- Urgency awareness: Prioritizing urgent victims
- Resource constraints: Limited supplies for each agent
- Communication: Agents share information to optimize strategy
The project evaluates how well different LLM models can handle these multi-agent collaborative planning tasks compared to traditional heuristic approaches.
