Task-driven object detection and segmentation
- What object should i use? - Task driven object detection
- TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
- VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation
- TOIST: Task oriented instance segmentation transformer with noun-pronoun distillation
- CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection
Affordance classification
- Visual object-action recognition: Inferring object affordances from human demonstration
- Learning visual object categories for robot affordance prediction
- High-level object affordance recognition
- Functional object descriptors for human activity modeling
- EGO-TOPO: Environment Affordances from Egocentric Video
Affordance detection and segmentation
- AffordanceNet: An end-to-end deep learning approach for object affordance detection
- Bayesian deep learning for affordance segmentation in images
- Learning affordance segmentation: An investigative study
- Are standard object segmentation models sufficient for learning affordance segmentation?
- Object affordance detection with boundary-preserving network for robotic manipulation task
- A new semantic edge aware network for object affordance detection
- Object-based affordances detection with convolutional neural networks and dense conditional random fields
- Weakly supervised affordance detection
- Adosmnet: a novel visual affordance detection network with object shape mask guided feature encoders
- Detecting object affordances with convolutional neural networks
- FPHA-Afford: A domain-specific benchmark dataset for occluded object affordance estimation in human-object-robot interaction
- Affordance segmentation of hand-occluded containers from exocentric images
- Visual affordance detection using an efficient attention convolutional neural network
- Multi-scale fusion and global semantic encoding for affordance detection
- Object affordance detection with relationship-aware network
- Strap: Structured object affordance segmentation with point supervision
- Segmenting object affordances: Reproducibility and sensitivity to scale
Affordance grounding
- Understanding 3d object interaction from a single image
- Leverage Interactive Affinity for Affordance Learning
- Locate: Localize and transfer object parts for weakly supervised affordance grounding
- One-shot transfer of affordance regions? affcorrs!
- Demo2vec: Reasoning object affordances from online video
- Grounded human-object interaction hotspots from video
- Oval-prompt: Open-vocabulary affordance localization for robot manipulation through LLM affordance-grounding
- What does clip know about peeling a banana?
- One-shot open affordance learning with foundation model
- Learning affordance grounding from exocentric image
- Affordancellm: Grounding affordance from vision language model
- Knowledge enhanced bottom-up affordance grounding for robotic interaction
Grasping detection
- Deep Learning for Detecting Robotic Grasps
- Real-Time Grasp Detection Using Convolutional Neural Networks
- Robotic Grasp Detection using Deep Convolutional Neural Networks
- GraspNet: An Efficient Convolutional Neural Network for Real-time Grasp Detection for Low-powered Devices
- Real-world Multi-object, Multi-grasp Detection
- ROI-based Robotic Grasp Detection for Object Overlapping Scenes
- End-to-end Trainable Deep Neural Network for Robotic Grasp Detection and Semantic Segmentation from RGB
- Jacquard: A Large Scale Dataset for Robotic Grasp Detection
If you find an error, if you want to suggest a new feature or a change, you can use the issues tab to raise an issue with the appropriate label.
Complete and full updates can be found in CHANGELOG.md. The file follows the guidelines of https://keepachangelog.com/en/1.1.0/.
T. Apicella, A. Xompero, A. Cavallaro, Visual Affordance Prediction: Survey and Reproducibility, arXiv:2505.05074 [cs.CV], 2025.
@misc{apicella2025visual,
title={Visual Affordance Prediction: Survey and Reproducibility},
author={Tommaso Apicella and Alessio Xompero and Andrea Cavallaro},
year={2025},
eprint={2505.05074},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2505.05074},
}
If you have any further enquiries, question, or comments, or you would like to file a bug report or a feature request, please use the Github issue tracker.
This work is licensed under the MIT License. To view a copy of this license, see LICENSE.