More than a Point: Adaptive Affordance Heatmaps as VLM Grounding Interfaces for Robotics

Accepted to CoRL 2026 (Poster)

Previously titled: More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks

Xinyu Shao1,2, TANG YANZHE1,2, Pengwei Xie2, Long ZENG1, Xiu Li1,*
1 Tsinghua Shenzhen International Graduate School, Tsinghua University
2 Huawei Technologies Co., Ltd.
Work completed during the internships of Xinyu Shao and Yanzhe Tang at Huawei Technologies Co., Ltd.
* Corresponding author: li.xiu@sz.tsinghua.edu.cn

Motivation

Teaser Image

Sparse vs. dense grounding. RoboMAP uses adaptive heatmaps to represent region support and multiple feasible candidates.

Framework Overview

Framework Overview

RoboMAP overview: heterogeneous supervision is synthesized into heatmap targets, and the Adaptive Heatmap Decoder produces a dense spatial interface for downstream robotic modules.

Abstract

Hierarchical VLM-based robotic systems often use sparse points or bounding boxes as intermediate spatial representations, which are inadequate for region-level grounding. This is particularly evident in instructions such as “place near the bowl,” which specify a feasible region rather than a single target point. We present RoboMAP, a VLM-based robotic grounding framework that uses adaptive affordance heatmaps as a language-conditioned intermediate interface between spatial reasoning and downstream modules. RoboMAP predicts dense score maps for continuous and object-free target regions, exposing feasible spatial support that downstream segmentation, grasping, and control modules can use beyond sparse prompts alone. To enable scalable training without manual dense labels, RoboMAP synthesizes heatmap supervision from heterogeneous annotations, including points, boxes, and robot trajectory data. Across four spatial grounding benchmarks evaluated with standard point-extraction metrics, RoboMAP obtains the best reported accuracy on three benchmarks while maintaining a 0.04 s grounding-stage forward pass. In 50 real-world dual-arm manipulation trials spanning five tabletop task types, it achieves an 82% success rate using heatmap-guided segmentation and grasp proposal. We further provide qualitative cross-embodiment demonstrations across manipulation and navigation scenarios.

Results and Analysis

RoboMAP generates coherent heatmaps for ambiguous and object-free spatial instructions, preserving feasible regions before downstream modules extract points, masks, or grasps. It achieves the best reported accuracy on three of four spatial grounding benchmarks and a 0.04 s grounding-stage forward pass; downstream perception and execution are excluded from this timing. In zero-shot SimplerEnv evaluation, RoboMAP reaches a 60.5% average success rate. In 50 real-world dual-arm manipulation trials across five tabletop task types, it achieves 41/50 successes (82%), compared with 78% without SAM and 66% for Embodied-R1 with SAM. The qualitative figures below show cross-embodiment manipulation and navigation demonstrations.

Spatial Grounding Results

Table 1: spatial grounding results

Table 1. Accuracy is computed from point predictions; grounding time measures only the VLM module.

SimplerEnv Evaluation

Table 5: SimplerEnv evaluation

Table 5. RoboMAP achieves a 60.5% average success rate across four zero-shot SimplerEnv tasks.

Real-World Manipulation

Table 6: real-world manipulation

Table 6. RoboMAP achieves 41/50 successes (82%) across five relational tabletop tasks.

›

Qualitative Grounding

Figure 4: qualitative grounding comparison

Figure 4. RoboMAP produces coherent heatmaps that preserve feasible regions and task-specific spatial extents.

Cross-Embodiment Demonstrations

Figure 5: cross-embodiment manipulation and navigation demonstrations

Figure 5. Qualitative zero-shot manipulation and navigation examples across multiple embodiments and environments.

Real-World Execution

The GIFs below are qualitative demonstrations and are separate from the 50-trial quantitative real-world evaluation reported above.

Place the right banana onto the right plate

Place the right banana onto the right plate.

Put the bitter gourd in the empty area on the left plate

Put the bitter gourd in the empty area on the left plate.

Move the bottle into the dustpan

Move the bottle into the dustpan.

Put the corn into the bottom basket

Put the corn into the bottom basket.

Place the corn in the upper basket

Place the corn in the upper basket.

Transfer the fish to the middle of the board

Move the fish to the middle of the board.

Lay the knife beside the plate

Lay the knife beside the plate.

Put the pepper in the bottom basket

Put the pepper in the bottom basket.

Move the pepper in the empty area on the left plate

Move the pepper in the empty area on the left plate.

Place the spoon to the left of the plate

Place the spoon to the left of the plate.

Put the spoon down near the plate

Put the spoon down near the plate.

Transfer the trash into the dustpan

Move the trash into the dustpan.

BibTeX

The citation below is provided for readers who want to cite this work.

@misc{shao2026robomap,
  title = {More than a Point: Adaptive Affordance Heatmaps as VLM Grounding Interfaces for Robotics},
  author = {Shao, Xinyu and Tang, Yanzhe and Xie, Pengwei and Zeng, Long and Li, Xiu},
  year = {2026},
  howpublished = {arXiv preprint arXiv:2510.10912},
  note = {Accepted to CoRL 2026, Poster},
  url = {https://arxiv.org/abs/2510.10912}
}