Abstract:
To address the failure of visual grasping strategy of robotic arms caused by non-rigid docking of bipedal robots, a semantic-aware token selection framework SATS (semantic-aware token selection) for robust control was proposed. This framework introduced the segmentation model SAM2 (segment anything model 2) as an offline teacher to implement explicit semantic supervision, guiding the online selector to focus on foreground features. By performing physical hard pruning to retain only the Top-K core features to isolate background noise, a high signal-to-noise ratio input in the representation space was ensured. Additionally, to handle target recognition ambiguity and end-effector trajectory oscillation caused by pose deviation induced by non-rigid docking, and the resulting accidental uncertainty grasping problem, the adaptive compensation mechanism was introduced. This mechanism linearly mapped the visual confidence defined by selective activation statistics to control guidance weights, thereby smoothing the command noise and enhancing the stability of the dynamic grasping process. The experiment shows that the SATS-RTC framework significantly increased the dynamic capture success rate to 82.5%, proving that the collaborative design of explicit semantic supervision and dynamic weight adjustment is the key path to achieving robust mobile capture.