基于跨任务交互与大核可分离注意力的交通环境感知算法

Traffic Environment Perception Algorithm Based on Cross-Task Interaction and Large Kernel Separable Attention

  • 摘要: 在基于车载摄像头的视觉感知任务中,多任务协同已成为智能驾驶感知的重要技术路径,但如何在单次神经网络推理过程中高效提取多任务特征并实现特征互补仍具挑战. 为此,面向目标检测、道路可行驶区域提取和车道线分割三项典型交通环境感知任务,提出一种多任务交通环境感知算法MtTEPNet. 自主设计了动态多尺度特征融合神经网络模块来整合多层级语义信息,提升模型对复杂交通场景信息的捕捉能力;提出一种跨任务交互注意力机制来建模目标检测与图像分割任务间的全局相关性,实现任务特征间的互补增强并减少相互干扰;构建了一种大核可分离注意力来提取图像特征的像素级长程上下文关系,从而提升对细长结构和边界目标的特征表征能力. 在BDD100K数据集上的实验表明,所提算法在道路可行驶区域提取任务上的平均交并比达到92.8%、车道线分割准确度为87.4%、目标检测平均精度为79.9%,参数量和计算量仅为8.3 M和13.8 GFLOPs,推理速度达到38 帧/s,综合性能优于所对比的主流多任务交通环境感知算法.

     

    Abstract: In visual perception tasks based on vehicle-mounted cameras, multi-task collaboration has become an important technical paradigm for intelligent driving. However, efficiently extracting multi-task features and achieving feature complementarity in a single inference pass remains challenging. To address this problem, MtTEPNet, a multi-task traffic environment perception algorithm, was proposed for object detection, road drivable area segmentation, and lane line segmentation.First, a dynamic multi-scale feature fusion neural module was designed to integrate hierarchical semantic information and enhance the model’s capability to capture information from complex traffic scenes.Then, a cross-task interactive attention mechanism was introduced to model global correlations between object detection and image segmentation tasks, achieving complementary feature enhancement and reducing mutual interference.Finally, a large-Kernel separable attention module was constructed to extract pixel-level long-range contextual relationships from image features, thereby enhancing feature representation for slender structures and boundary objects.Experiments on the BDD100K dataset show that the proposed algorithm achieves a mean Intersection-over-Union of 92.8% for drivable area extraction, an accuracy of 87.4% for lane line segmentation, and a mean average precision of 79.9% for object detection, with only 8.3 M parameters, 13.8 GFLOPs, and an inference speed of 38 frames per second. The proposed method outperforms mainstream multi-task traffic environment perception algorithms in overall performance.

     

/

返回文章
返回
Baidu
map