CLFT: Camera-LiDAR Fusion Transformer for Semantic Segmentation in Autonomous Driving

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Critical research about camera-and-LiDAR-based semantic object segmentation for autonomous driving significantly benefited from the recent development of deep learning. Specifically, the vision transformer is the novel ground-breaker that successfully brought the multi-head-attention mechanism to computer vision applications. Therefore, we propose a vision-transformer-based network to carry out camera-LiDAR fusion for semantic segmentation applied to autonomous driving. Our proposal uses the novel progressive-assemble strategy of vision transformers on a double-direction network and then integrates the results in a cross-fusion strategy over the transformer decoder layers. Unlike other works in the literature, our camera-LiDAR fusion transformers have been evaluated in challenging conditions like rain and low illumination, showing robust performance. The paper reports the segmentation results over the vehicle and human classes in different modalities: camera-only, LiDAR-only, and camera-LiDAR fusion. We perform coherent controlled benchmark experiments of CLFT against other networks that are also designed for semantic segmentation. The experiments aim to evaluate the performance of CLFT independently from two perspectives: multimodal sensor fusion and backbone architectures. The quantitative assessments show our CLFT networks yield an improvement of up to 10\% for challenging dark-wet conditions when comparing with Fully-Convolutional-Neural-Network-based (FCN) camera-LiDAR fusion neural network. Contrasting to the network with transformer backbone but using single modality input, the all-around improvement is 5-10\%.

Related collections

Author and article information

Journal

Publication date Created: 27 April 2024

Article

ArXiV ID: 2404.17793

SO-VID: cb5ec343-7721-44f6-95a8-485557b62961

License:

http://creativecommons.org/licenses/by/4.0/

History

Custom metadata

Comments Submitted to IEEE Transactions on Intelligent Vehicles

Categories cs.CV cs.RO

ScienceOpen disciplines: Computer vision & Pattern recognition,Robotics

Data availability:

ScienceOpen disciplines: Computer vision & Pattern recognition, Robotics

CLFT: Camera-LiDAR Fusion Transformer for Semantic Segmentation in Autonomous Driving

Read this article at

Abstract

Related collections

Semantic Knowledge Base

Author and article information

Journal

Article

History

Custom metadata

Comments

Comment on this article

Similar content 92