Skip to main navigation Skip to search Skip to main content

Weak calibration cross-fusion framework for multi-modal 3D object detection on unmanned surface vehicles

  • Yong Li*
  • , Dehang Lian
  • , Jialong Du
  • , Dongxu Gao
  • , Xiangrong Xu
  • , Xiang Gong
  • *Corresponding author for this work

    Research output: Contribution to journalArticlepeer-review

    16 Downloads (Pure)

    Abstract

    The field of intelligent transportation on inland waterways is experiencing rapid growth, driven by the global pursuit of enhanced waterway safety, operational efficiency, and environmental sustainability. In real-world autonomous operation scenarios of unmanned surface vehicles (USVs), image-based 2D object detection methods are insufficient to meet the demands of 3D environmental modeling and accurate perception of dynamic objects. Existing 3D perception systems for USVs depend heavily on precise sensor calibration. However, projection offsets between point clouds and images—caused by water surface fluctuations and complex outdoor environments—hinder the practical deployment of these methods. To address these limitations, we propose a weak calibration multi-modal 3D object detection algorithm based on cross-view fusion, termed RCF-Free (Radar-Camera Fusion, Free from precise calibration). Inspired by autonomous driving solutions, we design a Triple-Path Cross-View Fusion module that achieves high-quality cross-view feature fusion without requiring accurate calibration parameters, while simultaneously detecting complete bird’s-eye view (BEV) bounding boxes. We further enhance the spatial layout comprehension of the visual branch through a Mobile Self-Attention Module (MAM) and effectively encode sparse point cloud features in BEV space using a dedicated BEV-Point feature encoder. Additionally, we reconstruct and introduce two water-related 3D object detection datasets, FloW-BEV and WaterScenes-BEV. Experimental results demonstrate that RCF-Free achieves (Formula presented.) scores of 60.5% and 69.3% on the FloW-BEV and WaterScenes-BEV datasets, respectively, showing the effectiveness in water surface object detection. Moreover, on the DAIR-V2X-I dataset for autonomous driving scenarios, the model attains (Formula presented.) scores of 73.3%, 61.2%, and 61.2% across three task difficulty levels, illustrating strong cross-domain generalization capability.

    Original languageEnglish
    Article number867
    Number of pages28
    JournalJournal of Marine Science and Engineering
    Volume14
    Issue number9
    DOIs
    Publication statusPublished - 6 May 2026

    UN SDGs

    This output contributes to the following UN Sustainable Development Goals (SDGs)

    1. SDG 11 - Sustainable Cities and Communities
      SDG 11 Sustainable Cities and Communities

    Keywords

    • 3D object detection
    • bird’s-eye view
    • separable attention
    • unmanned surface vehicles
    • weak calibration

    Fingerprint

    Dive into the research topics of 'Weak calibration cross-fusion framework for multi-modal 3D object detection on unmanned surface vehicles'. Together they form a unique fingerprint.

    Cite this