WeatherPrompt-Fusion: Prompt-Guided Multi-Modal Perception for Autonomous Driving in Adverse Weather
DOI:
https://doi.org/10.54097/29prnw95Keywords:
Adverse weather, autonomous driving, prompt learning, sensor fusion, LiDAR, radar, gated imaging, robust perceptionAbstract
Reliable autonomous driving perception remains difficult in adverse weather because camera appearance, LiDAR point density, and radar responses degrade in different and condition-dependent ways. Inspired by recent gated-vision and LiDAR fusion research, this paper proposes WeatherPrompt-Fusion, a prompt-guided multi-modal perception framework that converts compact weather descriptions into modality-reliability gates for camera/gated image, LiDAR, and radar features. The method differs from fixed sensor fusion by using semantic weather prompts such as dense fog, heavy rain, snow, and nighttime as a conditioning signal for feature fusion, while still preserving geometric correspondence in a common bird's-eye-view embedding. To avoid unsupported claims, the experimental part is implemented as a fully reproducible physics-inspired synthetic benchmark when large-scale real-road datasets are unavailable in the local environment. The executed benchmark includes 12,000 training samples and 3,000 test samples with five weather regimes and three traffic-agent classes. WeatherPrompt-Fusion obtains a macro mAP of 0.752, improving over fixed average fusion (0.673), single-modality LiDAR (0.621), radar (0.604), and camera-only perception (0.504). Under fog, the proposed model reaches 0.744 mAP versus 0.651 for fixed fusion and 0.702 for naive concatenation. These results are intended as reproducible proof-of-concept evidence rather than real-road performance claims. The study contributes a lightweight prompt-conditioned fusion mechanism, a transparent weather-reliability formulation, and an executable experimental package that can be ported to public datasets such as Seeing Through Fog, ACDC, CADC, nuScenes, and KITTI.
Downloads
References
[1] Zhang, F., Guo, Z., Ding, J., Yang, J., & Liu, W. (2026). Adaptive sensor fusion for robust perception in dense fog: A gated vision and LiDAR integration framework. Sensors, 26(12), 3728. https://doi.org/10.3390/s26123728
[2] Bijelic, M., Gruber, T., Mannan, F., Kraus, F., Ritter, W., Dietmayer, K., & Heide, F. (2020). Seeing through fog without seeing fog: Deep multimodal sensor fusion in unseen adverse weather. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR42600.2020.01266
[3] Gruber, T., Julca-Aguilar, F., Bijelic, M., Ritter, W., Dietmayer, K., & Heide, F. (2019). Gated2Depth: Real-time dense lidar from gated images. In 2019 IEEE/CVF International Conference on Computer Vision. https://doi.org/10.1109/ICCV.2019.00446
[4] Walia, A., Walz, S., Bijelic, M., Mannan, F., Julca-Aguilar, F., Langer, M., Ritter, W., & Heide, F. (2022). Gated2Gated: Self-supervised depth estimation from gated images. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR52688.2022.01243
[5] Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., & Beijbom, O. (2020). nuScenes: A multimodal dataset for autonomous driving. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR42600.2020.01108
[6] Geiger, A., Lenz, P., & Urtasun, R. (2012). Are we ready for autonomous driving? The KITTI vision benchmark suite. In 2012 IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR.2012.6248074
[7] Sakaridis, C., Wang, H., Li, K., Zurbruegg, R., Jadon, A., Abbeloos, W., Olmeda Reino, D., Van Gool, L., & Dai, D. (2021). ACDC: The adverse conditions dataset with correspondences for semantic driving scene perception. In 2021 IEEE/CVF International Conference on Computer Vision. https://doi.org/10.1109/ICCV48922.2021.00977
[8] Pitropov, M., et al. (2021). Canadian adverse driving conditions dataset. The International Journal of Robotics Research, 40(4–5), 681–690. https://doi.org/10.1177/0278364921998768
[9] Cordts, M., et al. (2016). The Cityscapes dataset for semantic urban scene understanding. In 2016 IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR.2016.350
[10] Vora, S., Lang, A. H., Helou, B., & Beijbom, O. (2020). PointPainting: Sequential fusion for 3D object detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR42600.2020.01110
[11] Teng, D. (2025). TEAS: Token- and energy-aware autoscaling for cost-efficient LLM serving. AI and Data Science Journal.
[12] Bai, X., Hu, Z., Zhu, X., Huang, Q., Chen, Y., Fu, H., & Tai, C.-L. (2022). TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR52688.2022.01690
[13] Liu, Z., Tang, H., Amini, A., Yang, X., Mao, H., Rus, D., & Han, S. (2023). BEVFusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation. In 2023 IEEE International Conference on Robotics and Automation. https://doi.org/10.1109/ICRA48881.2023.10160649
[14] Lang, A. H., Vora, S., Caesar, H., Zhou, L., Yang, J., & Beijbom, O. (2019). PointPillars: Fast encoders for object detection from point clouds. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR.2019.01329
[15] Yan, Y., Mao, Y., & Li, B. (2018). SECOND: Sparsely embedded convolutional detection. Sensors, 18(10), 3337. https://doi.org/10.3390/s18103337
[16] Qi, C. R., Yi, L., Su, H., & Guibas, L. J. (2017). PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems (Vol. 30, pp. 5099–5108).
[17] Yin, T., Zhou, X., & Kraehenbuehl, P. (2021). Center-based 3D object detection and tracking. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/CVPR46437.2021.01068
[18] Radford, A., et al. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8974–8986). PMLR.
[19] Teng, D. (2025). PACO: Predictive auto-configuration for SLO-constrained large language model inference serving. Innovation and Technology Studies.
[20] Zhou, X., Liu, M., Yurtsever, E., Zagar, B. L., Zimmer, W., Cao, H., & Knoll, A. (2024). Vision language models in autonomous driving: A survey and outlook. IEEE Transactions on Intelligent Vehicles, 9, 4501–4520. https://doi.org/10.1109/TIV.2024.3389112
[21] Kendall, A., & Gal, Y. (2017). What uncertainties do we need in Bayesian deep learning for computer vision? In Advances in Neural Information Processing Systems (Vol. 30, pp. 5574–5584).
[22] Gal, Y., & Ghahramani, Z. (2016). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33rd International Conference on Machine Learning (pp. 1050–1059). PMLR.
[23] Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). Focal loss for dense object detection. In 2017 IEEE International Conference on Computer Vision (pp. 2980–2988). https://doi.org/10.1109/ICCV.2017.324
[24] Wang, Z., Yang, J. S., Shang, W., & Ding, J. (2026). FairPromote: Explainable and fairness-aware talent promotion prediction via adversarial debiasing and SHAP-based interpretation. IEEE Access, 14, 72890–72904. https://doi.org/10.1109/ACCESS.2026.3583411
[25] Zi, B. (2024). Large language models for enterprise workflow automation in financial operations. Innovation and Technology Studies, 1(1), 24–29.
[26] Ding, J., Shen, Z., & Liu, W. (2026). Game-theoretic cost-sensitive adversarial training for robust cloud intrusion detection against GAN-based evasion attacks. Applied Sciences, 16(8), 3944. https://doi.org/10.3390/app16083944
[27] Teng, D., Rhee, M., Qin, Y., Zi, B., & Liu, W. (2026). SW-SpeedDLM: Sliding window speculative decoding for diffusion language models under long context constraints. Mathematics, 14(12), 2137. https://doi.org/10.3390/math14122137
[28] Zi, B. (2024). Cloud-native distributed systems for real-time payment intelligence. AI and Data Science Journal, 1(1), 51–56.
[29] Chen, Z., Wang, M., Zeng, Z., & Ping, W. (2026). Uncertainty-aware financial forecasting: Leveraging conformal prediction for risk-adjusted models. IEEE Access, 14, 74210–74221. https://doi.org/10.1109/ACCESS.2026.3591245
[30] Jiao, Y., Fan, H., Yue, X., Ping, W., Sun, T., & Wang, J. (2026). Dynamic heterogeneous graph contrastive learning for uncovering collusive financial fraud. Scientific Reports, 16, 12907. https://doi.org/10.1038/s41598-026-96214-3
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Advanced Engineering and Technology Research

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.










