Publication
2026
Yuze Jiang, Ehsan Javanmardi, Bo Qian, Manabu Tsukada, Hiroshi Esaki, "CacheComm: A Motion-Aware Cache System for Cooperative Perception", In: 2026 IEEE 104th Vehicular Technology Conference (VTC2026-Fall), Boston, MA, USA, 2026.Proceedings Article | Abstract | BibTeX
@inproceedings{Jiang2026b,
title = {CacheComm: A Motion-Aware Cache System for Cooperative Perception},
author = {Yuze Jiang and Ehsan Javanmardi and Bo Qian and Manabu Tsukada and Hiroshi Esaki},
year = {2026},
date = {2026-09-06},
booktitle = {2026 IEEE 104th Vehicular Technology Conference (VTC2026-Fall)},
address = {Boston, MA, USA},
abstract = {Cooperative perception enables connected and automated vehicles (CAVs) to overcome the inherent limitations of single-agent sensing by sharing perceptual information through vehicular communications. However, existing intermediate-fusion approaches treat each transmission cycle independently, re-transmitting features for the entire scene regardless of how much has actually changed, resulting in significant bandwidth waste. We propose CacheComm, a motion-aware caching framework that exploits temporal redundancy in cooperative perception by utilizing an RSU as an intermediate feature cache. The RSU maintains an up-to-date feature map of its coverage area and uses lightweight motion detection to identify scene changes. It selectively updates the cached feature map and serve the cache to CAVs based on the motion map to save communication resources. Experiments show that our system saves 10% V2X communication load under sparse CAV and limited RSU range, and up to 77% with full RSU coverage. Our method can be easily expanded to cover more area to achieve more reduction on communication cost. Saved bandwidth can be used to transmit more critical information or further increase the perception range for a CAV.},
keywords = {},
pubstate = {published},
tppubtype = {inproceedings}
}
Cooperative perception enables connected and automated vehicles (CAVs) to overcome the inherent limitations of single-agent sensing by sharing perceptual information through vehicular communications. However, existing intermediate-fusion approaches treat each transmission cycle independently, re-transmitting features for the entire scene regardless of how much has actually changed, resulting in significant bandwidth waste. We propose CacheComm, a motion-aware caching framework that exploits temporal redundancy in cooperative perception by utilizing an RSU as an intermediate feature cache. The RSU maintains an up-to-date feature map of its coverage area and uses lightweight motion detection to identify scene changes. It selectively updates the cached feature map and serve the cache to CAVs based on the motion map to save communication resources. Experiments show that our system saves 10% V2X communication load under sparse CAV and limited RSU range, and up to 77% with full RSU coverage. Our method can be easily expanded to cover more area to achieve more reduction on communication cost. Saved bandwidth can be used to transmit more critical information or further increase the perception range for a CAV.
Hanlin Wu, Pengfei Lin, Ehsan Javanmardi, Naren Bao, Bo Qian, Hao Si, Manabu Tsukada, "A Synthetic Benchmark for Collaborative 3D Semantic Occupancy Prediction in V2X-Enabled Autonomous Driving ", In: IEEE International Conference on Robotics & Automation (ICRA 2026), Vienna, Austria, 2026.Proceedings Article | Abstract | BibTeX | Links:
@inproceedings{Wu2026,
title = {A Synthetic Benchmark for Collaborative 3D Semantic Occupancy Prediction in V2X-Enabled Autonomous Driving },
author = {Hanlin Wu and Pengfei Lin and Ehsan Javanmardi and Naren Bao and Bo Qian and Hao Si and Manabu Tsukada},
url = {https://arxiv.org/abs/2506.17004
https://github.com/tlab-wide/Co3SOP
https://youtu.be/VlrqcdkMyn8},
year = {2026},
date = {2026-06-01},
urldate = {2026-06-01},
booktitle = {IEEE International Conference on Robotics & Automation (ICRA 2026)},
address = {Vienna, Austria},
abstract = {3D semantic occupancy prediction is an emerging perception paradigm in autonomous driving, providing a voxel-level representation of both geometric details and semantic categories. However, its effectiveness is inherently constrained in single-vehicle setups by occlusions, restricted sensor range, and narrow viewpoints. To address these limitations, collaborative perception enables the exchange of complementary information, thereby enhancing
the completeness and accuracy of predictions. Despite its potential, research on collaborative 3D semantic occupancy prediction is hindered by the lack of dedicated datasets. To bridge this gap, we design a high-resolution semantic voxel sensor in CARLA to produce dense and comprehensive annotations. We further develop a baseline model that performs inter-agent feature fusion via spatial alignment and attention aggregation. In addition, we establish
benchmarks with varying prediction ranges designed to systematically assess the impact of spatial extent on collaborative prediction. Experimental results demonstrate the superior performance of our baseline, with increasing gains observed as range expands. },
keywords = {},
pubstate = {published},
tppubtype = {inproceedings}
}
3D semantic occupancy prediction is an emerging perception paradigm in autonomous driving, providing a voxel-level representation of both geometric details and semantic categories. However, its effectiveness is inherently constrained in single-vehicle setups by occlusions, restricted sensor range, and narrow viewpoints. To address these limitations, collaborative perception enables the exchange of complementary information, thereby enhancing
the completeness and accuracy of predictions. Despite its potential, research on collaborative 3D semantic occupancy prediction is hindered by the lack of dedicated datasets. To bridge this gap, we design a high-resolution semantic voxel sensor in CARLA to produce dense and comprehensive annotations. We further develop a baseline model that performs inter-agent feature fusion via spatial alignment and attention aggregation. In addition, we establish
benchmarks with varying prediction ranges designed to systematically assess the impact of spatial extent on collaborative prediction. Experimental results demonstrate the superior performance of our baseline, with increasing gains observed as range expands.
the completeness and accuracy of predictions. Despite its potential, research on collaborative 3D semantic occupancy prediction is hindered by the lack of dedicated datasets. To bridge this gap, we design a high-resolution semantic voxel sensor in CARLA to produce dense and comprehensive annotations. We further develop a baseline model that performs inter-agent feature fusion via spatial alignment and attention aggregation. In addition, we establish
benchmarks with varying prediction ranges designed to systematically assess the impact of spatial extent on collaborative prediction. Experimental results demonstrate the superior performance of our baseline, with increasing gains observed as range expands.


