Intelligent Vehicular Routing and Optimal Resource Allocation Using Deep Reinforcement Learning for Urban VANETs
Abstract
The proliferation of connected and autonomous vehicles has intensified the demand for robust Vehicular Ad-hoc Networks (VANETs). A critical challenge in urban VANETs is the co-design of efficient routing protocols and optimal communication resource allocation under highly dynamic and resource-constrained conditions. Traditional routing algorithms often fail to adapt to rapid topological changes, while conventional resource allocation schemes are not cognizant of the specific data flow requirements of multi-hop routes. This paper proposes a novel integrated framework, Deep Reinforcement Learning-based Vehicular Routing with Optimal Resource Allocation (DRL-VRORL), to address this joint problem. We model the urban VANET environment as a Markov Decision Process (MDP) where intelligent agents on vehicles collaboratively learn routing and resource allocation policies. The core of DRL-VRORL is a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) architecture, enhanced with a centralized critic for coordinated learning and a distributed actor for scalable execution. The framework simultaneously optimizes for end-to-end packet delivery delay, packet delivery ratio (PDR), and network throughput. For resource allocation, we formulate a convex optimization problem for power and sub-channel allocation in Vehicle-to-Infrastructure (V2I) and Vehicle-to-Vehicle (V2V) links, which is solved efficiently using the Lagrange duality method, guided by the DRL agent's routing decisions. Extensive simulations conducted in SUMO and NS-3 demonstrate that DRL-VRORL significantly outperforms state-of-the-art protocols like AODV, DSR, and Q-learning-based routing. Specifically, DRL-VRORL achieves up to a 32% higher PDR, 45% lower average end-to-end delay, and 28% better aggregate network throughput while maintaining superior resource utilization efficiency.
References
M. H. Eiza, T. Owens, Q. Ni, and Q. Shi, "Situation-aware QoS routing algorithm for vehicular ad hoc networks," IEEE Trans. Veh. Technol., vol. 64, no. 12, pp. 5520-5535, Dec. 2015.
J. Mei, K. Zheng, L. Zhao, Y. Lei, and L. Wang, "A latency and reliability guaranteed resource allocation scheme for LTE-V2V communication," IEEE Trans. Wireless Commun., vol. 17, no. 6, pp. 3850-3860, Jun. 2018.
C. Perkins, E. Belding-Royer, and S. Das, "Ad hoc On-Demand Distance Vector (AODV) Routing," IETF RFC 3561, 2003.
B. Karp and H. T. Kung, "GPSR: Greedy perimeter stateless routing for wireless networks," in Proc. ACM MobiCom, 2000, pp. 243-254.
L. Liang, G. Y. Li, and W. Xu, "Resource allocation for D2D-enabled vehicular communications," IEEE Trans. Commun., vol. 65, no. 7, pp. 3186-3197, Jul. 2017.
V. Mnih et al., "Human-level control through deep reinforcement learning," Nature, vol. 518, no. 7540, pp. 529-533, 2015.
J. Liu, K. Zhang, Y. Liu, and J. Li, "A Q-learning-based routing protocol for vehicular ad hoc networks," in Proc. IEEE IV, 2016, pp. 719-724.
D. B. Johnson, D. A. Maltz, and J. Broch, "DSR: The Dynamic Source Routing Protocol for Multi-Hop Wireless Ad Hoc Networks," in Ad Hoc Networking, 2001, pp. 139-172.
M. S. K. L. A. Littman and J. Boyan, "A reinforcement learning routing protocol for wireless networks," in Proc. ICAC, 2001.
C. J. C. H. Watkins and P. Dayan, "Q-learning," Machine Learning, vol. 8, no. 3-4, pp. 279-292, 1992.
Refbacks
- There are currently no refbacks.