Open Access Open Access  Restricted Access Subscription Access

Intelligent Vehicular Routing and Optimal Resource Allocation Using Deep Reinforcement Learning for Urban VANETs

K. Thamizhmaran

Abstract


The proliferation of connected and autonomous vehicles has intensified the demand for robust Vehicular Ad-hoc Networks (VANETs). A critical challenge in urban VANETs is the co-design of efficient routing protocols and optimal communication resource allocation under highly dynamic and resource-constrained conditions. Traditional routing algorithms often fail to adapt to rapid topological changes, while conventional resource allocation schemes are not cognizant of the specific data flow requirements of multi-hop routes. This paper proposes a novel integrated framework, Deep Reinforcement Learning-based Vehicular Routing with Optimal Resource Allocation (DRL-VRORL), to address this joint problem. We model the urban VANET environment as a Markov Decision Process (MDP) where intelligent agents on vehicles collaboratively learn routing and resource allocation policies. The core of DRL-VRORL is a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) architecture, enhanced with a centralized critic for coordinated learning and a distributed actor for scalable execution. The framework simultaneously optimizes for end-to-end packet delivery delay, packet delivery ratio (PDR), and network throughput. For resource allocation, we formulate a convex optimization problem for power and sub-channel allocation in Vehicle-to-Infrastructure (V2I) and Vehicle-to-Vehicle (V2V) links, which is solved efficiently using the Lagrange duality method, guided by the DRL agent's routing decisions. Extensive simulations conducted in SUMO and NS-3 demonstrate that DRL-VRORL significantly outperforms state-of-the-art protocols like AODV, DSR, and Q-learning-based routing. Specifically, DRL-VRORL achieves up to a 32% higher PDR, 45% lower average end-to-end delay, and 28% better aggregate network throughput while maintaining superior resource utilization efficiency.


Full Text:

PDF

References


M. H. Eiza, T. Owens, Q. Ni, and Q. Shi, "Situation-aware QoS routing algorithm for vehicular ad hoc networks," IEEE Trans. Veh. Technol., vol. 64, no. 12, pp. 5520-5535, Dec. 2015.

J. Mei, K. Zheng, L. Zhao, Y. Lei, and L. Wang, "A latency and reliability guaranteed resource allocation scheme for LTE-V2V communication," IEEE Trans. Wireless Commun., vol. 17, no. 6, pp. 3850-3860, Jun. 2018.

C. Perkins, E. Belding-Royer, and S. Das, "Ad hoc On-Demand Distance Vector (AODV) Routing," IETF RFC 3561, 2003.

B. Karp and H. T. Kung, "GPSR: Greedy perimeter stateless routing for wireless networks," in Proc. ACM MobiCom, 2000, pp. 243-254.

L. Liang, G. Y. Li, and W. Xu, "Resource allocation for D2D-enabled vehicular communications," IEEE Trans. Commun., vol. 65, no. 7, pp. 3186-3197, Jul. 2017.

V. Mnih et al., "Human-level control through deep reinforcement learning," Nature, vol. 518, no. 7540, pp. 529-533, 2015.

J. Liu, K. Zhang, Y. Liu, and J. Li, "A Q-learning-based routing protocol for vehicular ad hoc networks," in Proc. IEEE IV, 2016, pp. 719-724.

D. B. Johnson, D. A. Maltz, and J. Broch, "DSR: The Dynamic Source Routing Protocol for Multi-Hop Wireless Ad Hoc Networks," in Ad Hoc Networking, 2001, pp. 139-172.

M. S. K. L. A. Littman and J. Boyan, "A reinforcement learning routing protocol for wireless networks," in Proc. ICAC, 2001.

C. J. C. H. Watkins and P. Dayan, "Q-learning," Machine Learning, vol. 8, no. 3-4, pp. 279-292, 1992.


Refbacks

  • There are currently no refbacks.