An end-to-end Approach to a Reinforcement Learning in Transport Logistics

The use of machine learning and reinforcement learning techniques has become increasingly important in enhancing the performance of transportation in supply chains. These techniques allow for real-time adaptation to changing conditions and optimization of decision-making, resulting in more efficient and cost-effective transportation routes. By incorporating machine learning and reinforcement learning, companies can improve their overall supply chain management and competitiveness in today's fast-paced business environment. In this paper, we proposed a multi-mode transportation and route planning using Reinforcement Learning (RL) algorithm. The algorithm showed good performance in multi-modal routing and transport selection based on cost functions through the evaluation of three trained agents in 100 different environments. However, a comparison with the Dijkstra algorithm revealed sub-optimal decisions with higher costs. Further training is needed to fully define the optimal policy” with the dynamic environment being a challenge

Paper

Full text

PDF

An end-to-end Approach to a Reinforcement Learning in Transport Logistics

Semantic Scholar · Computer Science · 2023

Abstract

The use of machine learning and reinforcement learning techniques has become increasingly important in enhancing the performance of transportation in supply chains. These techniques allow for real-time adaptation to changing conditions and optimization of decision-making, resulting in more efficient and cost-effective transportation routes. By incorporating machine learning and reinforcement learning, companies can improve their overall supply chain management and competitiveness in today's fast-paced business environment. In this paper, we proposed a multi-mode transportation and route planning using Reinforcement Learning (RL) algorithm. The algorithm showed good performance in multi-modal routing and transport selection based on cost functions through the evaluation of three trained agents in 100 different environments. However, a comparison with the Dijkstra algorithm revealed sub-optimal decisions with higher costs. Further training is needed to fully define the optimal policy” with the dynamic environment being a challenge

Similar papers

© 2026 NYSGPT2525 LLC