Fault Tolerance using Reinforcement Learning for Cloud Resource Management: Fault Tolerance using RL for Cloud Resource Management
The cloud environment has become an essential platform due to its computing abilities and is being used in various fields and sectors all around the globe. Users from all over the globe use this computing platform to process their challenging tasks. The cloud computes these tasks on its Virtual Machines (VM) using the appropriate resource scheduling algorithms. While a particular task is being computed, there is always a chance that the cloud suffers damages due to the dynamically generated faults of the task. The cloud also needs better performance with proper resource scheduling, leading to increased costs. To focus on these problems and provide an intelligence mechanism to the cloud, an algorithm named Reinforcement Learning – First Come, First Serve (RL – FCFS) has been designed and implemented by combining the Reinforcement Learning (RL) technique with the existing resource scheduling algorithm First Come First Serve (FCFS) to handle the dynamic faults and provide better cost by improving the resource scheduling at its end. This RL – FCFS algorithm provides a fault-tolerance mechanism at the cloud's end by computing 55.5 % of tasks aggregately compared to an aggregate of 11.1 % for the FCFS. Also, it aggregately improves the cost by 18.50 % across all scenarios. With the RL – FCFS algorithm, the cloud will be in a learning phase at the beginning. With RL rewards and feedback, the cloud will adapt and begin to handle these dynamic faults over time and improve its resource scheduling process, ultimately providing the best Quality of Service (QoS).
Paper
Full text
Fault Tolerance using Reinforcement Learning for Cloud Resource Management: Fault Tolerance using RL for Cloud Resource Management
Semantic Scholar · Computer Science · 2023
Abstract
The cloud environment has become an essential platform due to its computing abilities and is being used in various fields and sectors all around the globe. Users from all over the globe use this computing platform to process their challenging tasks. The cloud computes these tasks on its Virtual Machines (VM) using the appropriate resource scheduling algorithms. While a particular task is being computed, there is always a chance that the cloud suffers damages due to the dynamically generated faults of the task. The cloud also needs better performance with proper resource scheduling, leading to increased costs. To focus on these problems and provide an intelligence mechanism to the cloud, an algorithm named Reinforcement Learning – First Come, First Serve (RL – FCFS) has been designed and implemented by combining the Reinforcement Learning (RL) technique with the existing resource scheduling algorithm First Come First Serve (FCFS) to handle the dynamic faults and provide better cost by improving the resource scheduling at its end. This RL – FCFS algorithm provides a fault-tolerance mechanism at the cloud's end by computing 55.5 % of tasks aggregately compared to an aggregate of 11.1 % for the FCFS. Also, it aggregately improves the cost by 18.50 % across all scenarios. With the RL – FCFS algorithm, the cloud will be in a learning phase at the beginning. With RL rewards and feedback, the cloud will adapt and begin to handle these dynamic faults over time and improve its resource scheduling process, ultimately providing the best Quality of Service (QoS).