Ransomwares spread rapidly over the past two years, which poses great security threats to users. DGA domain name detection is one of the key technologies in detecting ransomwares. Existing detection methods are usually based on machine learning, which needs large amounts of training data. However, it is difficult to collect enough training samples for a specific DGA family in a short time, and few training samples would lead to overfitting of the detection model. A LSTM DGA generation model can obtain a lot of new data learned from few real DGA samples. Reinforcement Learning guides this LSTM generation model to be improved by evaluating its generated domain name, which is proposed as RL-LSTM DGA generation model. In experiments, a DGA domain name detection model (ATT-GRU model) trained by the generated DGAs, is used to compute accuracies of real DGAs (as test dataset) to assess the validity of generated domain names. Experiments show that the distribution of generated DGAs is sufficiently close to real DGAs, and can play an alternative role of real DGAs in detection model training.
Paper
Full text
Detecting Domain Generation Algorithms Based on Reinforcement Learning
Semantic Scholar · Computer Science · 2019
Abstract
Ransomwares spread rapidly over the past two years, which poses great security threats to users. DGA domain name detection is one of the key technologies in detecting ransomwares. Existing detection methods are usually based on machine learning, which needs large amounts of training data. However, it is difficult to collect enough training samples for a specific DGA family in a short time, and few training samples would lead to overfitting of the detection model. A LSTM DGA generation model can obtain a lot of new data learned from few real DGA samples. Reinforcement Learning guides this LSTM generation model to be improved by evaluating its generated domain name, which is proposed as RL-LSTM DGA generation model. In experiments, a DGA domain name detection model (ATT-GRU model) trained by the generated DGAs, is used to compute accuracies of real DGAs (as test dataset) to assess the validity of generated domain names. Experiments show that the distribution of generated DGAs is sufficiently close to real DGAs, and can play an alternative role of real DGAs in detection model training.