When does deep learning fail and how to tackle it? A critical analysis on polymer sequence-property surrogate models
: Deep learning models are gaining popularity and potency in predicting polymer properties. These models can be built using pre-existing data and are useful for the rapid prediction of polymer properties. However, the performance of a deep learning model is intricately connected to its topology and the volume of training data. There is no facile protocol available to select a deep learning architecture, and there is a lack of a large volume of homogeneous sequence-property data of polymers. These two factors are the primary bottleneck for the efficient development of deep learning models. Here we assess the severity of these factors and propose new algorithms to address them. We show that a linear layer-by-layer expansion of a neural network can help in identifying the best neural network topology for a given problem. Moreover, we map the discrete sequence space of a polymer to a continuous one-dimensional latent space using a machine learning pipeline to identify minimal data points for building a universal deep learning model. We implement these approaches for three representative cases of building sequence-property surrogate models, viz., the single-molecule radius of gyration of a copolymer, adhesive free energy of a copolymer, and copolymer compatibilizer, demonstrating the generality of the proposed strategies. This work establishes efficient methods for building universal deep learning models with minimal data and hyperparameters for predicting sequence-defined properties of polymers. This approach does not require any special optimization algorithm to explore enormously large possibilities of a DNN topology. We use this protocol to develop DNN models that predict sequence-defined properties of polymers with more than 95% accuracy. Secondly, we build a DNN model using training data that represent a specific range of property and test this model's ability to predict the property that is outside the training data. We show that the performance of a DNN declines when the target property is outside the known range of property. We propose a new framework to tackle the transferability problem of ML by leveraging the power of convolution DNN autoencoder that automatically extracts features of a molecular system. We construct a one-dimensional sequence space and sample the sequences uniformly covering the entire space. This collection of points serves as the training data for our DNN model. We show that a model based on ~500 data points, which are selected intelligently, can predict the properties of ~40000 sequences very accurately. We expect this model to predict the properties of all possible sequences of a copolymer, which is ~10 30 for a binary copolymer of chain length 100. Although the current study focuses on sequence-property ML models, these methods are extensible for other classes of properties and materials. We expect that these new approaches to data and hyperparameter selections will accelerate the progress of ML model development.