A key desiderata for inclusive and accessible speech recognition technology\nis ensuring its robust performance to children's speech. Notably, this includes\nthe rapidly advancing neural network based end-to-end speech recognition\nsystems. Children speech recognition is more challenging due to the larger\nintra-inter speaker variability in terms of acoustic and linguistic\ncharacteristics compared to adult speech. Furthermore, the lack of adequate and\nappropriate children speech resources adds to the challenge of designing robust\nend-to-end neural architectures. This study provides a critical assessment of\nautomatic children speech recognition through an empirical study of\ncontemporary state-of-the-art end-to-end speech recognition systems. Insights\nare provided on the aspects of training data requirements, adaptation on\nchildren data, and the effect of children age, utterance lengths, different\narchitectures and loss functions for end-to-end systems and role of language\nmodels on the speech recognition performance.\n