This paper makes several contributions to automatic lyrics transcription\n(ALT) research. Our main contribution is a novel variant of the Multistreaming\nTime-Delay Neural Network (MTDNN) architecture, called MSTRE-Net, which\nprocesses the temporal information using multiple streams in parallel with\nvarying resolutions keeping the network more compact, and thus with a faster\ninference and an improved recognition rate than having identical TDNN streams.\nIn addition, two novel preprocessing steps prior to training the acoustic model\nare proposed. First, we suggest using recordings from both monophonic and\npolyphonic domains during training the acoustic model. Second, we tag\nmonophonic and polyphonic recordings with distinct labels for discriminating\nnon-vocal silence and music instances during alignment. Moreover, we present a\nnew test set with a considerably larger size and a higher musical variability\ncompared to the existing datasets used in ALT literature, while maintaining the\ngender balance of the singers. Our best performing model sets the\nstate-of-the-art in lyrics transcription by a large margin. For\nreproducibility, we publicly share the identifiers to retrieve the data used in\nthis paper.\n