Visual object tracking remains an active research field in computer vision\ndue to persisting challenges with various problem-specific factors in\nreal-world scenes. Many existing tracking methods based on discriminative\ncorrelation filters (DCFs) employ feature extraction networks (FENs) to model\nthe target appearance during the learning process. However, using deep feature\nmaps extracted from FENs based on different residual neural networks (ResNets)\nhas not previously been investigated. This paper aims to evaluate the\nperformance of twelve state-of-the-art ResNet-based FENs in a DCF-based\nframework to determine the best for visual tracking purposes. First, it ranks\ntheir best feature maps and explores the generalized adoption of the best\nResNet-based FEN into another DCF-based method. Then, the proposed method\nextracts deep semantic information from a fully convolutional FEN and fuses it\nwith the best ResNet-based feature maps to strengthen the target representation\nin the learning process of continuous convolution filters. Finally, it\nintroduces a new and efficient semantic weighting method (using semantic\nsegmentation feature maps on each video frame) to reduce the drift problem.\nExtensive experimental results on the well-known OTB-2013, OTB-2015, TC-128 and\nVOT-2018 visual tracking datasets demonstrate that the proposed method\neffectively outperforms state-of-the-art methods in terms of precision and\nrobustness of visual tracking.\n