This paper presents a video coding scheme that combines traditional optimization methods with deep learning methods based on the Enhanced Compression Model (ECM). In order to achieve better subjective quality, we adjust the quantization parameter (QP) of every coding tree unit (CTU) adaptively according to the spatial-temporal perception information and block importance. In addition, the QP of key frame is individually adjusted according to the video content characteristics. Meanwhile, the deep learning methods propose a convolutional neural network-based loop filter (CNNLF), which is turned on/off based on the rate-distortion optimization at the CTU and frame level. Besides, intra-prediction using neural networks (NN-intra) is proposed to further improve compression quality, where 8 neural networks are used for predicting blocks of different sizes. On the validation set of the Challenge for Learning Image Compression (CLIC), the proposed method achieves remarkable higher compression quality than the conventional ECM 3.0 in terms of PSNR and subjective evaluation.