Most existing symbolic music generation methods focus on generating short pieces, typically less than 8 bars and occasionally up to 32 bars. Generating long music sequences requires effective representation of coherent musical structures. Vanilla self-attention face challenges in capturing subtle long-term musical structures. We propose an approach to transfer the structural characteristics of training samples for generating music. We introduce a separable self-attention-based model that facilitates the learning and transfer of structural embeddings. It can generate music sequences of up to 100 bars, producing compositions with interpretable structures that closely resemble the structural and compositional techniques of the training set. Experiments show the model’s ability to generate music with targeted structures while maintaining good diversity.