This paper introduces a new framework to predict visual attention of\nomnidirectional images. The key setup of our architecture is the simultaneous\nprediction of the saliency map and a corresponding scanpath for a given\nstimulus. The framework implements a fully encoder-decoder convolutional neural\nnetwork augmented by an attention module to generate representative saliency\nmaps. In addition, an auxiliary network is employed to generate probable\nviewport center fixation points through the SoftArgMax function. The latter\nallows to derive fixation points from feature maps. To take advantage of the\nscanpath prediction, an adaptive joint probability distribution model is then\napplied to construct the final unbiased saliency map by leveraging the encoder\ndecoder-based saliency map and the scanpath-based saliency heatmap. The\nproposed framework was evaluated in terms of saliency and scanpath prediction,\nand the results were compared to state-of-the-art methods on Salient360!\ndataset. The results showed the relevance of our framework and the benefits of\nsuch architecture for further omnidirectional visual attention prediction\ntasks.\n