Automating the analysis of imagery of the Gastrointestinal (GI) tract\ncaptured during endoscopy procedures has substantial potential benefits for\npatients, as it can provide diagnostic support to medical practitioners and\nreduce mistakes via human error. To further the development of such methods, we\npropose a two-stream model for endoscopic image analysis. Our model fuses two\nstreams of deep feature inputs by mapping their inherent relations through a\nnovel relational network model, to better model symptoms and classify the\nimage. In contrast to handcrafted feature-based models, our proposed network is\nable to learn features automatically and outperforms existing state-of-the-art\nmethods on two public datasets: KVASIR and Nerthus. Our extensive evaluations\nillustrate the importance of having two streams of inputs instead of a single\nstream and also demonstrates the merits of the proposed relational network\narchitecture to combine those streams.\n