Surgical tool segmentation in endoscopic videos is an important component of\ncomputer assisted interventions systems. Recent success of image-based\nsolutions using fully-supervised deep learning approaches can be attributed to\nthe collection of big labeled datasets. However, the annotation of a big\ndataset of real videos can be prohibitively expensive and time consuming.\nComputer simulations could alleviate the manual labeling problem, however,\nmodels trained on simulated data do not generalize to real data. This work\nproposes a consistency-based framework for joint learning of simulated and real\n(unlabeled) endoscopic data to bridge this performance generalization issue.\nEmpirical results on two data sets (15 videos of the Cholec80 and EndoVis'15\ndataset) highlight the effectiveness of the proposed \\emph{Endo-Sim2Real}\nmethod for instrument segmentation. We compare the segmentation of the proposed\napproach with state-of-the-art solutions and show that our method improves\nsegmentation both in terms of quality and quantity.\n