OPA2Vec: combining formal and informal content of biomedical ontologies to improve similarity-based prediction
Motivation: Ontologies are widely used in biology for data annotation,\nintegration, and analysis. In addition to formally structured axioms,\nontologies contain meta-data in the form of annotation axioms which provide\nvaluable pieces of information that characterize ontology classes. Annotations\ncommonly used in ontologies include class labels, descriptions, or synonyms.\nDespite being a rich source of semantic information, the ontology meta-data are\ngenerally unexploited by ontology-based analysis methods such as semantic\nsimilarity measures. Results: We propose a novel method, OPA2Vec, to generate\nvector representations of biological entities in ontologies by combining formal\nontology axioms and annotation axioms from the ontology meta-data. We apply a\nWord2Vec model that has been pre-trained on PubMed abstracts to produce feature\nvectors from our collected data. We validate our method in two different ways:\nfirst, we use the obtained vector representations of proteins as a similarity\nmeasure to predict protein-protein interaction (PPI) on two different datasets.\nSecond, we evaluate our method on predicting gene-disease associations based on\nphenotype similarity by generating vector representations of genes and diseases\nusing a phenotype ontology, and applying the obtained vectors to predict\ngene-disease associations. These two experiments are just an illustration of\nthe possible applications of our method. OPA2Vec can be used to produce vector\nrepresentations of any biomedical entity given any type of biomedical ontology.\nAvailability: https://github.com/bio-ontology-research-group/opa2vec Contact:\nrobert.hoehndorf@kaust.edu.sa and xin.gao@kaust.edu.sa.\n