A Network-based Machine Learning Approach for Identifying Biomarkers of Breast Cancer Survivability

Identifying biomarkers for better diagnosis or prognosis of breast cancer is in demand but presents many challenges. In this study, we introduced a data-integration approach to identify sub-network biomarkers capable of predicting breast cancer treatment outcomes including disease-free survival, and overall survival at five years and long-term. A gene expression data is used for evaluating the predictive power of sub-networks of genes, while the protein-protein interaction network is to guide the search for the candidate sub-networks. To reduce the search space, we proposed a score to estimate the predictive ability of a set of genes, thus, only the candidates with the high score are evaluated by Support Vector Machine classifier during the search. After the sub-networks with highest classification performance were selected for all seed genes, they were further analyzed with pathway data and cancer-related genes from literature for their biological meaning. The selected sub-networks yielded highly accurate and contain genes associated with many cancer pathways, including breast cancer.

Paper

Full text

PDF

A Network-based Machine Learning Approach for Identifying Biomarkers of Breast Cancer Survivability

Semantic Scholar · Medicine · 2019

Abstract

Identifying biomarkers for better diagnosis or prognosis of breast cancer is in demand but presents many challenges. In this study, we introduced a data-integration approach to identify sub-network biomarkers capable of predicting breast cancer treatment outcomes including disease-free survival, and overall survival at five years and long-term. A gene expression data is used for evaluating the predictive power of sub-networks of genes, while the protein-protein interaction network is to guide the search for the candidate sub-networks. To reduce the search space, we proposed a score to estimate the predictive ability of a set of genes, thus, only the candidates with the high score are evaluated by Support Vector Machine classifier during the search. After the sub-networks with highest classification performance were selected for all seed genes, they were further analyzed with pathway data and cancer-related genes from literature for their biological meaning. The selected sub-networks yielded highly accurate and contain genes associated with many cancer pathways, including breast cancer.

Similar papers

© 2026 NYSGPT2525 LLC