AUTHOR=Le Phi , Gong Xingyue , Ung Leah , Yang Hai , Keenan Bridget P. , Zhang Li , He Tao TITLE=A robust ensemble feature selection approach to prioritize genes associated with survival outcome in high-dimensional gene expression data JOURNAL=Frontiers in Systems Biology VOLUME=4 YEAR=2024 URL=https://www.frontiersin.org/journals/systems-biology/articles/10.3389/fsysb.2024.1355595 DOI=10.3389/fsysb.2024.1355595 ISSN=2674-0702 ABSTRACT=
Exploring features associated with the clinical outcome of interest is a rapidly advancing area of research. However, with contemporary sequencing technologies capable of identifying over thousands of genes per sample, there is a challenge in constructing efficient prediction models that balance accuracy and resource utilization. To address this challenge, researchers have developed feature selection methods to enhance performance, reduce overfitting, and ensure resource efficiency. However, applying feature selection models to survival analysis, particularly in clinical datasets characterized by substantial censoring and limited sample sizes, introduces unique challenges. We propose a robust ensemble feature selection approach integrated with group Lasso to identify compelling features and evaluate its performance in predicting survival outcomes. Our approach consistently outperforms established models across various criteria through extensive simulations, demonstrating low false discovery rates, high sensitivity, and high stability. Furthermore, we applied the approach to a colorectal cancer dataset from The Cancer Genome Atlas, showcasing its effectiveness by generating a composite score based on the selected genes to correctly distinguish different subtypes of the patients. In summary, our proposed approach excels in selecting impactful features from high-dimensional data, yielding better outcomes compared to contemporary state-of-the-art models.