AUTHOR=Štancl Paula , Karlić Rosa TITLE=Machine learning for pan-cancer classification based on RNA sequencing data JOURNAL=Frontiers in Molecular Biosciences VOLUME=10 YEAR=2023 URL=https://www.frontiersin.org/journals/molecular-biosciences/articles/10.3389/fmolb.2023.1285795 DOI=10.3389/fmolb.2023.1285795 ISSN=2296-889X ABSTRACT=

Despite recent improvements in cancer diagnostics, 2%-5% of all malignancies are still cancers of unknown primary (CUP), for which the tissue-of-origin (TOO) cannot be determined at the time of presentation. Since the primary site of cancer leads to the choice of optimal treatment, CUP patients pose a significant clinical challenge with limited treatment options. Data produced by large-scale cancer genomics initiatives, which aim to determine the genomic, epigenomic, and transcriptomic characteristics of a large number of individual patients of multiple cancer types, have led to the introduction of various methods that use machine learning to predict the TOO of cancer patients. In this review, we assess the reproducibility, interpretability, and robustness of results obtained by 20 recent studies that utilize different machine learning methods for TOO prediction based on RNA sequencing data, including their reported performance on independent data sets and identification of important features. Our review investigates the strengths and weaknesses of different methods, checks the correspondence of their results, and identifies potential issues with datasets used for model training and testing, assessing their potential usefulness in a clinical setting and suggesting future improvements.