상세 보기
Pre-Training and Ensembling of Deep Neural Networks for Target Gene Expression Prediction From Landmark Genes
- Lee, Da-Bin;
- Hwang, Kyu-Baek
WEB OF SCIENCE
0초록
Gene expression profiling is used in many biological and biomedical studies. The L1000 assay is an efficient profiling method that measures the expression of a small number of "landmark" genes and predicts the expression of the remaining "target" genes. Extreme gradient boosting and deep neural networks (DNNs) have shown high accuracy in predicting target gene expression based on landmark genes. We propose a pre-training and ensemble method to improve the accuracy of DNNs. Specifically, we apply autoencoder-based pre-training not only to the input but also to the output of DNNs. Experimental results on gene expression profile data show that the proposed method is more accurate than the state-of-the-art methods. Pre-training with multiple autoencoders, obtained using multiple random seeds, improves accuracy by positioning DNNs in a region favorable for optimization in the function space, while still positioning them far enough apart to take advantage of ensembling. Pre-training also has the effect of diversifying the number of genes associated through the hidden nodes and associating more biologically relevant genes, which is more pronounced when the number of hidden layers is larger. We expect our proposed method to be effective not only for gene expression inference, but also for general multi-target regression problems.
키워드
- 제목
- Pre-Training and Ensembling of Deep Neural Networks for Target Gene Expression Prediction From Landmark Genes
- 저자
- Lee, Da-Bin; Hwang, Kyu-Baek
- 발행일
- 2025-07
- 유형
- Article
- 저널명
- IEEE TRANSACTIONS ON COMPUTATIONAL BIOLOGY AND BIOINFORMATICS
- 권
- 22
- 호
- 4
- 페이지
- 1574 ~ 1586