Skip to main navigation Skip to search Skip to main content

SubOmiEmbed: Self-supervised Representation Learning of Multi-omics Data for Cancer Type Classification

  • Mohamed Bin Zayed University of Artificial Intelligence

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

4 Scopus citations

Abstract

For personalized medicines, very crucial intrinsic information is present in high dimensional omics data which is difficult to capture due to the large number of molecular features and small number of available samples. Different types of omics data show various aspects of samples. Integration and analysis of multi-omics data give us a broad view of tumours, which can improve clinical decision making. Omics data, mainly DNA methylation and gene expression profiles are usually high dimensional data with a lot of molecular features. In recent years, variational autoencoders (VAE) [1] have been extensively used in embedding image and text data into lower dimensional latent spaces. In our work, we extend the idea of using a VAE model for low dimensional latent space extraction with the self-supervised learning technique of feature subsetting. With VAEs, the key idea is to make the model learn meaningful representations from different types of omics data, which could then be used for downstream tasks such as cancer type classification. The main goals are to overcome the curse of dimensionality and integrate methylation and expression data to combine information about different aspects of same tissue samples, and hopefully extract biologically relevant features. Our extension involves training encoder and decoder to reconstruct the data from just a subset of it. By doing this, we force the model to encode most important information in the latent representation. We also added an identity to the subsets so that the model knows which subset is being fed into it during training and testing. We experimented with our approach and found that SubOmiEmbed produces comparable results to the baseline OmiEmbed [2] with a much smaller network and by using just a subset of the data. This work can be improved to integrate mutation-based genomic data as well.

Original languageEnglish
Title of host publication2022 10th International Conference on Bioinformatics and Computational Biology, ICBCB 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages66-72
Number of pages7
ISBN (Electronic)9781665401081
DOIs
StatePublished - 2022
Event10th International Conference on Bioinformatics and Computational Biology, ICBCB 2022 - Virtual, Hangzhou, China
Duration: 13 May 202215 May 2022

Publication series

Name2022 10th International Conference on Bioinformatics and Computational Biology, ICBCB 2022

Conference

Conference10th International Conference on Bioinformatics and Computational Biology, ICBCB 2022
Country/TerritoryChina
CityVirtual, Hangzhou
Period13/05/2215/05/22

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • cancer type classification
  • feature subsetting
  • self-supervised learning

Fingerprint

Dive into the research topics of 'SubOmiEmbed: Self-supervised Representation Learning of Multi-omics Data for Cancer Type Classification'. Together they form a unique fingerprint.

Cite this