Skip to main navigation Skip to search Skip to main content

RGB-D Scene Classification via Multi-modal Feature Learning

  • Northwestern Polytechnical University Xian
  • Nanjing University of Posts and Telecommunications

Research output: Contribution to journalArticlepeer-review

16 Scopus citations

Abstract

Most of the past deep learning methods which are proposed for RGB-D scene classification use global information and directly consider all pixels in the whole image for high-level tasks. Such methods cannot hold much information about local feature distributions, and simply concatenate RGB and depth features without exploring the correlation and complementarity between raw RGB and depth images. From the human vision perspective, we recognize the category of one unknown scene mainly relying on the object-level information in the scene which includes the appearance, texture, shape, and depth. The structural distribution of different objects is also taken into consideration. Based on this observation, constructing mid-level representations with discriminative object parts would generally be more attractive for scene analysis. In this paper, we propose a new Convolutional Neural Networks (CNNs)-based local multi-modal feature learning framework (LM-CNN) for RGB-D scene classification. This method can effectively capture much of the local structure from the RGB-D scene images and automatically learn a fusion strategy for the object-level recognition step instead of simply training a classifier on top of features extracted from both modalities. The experimental results on two popular datasets, i.e., NYU v1 depth dataset and SUN RGB-D dataset, show that our method with local multi-modal CNNs outperforms state-of-the-art methods.

Original languageEnglish
Pages (from-to)825-840
Number of pages16
JournalCognitive Computation
Volume11
Issue number6
DOIs
StatePublished - 1 Dec 2019

Keywords

  • Convolutional neural networks
  • Deep learning
  • Local fine-tuning
  • RGB-D scene classification

Fingerprint

Dive into the research topics of 'RGB-D Scene Classification via Multi-modal Feature Learning'. Together they form a unique fingerprint.

Cite this