Multisource region attention network for fine-grained object recognition in remote sensing imagery

Sümbül, Gencer; Cinbiş, Ramazan Gökberk; Aksoy, Selim

Multisource region attention network for fine-grained object recognition in remote sensing imagery

buir.contributor.author	Sümbül, Gencer
buir.contributor.author	Cinbiş, Ramazan Gökberk
buir.contributor.author	Aksoy, Selim
dc.citation.epage	4937	en_US
dc.citation.issueNumber	7	en_US
dc.citation.spage	4929	en_US
dc.citation.volumeNumber	57	en_US
dc.contributor.author	Sümbül, Gencer	en_US
dc.contributor.author	Cinbiş, Ramazan Gökberk	en_US
dc.contributor.author	Aksoy, Selim	en_US
dc.date.accessioned	2020-02-04T11:07:07Z
dc.date.available	2020-02-04T11:07:07Z
dc.date.issued	2019-07
dc.department	Department of Computer Engineering	en_US
dc.description.abstract	Fine-grained object recognition concerns the identification of the type of an object among a large number of closely related subcategories. Multisource data analysis that aims to leverage the complementary spectral, spatial, and structural information embedded in different sources is a promising direction toward solving the fine-grained recognition problem that involves low between-class variance, small training set sizes for rare classes, and class imbalance. However, the common assumption of coregistered sources may not hold at the pixel level for small objects of interest. We present a novel methodology that aims to simultaneously learn the alignment of multisource data and the classification model in a unified framework. The proposed method involves a multisource region attention network that computes per-source feature representations, assigns attention scores to candidate regions sampled around the expected object locations by using these representations, and classifies the objects by using an attention-driven multisource representation that combines the feature representations and the attention scores from all sources. All components of the model are realized using deep neural networks and are learned in an end-to-end fashion. Experiments using RGB, multispectral, and LiDAR elevation data for classification of street trees showed that our approach achieved 64.2% and 47.3% accuracies for the 18-class and 40-class settings, respectively, which correspond to 13% and 14.3% improvement relative to the commonly used feature concatenation approach from multiple sources.	en_US
dc.identifier.doi	10.1109/TGRS.2019.2894425	en_US
dc.identifier.issn	0196-2892	en_US
dc.identifier.uri	http://hdl.handle.net/11693/53050	en_US
dc.language.iso	English	en_US
dc.publisher	IEEE	en_US
dc.relation.isversionof	https://doi.org/10.1109/TGRS.2019.2894425	en_US
dc.source.title	IEEE Transactions on Geoscience and Remote Sensing	en_US
dc.subject	Deep learning	en_US
dc.subject	Fine-grained classification	en_US
dc.subject	Image alignment	en_US
dc.subject	Multisource classification	en_US
dc.subject	Object recognition	en_US
dc.title	Multisource region attention network for fine-grained object recognition in remote sensing imagery	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: Multisource_Region_Attention_Network_for_Fine-Grained_Object_Recognition_in_Remote_Sensing_Imagery.pdf
Size:: 2.67 MB
Format:: Adobe Portable Document Format
Description:

Download

Collections

Scholarly Publications - Computer Engineering