Top

Multimedia Systems

Published in:

01-02-2024 | Regular Paper

Depth alignment interaction network for camouflaged object detection

Authors: Hongbo Bi, Yuyu Tong, Jiayuan Zhang, Cong Zhang, Jinghui Tong, Wei Jin

Published in: Multimedia Systems | Issue 1/2024

Activate our intelligent search to find suitable subject content or patents.

search-config

AI-assisted search

Off

Abstract

Many animals actively change their own characteristics, such as color and texture, through camouflage, a natural defense mechanism, making them difficult to be detected in the natural environment, which makes the task of camouflaged object detection extremely challenging. Biological research shows that the eyes of animals have three-dimensional perception ability, and the obtained depth information can provide useful object positioning clues for finding camouflaged objects. However, almost all the current studies for camouflaged object detection do not combine depth maps with RGB images. Therefore, combining depth maps with traditional unimodal RGB images is of great research significance to improve the accuracy of camouflaged object detection. In this paper, we propose a depth alignment interaction network for camouflaged object detection in which the depth maps used are generated from existing monocular depth estimation networks. To address the problem that the quality of the generated depth maps varies, we propose a depth alignment index method to evaluate the quality of the depth maps. The method dynamically assigns the proportion of depth maps in the fusion process to depth maps of different quality according to their alignment with RGB images. Then, to fully extract the fused artifact features, we design an expanded pyramid interaction module, which first expands the receptive field of the features in each layer. Then, the features at the higher levels interacted with the features at the lower levels by connecting them step-by-step to further refine the predicted camouflaged area. Extensive experiments on 4 camouflaged object detection datasets demonstrate the effectiveness of our solution for camouflaged object detection.

previous article GVA: guided visual attention approach for automatic image caption generation

next article A multi-layer mesh synchronized reversible data hiding algorithm on the 3D model

Dont have a licence yet? Then find out more about our products and how to get one now:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

über 102.000 Bücher
über 537 Zeitschriften

aus folgenden Fachgebieten:

Automobil + Motoren
Bauwesen + Immobilien
Business IT + Informatik
Elektrotechnik + Elektronik
Energie + Nachhaltigkeit
Finance + Banking
Management + Führung
Marketing + Vertrieb
Maschinenbau + Werkstoffe
Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

inform now

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

über 67.000 Bücher
über 390 Zeitschriften

aus folgenden Fachgebieten:

Automobil + Motoren
Bauwesen + Immobilien
Business IT + Informatik
Elektrotechnik + Elektronik
Energie + Nachhaltigkeit
Maschinenbau + Werkstoffe

Jetzt Wissensvorsprung sichern!

inform now

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

über 67.000 Bücher
über 340 Zeitschriften

aus folgenden Fachgebieten:

Bauwesen + Immobilien
Business IT + Informatik
Finance + Banking
Management + Führung
Marketing + Vertrieb
Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

inform now

Han, J., Chen, H., Liu, N., Yan, C., Li, X.: Cnns-based rgb-d saliency detection via cross-view transfer and multiview fusion. IEEE Trans. Cybern. 48(11), 3171–3183 (2017)CrossRef

Hazirbas, C., Ma, L., Domokos, C., Cremers, D.: Fusenet: incorporating depth into semantic segmentation via fusion-based cnn architecture. In: Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20–24, 2016, Revised Selected Papers, Part I 13, pp. 213–228. Springer (2017)

Adams, W.J., Graf, E.W., Anderson, M.: Disruptive coloration and binocular disparity: breaking camouflage. Proc. R. Soc. B 286(1896), 20182045 (2019)CrossRefPubMedPubMedCentral

Penacchio, O., Lovell, P.G., Cuthill, I.C., Ruxton, G.D., Harris, J.M.: Three-dimensional camouflage: exploiting photons to conceal form. Am. Nat. 186(4), 553–563 (2015)CrossRefPubMed

Xiang, M., Zhang, J., Lv, Y., Li, A., Zhong, Y., Dai, Y.: Exploring depth contribution for camouflaged object detection (2021). arXiv:2106.13217

Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., Koltun, V.: Towards robust monocular depth estimation: mixing datasets for zero-shot cross-dataset transfer. IEEE Trans. Pattern Anal. Mach. Intell. 44(3), 1623–1637 (2020)CrossRef

Godard, C., Mac Aodha, O., Firman, M., Brostow, G.J.: Digging into self-supervised monocular depth estimation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3828–3838 (2019)

Li, Z., Dekel, T., Cole, F., Tucker, R., Snavely, N., Liu, C., Freeman, W.T.: Learning the depths of moving people by watching frozen people. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4521–4530 (2019)

Yang, M., Yu, K., Zhang, C., Li, Z., Yang, K.: Denseaspp for semantic segmentation in street scenes. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3684–3692 (2018)

10.

Fan, D.-P., Ji, G.-P., Cheng, M.-M., Shao, L.: Concealed object detection. IEEE Trans. Pattern Anal. Mach. Intell. (2021). https://doi.org/10.1109/TPAMI.2021.3085766CrossRefPubMed

11.

Copeland, A.C., Trivedi, M.M.: Models and metrics for signature strength evaluation of camouflaged targets. In: Algorithms for Synthetic Aperture Radar Imagery IV, vol. 3070, pp. 194–199. SPIE (1997)

12.

Bhajantri, N.U., Nagabhushan, P.: Camouflage defect identification: a novel approach. In: 9th International Conference on Information Technology (ICIT’06), pp. 145–148. IEEE (2006)

13.

Feng, X., Guoying, C., Wei, S.: Camouflage texture evaluation using saliency map. In: Proceedings of the Fifth International Conference on Internet Multimedia Computing and Service, pp. 93–96 (2013)

14.

Li, S., Florencio, D., Li, W., Zhao, Y., Cook, C.: A fusion framework for camouflaged moving foreground detection in the wavelet domain. IEEE Trans. Image Process. 27(8), 3918–3930 (2018)ADSMathSciNetCrossRef

15.

Pike, T.W.: Quantifying camouflage and conspicuousness using visual salience. Methods Ecol. Evol. 9(8), 1883–1895 (2018)CrossRef

16.

Xue, F., Yong, C., Xu, S., Dong, H., Luo, Y., Jia, W.: Camouflage performance analysis and evaluation framework based on features fusion. Multimed. Tools Appl. 75(7), 4065–4082 (2016)CrossRef

17.

Zheng, Y., Zhang, X., Wang, F., Cao, T., Sun, M., Wang, X.: Detection of people with camouflage pattern via dense deconvolution network. IEEE Signal Process. Lett. 26(1), 29–33 (2018)ADSCrossRef

18.

Kendall, A., Gal, Y.: What uncertainties do we need in bayesian deep learning for computer vision? In: Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 5580–5590 (2017)

19.

He, C., Li, K., Zhang, Y., Tang, L., Zhang, Y., Guo, Z., Li, X.: Camouflaged object detection with feature decomposition and edge reconstruction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22046–22055 (2023)

20.

Ji, G.-P., Fan, D.-P., Chou, Y.-C., Dai, D., Liniger, A., Van Gool, L.: Deep gradient learning for efficient camouflaged object detection. Mach. Intell. Res. 20(1), 92–108 (2023)CrossRef

21.

Cheng, Y., Fu, H., Wei, X., Xiao, J., Cao, X.: Depth enhanced saliency detection method. In: Proceedings of International Conference on Internet Multimedia Computing and Service, pp. 23–27 (2014)

22.

Feng, D., Barnes, N., You, S., McCarthy, C.: Local background enclosure for rgb-d salient object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2343–2350 (2016)

23.

Peng, H., Li, B., Xiong, W., Hu, W., Ji, R.: Rgbd salient object detection: a benchmark and algorithms. In: Computer Vision—ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6–12, 2014, Proceedings, Part III 13, pp. 92–109. Springer (2014)

24.

Ren, J., Gong, X., Yu, L., Zhou, W., Ying Yang, M.: Exploiting global priors for rgb-d saliency detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 25–32 (2015)

25.

Zhu, C., Li, G., Wang, W., Wang, R.: An innovative salient object detection using center-dark channel prior. In: Proceedings of the IEEE International Conference on Computer Vision Workshops, pp. 1509–1515 (2017)

26.

Wang, X., Zhu, L., Tang, S., Fu, H., Li, P., Wu, F., Yang, Y., Zhuang, Y.: Boosting rgb-d saliency detection by leveraging unlabeled rgb images. IEEE Trans. Image Process. 31, 1107–1119 (2022)ADSCrossRefPubMed

27.

Wu, Y.-H., Liu, Y., Xu, J., Bian, J.-W., Gu, Y.-C., Cheng, M.-M.: Mobilesal: extremely efficient rgb-d salient object detection. IEEE Trans. Pattern Anal. Mach. Intell. 44(12), 10261–10269 (2021)CrossRef

28.

Zhou, T., Fu, H., Chen, G., Zhou, Y., Fan, D.-P., Shao, L.: Specificity-preserving rgb-d saliency detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4681–4691 (2021)

29.

Li, C., Cong, R., Kwong, S., Hou, J., Fu, H., Zhu, G., Zhang, D., Huang, Q.: Asif-net: attention steered interweave fusion network for rgb-d salient object detection. IEEE Trans. Cybern. 51(1), 88–100 (2020)CrossRefPubMed

30.

Chen, Q., Fu, K., Liu, Z., Chen, G., Du, H., Qiu, B., Shao, L.: Ef-net: a novel enhancement and fusion network for rgb-d saliency detection. Pattern Recogn. 112, 107740 (2021)CrossRef

31.

Piao, Y., Ji, W., Li, J., Zhang, M., Lu, H.: Depth-induced multi-scale recurrent attention network for saliency detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7254–7263 (2019)

32.

Zhang, M., Ren, W., Piao, Y., Rong, Z., Lu, H.: Select, supplement and focus for rgb-d saliency detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3472–3481 (2020)

33.

Deng, Z., Todorovic, S., Jan Latecki, L.: Semantic segmentation of rgbd images with mutex constraints. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 1733–1741 (2015)

34.

Gupta, S., Arbelaez, P., Malik, J.: Perceptual organization and recognition of indoor scenes from rgb-d images. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 564–571 (2013)

35.

Ren, X., Bo, L., Fox, D.: Rgb-(d) scene labeling: Features and algorithms. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2759–2766. IEEE (2012)

36.

Silberman, N., Fergus, R.: Indoor scene segmentation using a structured light sensor. In: 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops), pp. 601–608. IEEE (2011)

37.

Couprie, C., Farabet, C., Najman, L., LeCun, Y.: Indoor semantic segmentation using depth information (2013). arXiv:1301.3572

38.

Chen, X., Lin, K.-Y., Wang, J., Wu, W., Qian, C., Li, H., Zeng, G.: Bi-directional cross-modality feature propagation with separation-and-aggregation gate for rgb-d semantic segmentation. In: European Conference on Computer Vision, pp. 561–577. Springer (2020)

39.

Gatys, L.A., Ecker, A.S., Bethge, M.: Image style transfer using convolutional neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2414–2423 (2016)

40.

Zhao, Z.-Q., Zheng, P., Xu, S.-T., Wu, X.: Object detection with deep learning: a review. IEEE Trans. Neural Netw. Learn. Syst. 30(11), 3212–3232 (2019)CrossRefPubMed

41.

Fan, D.P., Gong, C., Cao, Y., Ren, B., Cheng, M.M., Borji, A.: Enhanced-alignment measure for binary foreground map evaluation. In: IJCAI. AAAI Press (2018)

42.

Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2117–2125 (2017)

43.

Ghiasi, G., Lin, T.-Y., Le, Q.V.: Nas-fpn: learning scalable feature pyramid architecture for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7036–7045 (2019)

44.

Tan, M., Pang, R., Le, Q.V.: Efficientdet: scalable and efficient object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10781–10790 (2020)

45.

Liu, S., Qi, L., Qin, H., Shi, J., Jia, J.: Path aggregation network for instance segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8759–8768 (2018)

46.

Gao, N., Shan, Y., Wang, Y., Zhao, X., Yu, Y., Yang, M., Huang, K.: Ssap: single-shot instance segmentation with affinity pyramid. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 642–651 (2019)

47.

Hu, M., Li, Y., Fang, L., Wang, S.: A2-fpn: attention aggregation based feature pyramid network for instance segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15343–15352 (2021)

48.

Lin, G., Milan, A., Shen, C., Reid, I.: Refinenet: multi-path refinement networks for high-resolution semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1925–1934 (2017)

49.

Nie, D., Lan, R., Wang, L., Ren, X.: Pyramid architecture for multi-scale processing in point cloud segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17284–17294 (2022)

50.

Qiu, S., Anwar, S., Barnes, N.: Semantic segmentation for real point cloud scenes via bilateral augmentation and adaptive fusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1757–1767 (2021)

51.

Kingma, D.P., Ba, J.: Adam: a method for stochastic optimization (2014). arXiv:1412.6980

52.

Le, T.-N., Nguyen, T.V., Nie, Z., Tran, M.-T., Sugimoto, A.: Anabranch network for camouflaged object segmentation. Comput. Vis. Image Underst. 184, 45–56 (2019)CrossRef

53.

Skurowski, P., Abdulameer, H., Błaszczyk, J., Depta, T., Kornacki, A., Kozieł, P.: Animal camouflage analysis: Chameleon database. Unpublished manuscript 2(6), 7 (2018)

54.

Fan, D.-P., Ji, G.-P., Sun, G., Cheng, M.-M., Shen, J., Shao, L.: Camouflaged object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2777–2787 (2020)

55.

Lv, Y., Zhang, J., Dai, Y., Li, A., Liu, B., Barnes, N., Fan, D.-P.: Simultaneously localize, segment and rank the camouflaged objects. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11591–11601 (2021)

56.

Bi, H., Zhang, C., Wang, K., Tong, J., Zheng, F.: Rethinking camouflaged object detection: models and datasets. IEEE Trans. Circuits Syst. Video Technol. 32(9), 5708–5724 (2021)CrossRef

57.

Fan, D.-P., Cheng, M.-M., Liu, Y., Li, T., Borji, A.: Structure-measure: a new way to evaluate foreground maps. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 4548–4557 (2017)

58.

Perazzi, F., Krähenbühl, P., Pritch, Y., Hornung, A.: Saliency filters: Contrast based filtering for salient region detection. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 733–740. IEEE (2012)

59.

Margolin, R., Zelnik-Manor, L., Tal, A.: How to evaluate foreground maps? In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255 (2014)

60.

Zhao, J.-X., Liu, J.-J., Fan, D.-P., Cao, Y., Yang, J., Cheng, M.-M.: Egnet: edge guidance network for salient object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8779–8788 (2019)

61.

Chen, S., Tan, X., Wang, B., Hu, X.: Reverse attention for salient object detection. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 234–250 (2018)

62.

Wu, Z., Su, L., Huang, Q.: Cascaded partial decoder for fast and accurate salient object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3907–3916 (2019)

63.

Fan, D.-P., Ji, G.-P., Zhou, T., Chen, G., Fu, H., Shen, J., Shao, L.: Pranet: Parallel reverse attention network for polyp segmentation. MICCAI (2020)

64.

Pang, Y., Zhao, X., Zhang, L., Lu, H.: Multi-scale interactive network for salient object detection (2020)

65.

Fan, D.-P., Ji, G.-P., Sun, G., Cheng, M.-M., Shen, J., Shao, L.: Camouflaged object detection. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)

66.

Mei, H., Ji, G.-P., Wei, Z., Yang, X., Wei, X., Fan, D.-P.: Camouflaged object segmentation with distraction mining. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021)

67.

Yang, F., Zhai, Q., Li, X., Huang, R., Cheng, H., Fan, D.-P.: Uncertainty-guided transformer reasoning for camouflaged object detection. In: IEEE International Conference on Computer Vision (ICCV) (2021)

68.

Ji, G.P., Zhu, L., Zhuge, M., Fu, K.: Fast camouflaged object detection via edge-based reversible re-calibration network. Pattern Recogn. 123, 108414 (2022)CrossRef

69.

Zhuge, M., Lu, X., Guo, Y., Cai, Z., Chen, S.: Cubenet: X-shape connection for camouflaged object detection. Pattern Recogn. 127, 108644 (2022)CrossRef

70.

Lv, Y., Zhang, J., Dai, Y., Li, A., Barnes, N., Fan, D.-P.: Towards deeper understanding of camouflaged object detection. IEEE Trans. Circuits Syst. Video Technol. 33, 3462–3476 (2023)CrossRef

71.

Khan, A., Khan, M., Gueaieb, W., El Saddik, A., De Masi, G., Karray, F.: Recod: resource-efficient camouflaged object detection for uav-based smart cities applications. In: 2023 IEEE International Smart Cities Conference (ISC2), pp. 1–5. IEEE (2023)

72.

He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)

73.

Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., Adam, H.: Mobilenets: efficient convolutional neural networks for mobile vision applications (2017). CoRR arXiv:1704.04861

Title: Depth alignment interaction network for camouflaged object detection
Authors: Hongbo Bi
Yuyu Tong
Jiayuan Zhang
Cong Zhang
Jinghui Tong
Wei Jin
Publication date: 01-02-2024
Publisher: Springer Berlin Heidelberg
Published in: Multimedia Systems / Issue 1/2024
Print ISSN: 0942-4962
Electronic ISSN: 1432-1882
DOI: https://doi.org/10.1007/s00530-023-01250-3

Springer Professional

Abstract

Please log in to get access to your license.

Dont have a licence yet? Then find out more about our products and how to get one now:

Springer Professional "Wirtschaft+Technik"

Springer Professional "Technik"

Springer Professional "Wirtschaft"

Other articles of this Issue 1/2024

Multi-label neural architecture search for chest radiography image classification

SwinCT: feature enhancement based low-dose CT images denoising with swin transformer

An automatic music generation method based on RSCLN_Transformer network

Universal unsupervised cross-domain 3D shape retrieval

AI and data-driven media analysis of TV content for optimised digital content marketing

A real-time camera-based gaze-tracking system involving dual interactive modes and its application in gaming