nach oben

Knowledge and Information Systems

Erschienen in:

01.06.2021 | Regular Paper

Revisiting data complexity metrics based on morphology for overlap and imbalance: snapshot, new overlap number of balls metrics and singular problems prospect

verfasst von: José Daniel Pascual-Triana, David Charte, Marta Andrés Arroyo, Alberto Fernández, Francisco Herrera

Erschienen in: Knowledge and Information Systems | Ausgabe 7/2021

Einloggen

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config

KI-gestützte Suche

Aus

Abstract

Data Science and Machine Learning have become fundamental assets for companies and research institutions alike. As one of its fields, supervised classification allows for class prediction of new samples, learning from given training data. However, some properties can cause datasets to be problematic to classify. In order to evaluate a dataset a priori, data complexity metrics have been used extensively. They provide information regarding different intrinsic characteristics of the data, which serve to evaluate classifier compatibility and a course of action that improves performance. However, most complexity metrics focus on just one characteristic of the data, which can be insufficient to properly evaluate the dataset towards the classifiers’ performance. In fact, class overlap, a very detrimental feature for the classification process (especially when imbalance among class labels is also present) is hard to assess. This research work focuses on revisiting complexity metrics based on data morphology. In accordance to their nature, the premise is that they provide both good estimates for class overlap, and great correlations with the classification performance. For that purpose, a novel family of metrics has been developed. Being based on ball coverage by classes, they are named after Overlap Number of Balls. Finally, some prospects for the adaptation of the former family of metrics to singular (more complex) problems are discussed.

Vorheriger Artikel Application of genetic algorithm-based intuitionistic fuzzy weighted c-ordered-means algorithm to cluster analysis

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

über 102.000 Bücher
über 537 Zeitschriften

aus folgenden Fachgebieten:

Automobil + Motoren
Bauwesen + Immobilien
Business IT + Informatik
Elektrotechnik + Elektronik
Energie + Nachhaltigkeit
Finance + Banking
Management + Führung
Marketing + Vertrieb
Maschinenbau + Werkstoffe
Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Jetzt informieren

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

über 67.000 Bücher
über 340 Zeitschriften

aus folgenden Fachgebieten:

Bauwesen + Immobilien
Business IT + Informatik
Finance + Banking
Management + Führung
Marketing + Vertrieb
Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Jetzt informieren

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

über 67.000 Bücher
über 390 Zeitschriften

aus folgenden Fachgebieten:

Automobil + Motoren
Bauwesen + Immobilien
Business IT + Informatik
Elektrotechnik + Elektronik
Energie + Nachhaltigkeit
Maschinenbau + Werkstoffe

Jetzt Wissensvorsprung sichern!

Jetzt informieren

Nur mit Berechtigung zugänglich

https://github.com/jdpastri/morphology-metrics.

Aggarwal C (2014) Data classification: algorithms and applications data classification: algorithms and applications. Chapman & Hall/CRC. https://doi.org/10.1201/b17320CrossRefMATH

Ahmed M (2019) Data summarization: a survey. Knowl Information Syst 58(2):249–273. https://doi.org/10.1007/s10115-018-1183-0CrossRef

Alejo R, Valdovinos RM, García V, Pacheco-Sanchez JH (2013) A hybrid method to face class overlap and class imbalance on neural networks and multi-class scenarios. Pattern Recognit Lett 34(4):380–388. https://doi.org/10.1016/j.patrec.2012.09.003CrossRef

Alpaydin E (2016) Machine learning: the new AI. MIT Press, Cambridge

Alshomrani S, Bawakid A, Shim SO, Fernández A, Herrera F (2015) A proposal for evolutionary fuzzy systems using feature weighting: dealing with overlapping in imbalanced datasets. Knowl-Based Syst 73:1–17. https://doi.org/10.1016/j.knosys.2014.09.002CrossRef

Anuradha Gupta G (2014) A self explanatory review of decision tree classifiers. ICRAIE. https://doi.org/10.1109/ICRAIE.2014.6909245CrossRef

Astorino A, Fuduli A, Gaudioso M, Vocaturo E (2019) Multiple instance learning algorithm for medical image classification. SEBD 2400:1–8

Barboza F, Kimura H, Altman E (2017) Machine learning models and bankruptcy prediction. Expert Syst Appl 83:405–417. https://doi.org/10.1016/j.eswa.2017.04.006CrossRef

Baumgartner R, Somorjai R (2006) Data complexity assesment in undersampled classification of high dimensional biomedical data. Pattern Recog Lett 27:1383–1389. https://doi.org/10.1016/j.patrec.2006.01.006CrossRef

10.

Ben-Israel D, Jacobs W, Casha S, Lang S, Ryu W, de Lotbiniere-Bassett M, Cadotte D (2020) The impact of machine learning on patient care: a systematic review. Artifi Intell Med. https://doi.org/10.1016/j.artmed.2019.101785CrossRef

11.

Bergstra J, Bengio Y (2012) Random search for hyper-parameter optimization. J Mach Learn Res 13(10):281–305MathSciNetMATH

12.

Bernadó-Mansilla E, Ho T (2005) Domain of competence of XCS classifier system in complexity measurement space. IEEE Trans Evol Comput 9(1):82–104. https://doi.org/10.1109/TEVC.2004.840153CrossRef

13.

Bielza C, Li G, Larrañaga P (2011) Multi-dimensional classification with bayesian networks. Int J Approx Reason 52(6):705–727. https://doi.org/10.1016/j.ijar.2011.01.007MathSciNetCrossRefMATH

14.

Borchani H, Varando G, Bielza C, Larrañaga P (2015) A survey on multi-output regression. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 5(5):216–233. https://doi.org/10.1002/widm.1157

15.

Boutell MR, Luo J, Shen X, Brown CM (2004) Learning multi-label scene classification. Pattern Recognit 37(9):1757–1771. https://doi.org/10.1016/j.patcog.2004.03.009CrossRef

16.

Cano JR (2013) Analysis of data complexity measures for classification. Expert Syst Appl 40(12):4820–4831. https://doi.org/10.1016/j.eswa.2013.02.025CrossRef

17.

Carbonneau MA, Cheplygina V, Granger E, Gagnon G (2016) Multiple instance learning: a survey of problem characteristics and applications. Pattern Recognit. https://doi.org/10.1016/j.patcog.2017.10.009CrossRef

18.

Charte D, Charte F, García S, Herrera F (2019) A snapshot on nonstandard supervised learning problems: taxonomy, relationships, problem transformations and algorithm adaptations. Prog Artif Intell 8(1):1–14. https://doi.org/10.1007/s13748-018-00167-7CrossRef

19.

Charte F, Rivera AJ, del Jesus MJ, Herrera F (2015) Addressing imbalance in multilabel classification: measures and random resampling algorithms. Neurocomputing 163:3–16. https://doi.org/10.1016/j.neucom.2014.08.091CrossRef

20.

Cózar J, Fernández A, Herrera F, Gámez JA (2019) A metahierarchical rule decision system to design robust fuzzy classifiers based on data complexity. IEEE Trans Fuzzy Syst 27(4):701–715. https://doi.org/10.1109/TFUZZ.2018.2866967CrossRef

21.

Das S, Datta S, Chaudhuri BB (2018) Handling data irregularities in classification: foundations, trends, and future challenges. Pattern Recognit. 81:674–693. https://doi.org/10.1016/j.patcog.2018.03.008

22.

Diedenhofen B, Musch J (2015) cocor: a comprehensive solution for the statistical comparison of correlations. PLOS ONE 10(4):1–12. https://doi.org/10.1371/journal.pone.0121945CrossRef

23.

Diederhofen B cocor function | R Documentation. URL https://www.rdocumentation.org/packages/cocor/versions/1.1-3/topics/cocor

24.

Fernández A, Carmona CJ, Del Jesus MJ, Herrera F (2017) A pareto based ensemble with feature and instance selection for learning from multi-class imbalanced datasets. Int J Neural Syst. https://doi.org/10.1142/S0129065717500289CrossRef

25.

Fernández A, García S, Galar M, Prati R, Krawczyk B, Herrera F (2018). Learning from Imbalanced Data Sets Springer. https://doi.org/10.1007/978-3-319-98074-4

26.

Feurer M, Hutter F (2019) Hyperparameter optimization. Springer, Berlin. https://doi.org/10.1007/978-3-030-05318-5_1CrossRef

27.

Galar M, Fernández A, Barrenechea E, Herrera F (2014) Empowering difficult classes with a similarity-based aggregation in multi-class classification problems. Inf Sci 264:135–157. https://doi.org/10.1016/j.ins.2013.12.053MathSciNetCrossRefMATH

28.

Galar M, Fernández A, Tartas EB, Sola HB, Herrera F (2011) An overview of ensemble methods for binary classifiers in multi-class problems: experimental study on one-vs-one and one-vs-all schemes. Pattern Recognit 44(8):1761–1776. https://doi.org/10.1016/j.patcog.2011.01.017

29.

Garcia LPF, Carvalho ACPdLFd, Lorena AC (2015) Effect of label noise in the complexity of classification problems. Neurocomputing. https://doi.org/10.1016/j.neucom.2014.10.085CrossRef

30.

García S, Luengo J, Herrera F (2016) Tutorial on practical tips of the most influential data preprocessing algorithms in data mining. Knowl-Based Syst 98:1–29. https://doi.org/10.1016/j.knosys.2015.12.006CrossRef

31.

Geng X (2016) Label distribution learning. IEEE Trans Knowl Data Eng 28(7):1734–1748. https://doi.org/10.1109/TKDE.2016.2545658CrossRef

32.

Gu B, Sheng V, Tay K, Romano W, Li S (2015) Incremental support vector learning for ordinal regression. IEEE Trans Neural Netw Learn Syst 26(7):1403–1416. https://doi.org/10.1109/TNNLS.2014.2342533MathSciNetCrossRef

33.

Gupta MR, Bengio S, Weston J (2014) Training highly multiclass classifiers. J Mach Learn Res 15(1):1461–1492. https://dl.acm.org/doi/10.5555/2627435.2638582

34.

Hand DJ, Till RJ (2001) A simple generalisation of the area under the ROC curve for multiple class classification problems. Mach Learn 45(2):171–186. https://doi.org/10.1023/A:1010920819831CrossRefMATH

35.

Herrera F, Charte F, Rivera AJ, Jesus MJd (2016) Multilabel Classification : Problem Analysis, Metrics and Techniques. Springer, Berlin. https://doi.org/10.1007/978-3-319-41111-8CrossRef

36.

Herrera F, Ventura S, Bello R, Cornelis C, Zafra A, Sánchez-Tarragó D, Vluymans S (2016) Multiple instance learning: foundations and algorithms. Springer, Berlin. https://doi.org/10.1007/978-3-319-47759-6CrossRefMATH

37.

Ho TK, Basu M (2002) Complexity measures of supervised classification problems. IEEE Trans Pattern Anal Mach Intell 24(3):289–300. https://doi.org/10.1109/34.990132CrossRef

38.

Hoekstra A, Duin R (1996) On the nonlinearity of pattern classifiers. In: Proceedings of 13th International Conference on Pattern Recognition, vol. 4, pp. 271–275 vol.4. https://doi.org/10.1109/ICPR.1996.547429. ISSN: 1051-4651

39.

Hornik K Weka\(\_\)classifier\(\_\)trees function | R Documentation. URL https://www.rdocumentation.org/packages/RWeka/versions/0.4-42/topics/Weka_classifier_trees

40.

Hüllermeier E, Fürnkranz J, Cheng W, Brinker K (2008) Label ranking by learning pairwise preferences. Artif Intell 172(16–17):1897–1916. https://doi.org/10.1016/j.artint.2008.08.002MathSciNetCrossRefMATH

41.

Hutter F, Kotthoff L, Vanschoren J (2019) Automated machine learning - methods, systems challenges. Springer, BerlinCrossRef

42.

Katakis I, Tsoumakas G, Vlahavas I (2008) Multilabel text classification for automated tag suggestion. Proc. ECML PKDD08 Discovery Challenge p. 9

43.

Krawczyk B, Triguero I, García S, Woźniak M, Herrera F (2019) Instance reduction for one-class classification. Knowl Inf Syst 59(3):601–628. https://doi.org/10.1007/s10115-018-1220-zCrossRef

44.

Leevy J, Khoshgoftaar T, Bauder R, Seliya N (2018) A survey on addressing high-class imbalance in big data. J Big Data. https://doi.org/10.1186/s40537-018-0151-6CrossRef

45.

Leyva E, González A, Pérez R (2015) A set of complexity measures designed for applying meta-learning to instance selection. IEEE Trans Knowl Data Eng 27(2):354–367. https://doi.org/10.1109/TKDE.2014.2327034CrossRef

46.

Lorena A, Costa I, Spolaôr N, de Souto M (2012) Analysis of complexity indices for classification problems: cancer gene expression data. Neurocomputing 75:33–42. https://doi.org/10.1016/j.neucom.2011.03.054CrossRef

47.

Lorena AC, Garcia LPF, Lehmann J, Souto MCP, Ho TK (2019) How Complex is your classification problem? A survey on measuring classification complexity. ACM Comput Surv 52(5):34. https://doi.org/10.1145/3347711

48.

Luengo J, Fernández A, García S, Herrera F (2011) Addressing data complexity for imbalanced data sets: analysis of SMOTE-based oversampling and evolutionary undersampling. Soft Comput 15(10):1909–1936. https://doi.org/10.1007/s00500-010-0625-8CrossRef

49.

Luengo J, García-Gil D, Ramírez-Gallego S, García S, Herrera F (2020) Big data preprocessing: enabling smart data. Springer, Berlin. https://doi.org/10.1007/978-3-030-39105-8CrossRef

50.

Luengo J, Herrera F (2010) Domains of competence of fuzzy rule based classification systems with data complexity measures: A case of study using a fuzzy hybrid genetic based machine learning method. Fuzzy Sets Syst 161(1):3–19. https://doi.org/10.1016/j.fss.2009.04.001MathSciNetCrossRef

51.

Luengo J, Herrera F (2012) Shared domains of competence of approximate learning models using measures of separability of classes. Inf Sci 185(1):43–65. https://doi.org/10.1016/j.ins.2011.09.022MathSciNetCrossRef

52.

Luengo J, Herrera F (2015) An automatic extraction method of the domains of competence for learning classifiers using data complexity measures. Knowl Inf Syst 42(1):147–180. https://doi.org/10.1007/s10115-013-0700-4CrossRef

53.

Luo G (2016) A review of automatic selection methods for machine learning algorithms and hyper-parameter values. Netw Model Anal Health Inform Bioinform. https://doi.org/10.1007/s13721-016-0125-6CrossRef

54.

Luque A, Carrasco A, Martín A, de lasde las Heras AA (2019) The impact of class imbalance in classification performance metrics based on the binary confusion matrix. Pattern Recognit 91:216–231. https://doi.org/10.1016/j.patcog.2019.02.023CrossRef

55.

López V, Fernández A, García S, Palade V, Herrera F (2013) An insight into classification with imbalanced data: empirical results and current trends on using data intrinsic characteristics. Inf Sci 250:113–141. https://doi.org/10.1016/j.ins.2013.07.007CrossRef

56.

Ma Y (2018) Data complexity analysis for software defect detection. Int J Perform Eng. https://doi.org/10.23940/ijpe.18.08.p5.16951704CrossRef

57.

Manukyan A, Ceyhan E (2016) Classification of Imbalanced Data with a Geometric Digraph Family. J. Mach. Learn. Res. https://dl.acm.org/doi/abs/10.5555/2946645.3053471

58.

Martínez Torres J, Iglesias Comesaña C, García-Nieto PJ (2019) Review: machine learning techniques applied to cybersecurity. Int J Mach Learn Cybern 10(10):2823–2836. https://doi.org/10.1007/s13042-018-00906-1CrossRef

59.

Mazurowski M, Malof J, Tourassi G (2011) Comparative analysis of instance selection algorithms for instance-based classifiers in the context of medical decision support. Phys Med Biol 56(2):473–489. https://doi.org/10.1088/0031-9155/56/2/012CrossRef

60.

Meyer D naiveBayes function | R Documentation. URL https://www.rdocumentation.org/packages/e1071/versions/1.7-2/topics/naiveBayes

61.

Morais G, Prati RC (2013) Complex Network Measures for Data Set Characterization. In: 2013 Brazilian Conference on Intelligent Systems, pp. 12–18. https://doi.org/10.1109/BRACIS.2013.11

62.

Morán-Fernández L, Bolón-Canedo V, Alonso-Betanzos A (2016) Can classification performance be predicted by complexity measures? a study using microarray data. Knowl Inf Syst. https://doi.org/10.1007/s10115-016-1003-3CrossRef

63.

Orriols-Puig A, Macia N, Ho TK (2010) Documentation for the data complexity library in C++. Universitat Ramon Llull, La Salle 196:1–40

64.

Prati RC, Luengo J, Herrera F (2019) Emerging topics and challenges of learning from noisy data in nonstandard classification: a survey beyond binary class noise. Knowl Inf Syst 60:63–97. https://doi.org/10.1007/s10115-018-1244-4

65.

Rodriguez D, Dolado J, Tuya J (2015) Bayesian concepts in software testing: An initial review. In: A-TEST 2015: Proceedings of the 6th International Workshop on Automating Test Case Design, Selection and Evaluation, pp. 41–46. https://doi.org/10.1145/2804322.2804329

66.

Schliep K kknn function | R Documentation. https://www.rdocumentation.org/packages/kknn/versions/1.3.1%20/topics/kknn

67.

Scopus: Document Search. URL https://www.scopus.com/search/form.uri?display=basic

68.

Shalev-Shwartz S, Ben-David S (2014) Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, USACrossRef

69.

Singh S (2003) Multiresolution estimates of classification complexity. IEEE Trans Pattern Anal Mach Intell. https://doi.org/10.1109/TPAMI.2003.1251146CrossRef

70.

Sun S, Mao L, Dong Z, Wu L (2019) Multiview Machine Learning, 1st edn. Springer, BerlinCrossRef

71.

Sáez JA, Luengo J, Herrera F (2013) Predicting noise filtering efficacy with data complexity measures for nearest neighbor classification. Pattern Recognit 46(1):355–364. https://doi.org/10.1016/j.patcog.2012.07.009CrossRef

72.

Tanwani AK, Farooq M (2010) Classification Potential vs. Classification Accuracy: A Comprehensive Study of Evolutionary Algorithms with Biomedical Datasets. In: Bacardit J, Browne W, Drugowitsch, J Bernadó-Mansilla E, Butz MV (eds) Learning Classifier Systems Lecture Notes in Computer Science, (pp. 127–144) Springer, Berlin. doi: https://doi.org/10.1007/978-3-642-17508-4_9

73.

Triguero I, González S, Moyano JM, García S, Alcalá-Fdez J, Luengo J, Fernández A, Jesús MJd, Sánchez L, Herrera F (2017) KEEL 3.0: an open source software for multi-stage analysis in data mining. Int J Comput Intell Syst 10(1):1238–1249. https://doi.org/10.2991/ijcis.10.1.82CrossRef

74.

Vuttipittayamongkol P, Elyan E (2020) Neighbourhood-based undersampling approach for handling imbalanced and overlapped data. Inf Sci 509:47–70. https://doi.org/10.1016/j.ins.2019.08.062CrossRef

75.

Wojciechowski S, Wilk S (2017) Difficulty factors and preprocessing in imbalanced data sets: an experimental study on artificial data. Found Comput Decis Sci 42(2):149–176. https://doi.org/10.1515/fcds-2017-0007CrossRefMATH

76.

Zhao J, Xie X, Xu X, Sun S (2017) Multi-view learning overview: recent progress and new challenges. Inf Fusion 38:43–54. https://doi.org/10.1016/j.inffus.2017.02.007CrossRef

77.

Zhu X (2005) Semi-supervised learning with graphs. phd, Carnegie Mellon University, USA. AAI3179046 ISBN-10: 0542190591

78.

Zou GY (2007) Toward using confidence intervals to compare correlations. Psychol Methods 12(4):399–413. https://doi.org/10.1037/1082-989X.12.4.399CrossRef

Titel: Revisiting data complexity metrics based on morphology for overlap and imbalance: snapshot, new overlap number of balls metrics and singular problems prospect
verfasst von: José Daniel Pascual-Triana
David Charte
Marta Andrés Arroyo
Alberto Fernández
Francisco Herrera
Publikationsdatum: 01.06.2021
Verlag: Springer London
Erschienen in: Knowledge and Information Systems / Ausgabe 7/2021
Print ISSN: 0219-1377
Elektronische ISSN: 0219-3116
DOI: https://doi.org/10.1007/s10115-021-01577-1

Springer Professional

Abstract

Bitte loggen Sie sich ein, um Zugang zu Ihrer Lizenz zu erhalten.

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Springer Professional "Wirtschaft"

Springer Professional "Technik"

Weitere Artikel der Ausgabe 7/2021

A meta-algorithm for finding large k-plexes

Anonymous location sharing in urban area mobility

A Nested Two-Stage Clustering Method for Structured Temporal Sequence Data

A novel cluster-based approach for keyphrase extraction from MOOC video lectures

GrAFCI+ A fast generator-based algorithm for mining frequent closed itemsets

An integrated approach using IF-TOPSIS, fuzzy DEMATEL, and enhanced CSA optimized ANFIS for software risk prediction