Skip to main content
Top

2013 | OriginalPaper | Chapter

MOCA-I: Discovering Rules and Guiding Decision Maker in the Context of Partial Classification in Large and Imbalanced Datasets

Authors : Julie Jacques, Julien Taillard, David Delerue, Laetitia Jourdan, Clarisse Dhaenens

Published in: Learning and Intelligent Optimization

Publisher: Springer Berlin Heidelberg

Activate our intelligent search to find suitable subject content or patents.

search-config
loading …

Abstract

This paper focuses on the modeling and the implementation as a multi-objective optimization problem of a Pittsburgh classification rule mining algorithm adapted to large and imbalanced datasets, as encountered in hospital data. We associate to this algorithm an original post-processing method based on ROC curve to help the decision maker to choose the most interesting rules. After an introduction to problems brought by hospital data such as class imbalance, volumetry or inconsistency, we present MOCA-I - a Pittsburgh modelization adapted to this kind of problems. We propose its implementation as a dominance-based local search in opposition to existing multi-objective approaches based on genetic algorithms. Then we introduce the post-processing method to sort and filter the obtained classifiers. Our approach is compared to state-of-the-art classification rule mining algorithms, giving as good or better results, using less parameters. Then it is compared to C4.5 and C4.5-CS on hospital data with a larger set of attributes, giving the best results.

Dont have a licence yet? Then find out more about our products and how to get one now:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 102.000 Bücher
  • über 537 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 390 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe




 

Jetzt Wissensvorsprung sichern!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 340 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko




Jetzt Wissensvorsprung sichern!

Footnotes
1
International classification of diseases; http://​www.​who.​int/​classifications/​icd/​en/​
 
Literature
1.
go back to reference Fernández, A., Garciá, S., Luengo, J., Bernadó-Mansilla, E., Herrera, F.: Genetics-based machine learning for rule induction: state of the art, taxonomy, and comparative study. IEEE Trans. Evol. Comput. 14(6), 913–941 (2010)CrossRef Fernández, A., Garciá, S., Luengo, J., Bernadó-Mansilla, E., Herrera, F.: Genetics-based machine learning for rule induction: state of the art, taxonomy, and comparative study. IEEE Trans. Evol. Comput. 14(6), 913–941 (2010)CrossRef
2.
go back to reference Quinlan, J.R.: C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers Inc., San Francisco (1993) Quinlan, J.R.: C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers Inc., San Francisco (1993)
3.
go back to reference Chawla, N.V.: Data mining for imbalanced datasets: an overview. In: Maimon, O., Rokach, L. (eds.) Data Mining and Knowledge Discovery Handbook, 2nd edn, pp. 875–886. Springer, New York (2010) Chawla, N.V.: Data mining for imbalanced datasets: an overview. In: Maimon, O., Rokach, L. (eds.) Data Mining and Knowledge Discovery Handbook, 2nd edn, pp. 875–886. Springer, New York (2010)
4.
5.
go back to reference Geng, L., Hamilton, H.J.: Interestingness measures for data mining: a survey. ACM Comput. Surv. (CSUR) 38(3), 1–32 (2006)CrossRef Geng, L., Hamilton, H.J.: Interestingness measures for data mining: a survey. ACM Comput. Surv. (CSUR) 38(3), 1–32 (2006)CrossRef
6.
go back to reference Ohsaki, M., Abe, H., Tsumoto, S., Yokoi, H., Yamaguchi, T.: Evaluation of rule interestingness measures in medical knowledge discovery in databases. Artif. Intell. Med. 41, 177–196 (2007)CrossRef Ohsaki, M., Abe, H., Tsumoto, S., Yokoi, H., Yamaguchi, T.: Evaluation of rule interestingness measures in medical knowledge discovery in databases. Artif. Intell. Med. 41, 177–196 (2007)CrossRef
7.
go back to reference Greco, S., Pawlak, Z., Slowiński, R.: Can bayesian confirmation measures be useful for rough set decision rules? Eng. Appl. Artif. Intell. 17(4), 345–361 (2004)CrossRef Greco, S., Pawlak, Z., Slowiński, R.: Can bayesian confirmation measures be useful for rough set decision rules? Eng. Appl. Artif. Intell. 17(4), 345–361 (2004)CrossRef
8.
go back to reference Bayardo, J., Agrawal, R.: Mining the most interesting rules. In: Proceedings of the Fifth ACM SIGKDD, ser. KDD ’99, pp. 145–154 (1999) Bayardo, J., Agrawal, R.: Mining the most interesting rules. In: Proceedings of the Fifth ACM SIGKDD, ser. KDD ’99, pp. 145–154 (1999)
9.
10.
go back to reference Reynolds, A., de la Iglesia, B.: Rule induction for classification using multi-objective genetic programming. In: Obayashi, S., Deb, K., Poloni, C., Hiroyasu, T., Murata, T. (eds.) EMO 2007. LNCS, vol. 4403, pp. 516–530. Springer, Heidelberg (2007) Reynolds, A., de la Iglesia, B.: Rule induction for classification using multi-objective genetic programming. In: Obayashi, S., Deb, K., Poloni, C., Hiroyasu, T., Murata, T. (eds.) EMO 2007. LNCS, vol. 4403, pp. 516–530. Springer, Heidelberg (2007)
11.
go back to reference Bacardit, J., Stout, M., Hirst, J.D., Sastry, K., Llorà, X., Krasnogor, N.: Automated alphabet reduction method with evolutionary algorithms for protein structure prediction. In: GECCO, pp. 346–353 (2007) Bacardit, J., Stout, M., Hirst, J.D., Sastry, K., Llorà, X., Krasnogor, N.: Automated alphabet reduction method with evolutionary algorithms for protein structure prediction. In: GECCO, pp. 346–353 (2007)
12.
go back to reference Corne, D., Dhaenens, C., Jourdan, L.: Synergies between operations research and data mining: the emerging use of multi-objective approaches. Eur. J. Oper. Res. 221(3), 469–479 (2012)MathSciNetCrossRefMATH Corne, D., Dhaenens, C., Jourdan, L.: Synergies between operations research and data mining: the emerging use of multi-objective approaches. Eur. J. Oper. Res. 221(3), 469–479 (2012)MathSciNetCrossRefMATH
13.
go back to reference Srinivasan, S., Ramakrishnan, S.: Evolutionary multi objective optimization for rule mining: a review. Artif. Intell. Rev. 36(3), 205–248 (2011) Srinivasan, S., Ramakrishnan, S.: Evolutionary multi objective optimization for rule mining: a review. Artif. Intell. Rev. 36(3), 205–248 (2011)
14.
go back to reference Coello Coello, C.A, Dhaenens, C., Jourdan, L. (eds.): Advances in Multi-Objective Nature Inspired Computing. SCI, vol. 272. Springer, Heidelberg (2010) Coello Coello, C.A, Dhaenens, C., Jourdan, L. (eds.): Advances in Multi-Objective Nature Inspired Computing. SCI, vol. 272. Springer, Heidelberg (2010)
15.
go back to reference Casillas, J., Martínez, P., Benítez, A.: Learning consistent, complete and compact sets of fuzzy rules in conjunctive normal form for regression problems. Soft Comput. (A Fusion of Foundations, Methodologies and Applications) 13, 451–465 (2009) Casillas, J., Martínez, P., Benítez, A.: Learning consistent, complete and compact sets of fuzzy rules in conjunctive normal form for regression problems. Soft Comput. (A Fusion of Foundations, Methodologies and Applications) 13, 451–465 (2009)
16.
go back to reference Liefooghe, A., Humeau, J., Mesmoudi, S., Jourdan, L., Talbi, E.-G.: On dominance-based multiobjective local search: design, implementation and experimental analysis on scheduling and traveling salesman problems. J. Heuristics 18, 317–352 (2012)CrossRef Liefooghe, A., Humeau, J., Mesmoudi, S., Jourdan, L., Talbi, E.-G.: On dominance-based multiobjective local search: design, implementation and experimental analysis on scheduling and traveling salesman problems. J. Heuristics 18, 317–352 (2012)CrossRef
17.
go back to reference Liefooghe, A., Jourdan, L., Talbi, E.-G.: A software framework based on a conceptual unified model for evolutionary multiobjective optimization: paradiseo-moeo. Eur. J. Oper. Res. 209(2), 104–112 (2011)MathSciNetCrossRef Liefooghe, A., Jourdan, L., Talbi, E.-G.: A software framework based on a conceptual unified model for evolutionary multiobjective optimization: paradiseo-moeo. Eur. J. Oper. Res. 209(2), 104–112 (2011)MathSciNetCrossRef
19.
go back to reference Alcalá-Fdez, J., et al.: Keel: a software tool to assess evolutionary algorithms for data mining problems. Soft Comput. (A Fusion of Foundations, Methodologies and Applications) 13, 307–318 (2009) Alcalá-Fdez, J., et al.: Keel: a software tool to assess evolutionary algorithms for data mining problems. Soft Comput. (A Fusion of Foundations, Methodologies and Applications) 13, 307–318 (2009)
20.
go back to reference Plantevit, M., Laurent, A., Laurent, D., Teisseire, M., Choong, Y.W.: Mining multidimensional and multilevel sequential patterns. ACM TKDD 4(1), 1–37 (2010)CrossRef Plantevit, M., Laurent, A., Laurent, D., Teisseire, M., Choong, Y.W.: Mining multidimensional and multilevel sequential patterns. ACM TKDD 4(1), 1–37 (2010)CrossRef
21.
go back to reference Zhang, J., Bala, J.W., Hadjarian, A., Han, B.: Learning to rank cases with classification rules. In: Fürnkranz, J., Hüllermeier, E. (eds.) Preference Learning, pp. 155–177. Springer, Heidelberg (2011) Zhang, J., Bala, J.W., Hadjarian, A., Han, B.: Learning to rank cases with classification rules. In: Fürnkranz, J., Hüllermeier, E. (eds.) Preference Learning, pp. 155–177. Springer, Heidelberg (2011)
Metadata
Title
MOCA-I: Discovering Rules and Guiding Decision Maker in the Context of Partial Classification in Large and Imbalanced Datasets
Authors
Julie Jacques
Julien Taillard
David Delerue
Laetitia Jourdan
Clarisse Dhaenens
Copyright Year
2013
Publisher
Springer Berlin Heidelberg
DOI
https://doi.org/10.1007/978-3-642-44973-4_5

Premium Partner