Skip to main content
Erschienen in: Data Mining and Knowledge Discovery 4/2016

01.07.2016

On the evaluation of unsupervised outlier detection: measures, datasets, and an empirical study

verfasst von: Guilherme O. Campos, Arthur Zimek, Jörg Sander, Ricardo J. G. B. Campello, Barbora Micenková, Erich Schubert, Ira Assent, Michael E. Houle

Erschienen in: Data Mining and Knowledge Discovery | Ausgabe 4/2016

Einloggen

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

Abstract

The evaluation of unsupervised outlier detection algorithms is a constant challenge in data mining research. Little is known regarding the strengths and weaknesses of different standard outlier detection models, and the impact of parameter choices for these algorithms. The scarcity of appropriate benchmark datasets with ground truth annotation is a significant impediment to the evaluation of outlier methods. Even when labeled datasets are available, their suitability for the outlier detection task is typically unknown. Furthermore, the biases of commonly-used evaluation measures are not fully understood. It is thus difficult to ascertain the extent to which newly-proposed outlier detection methods improve over established methods. In this paper, we perform an extensive experimental study on the performance of a representative set of standard k nearest neighborhood-based methods for unsupervised outlier detection, across a wide variety of datasets prepared for this purpose. Based on the overall performance of the outlier detection methods, we provide a characterization of the datasets themselves, and discuss their suitability as outlier detection benchmark sets. We also examine the most commonly-used measures for comparing the performance of different methods, and suggest adaptations that are more suitable for the evaluation of outlier detection results.

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 102.000 Bücher
  • über 537 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 390 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe




 

Jetzt Wissensvorsprung sichern!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 340 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko




Jetzt Wissensvorsprung sichern!

Fußnoten
1
While only recently defined formally, SimplifiedLOF has been implicitly used (and adapted), often presumably unintentionally [i.e., not being aware of the special definition of the reachability distance (Eq. 2)], in many earlier variants of LOF. Here, for the first time, it is evaluated explicitly.
 
2
In fact, the number of true outliers expected to be ranked by chance among the top n positions is a fraction n / N of |O|, which yields \(P@n = \frac{n \cdot |O|}{N}\big / n = \frac{|O|}{N}\).
 
4
Available at: http://​www.​ipd.​kit.​edu/​~muellere/​HiCS/​realworld.​zip. Note that we have supplemented our collection with some of these datasets, without further preprocessing.
 
5
For unsupervised learning, both training and test sets can be used together, and we assume this is the case unless otherwise specified.
 
6
FastABOD requires at least a set of 3 neighbors, as it computes variances of angles to neighbors. LDOF, KDEOS, and ODIN require at least 2 neighbors.
 
8
We see the same overall tendency (although much weaker due to overall low values) if we use \(P@n\) and \({{\mathrm{AP}}}\) (both adjusted and unadjusted) instead of ROC AUC. This is expected since (Adjusted) \(P@n\) and (Adjusted) \({{\mathrm{AP}}}\) can yield additional insights when comparing results that are very good in terms of ROC AUC. In this aggregated evaluation, however, many results with weak scores are included. The corresponding plots are available online.
 
9
Therefore, as a side effect, such heat maps can also serve to visualize the profile of performance in terms of \(P@(x \cdot n)\) for \(x=1,\ldots ,9\).
 
10
This is not surprising given the relatively large amount of outliers (\(\approx \)75 %) in the base dataset.
 
11
Prima facie, this conclusion is valid, based on our experiments, for the dependency of related methods on a parameter choice regarding cardinality of a local neighborhood. Common sense suggests that we can have a similar expectation, mutatis mutandis, for other types of parameters for other kinds of methods.
 
Literatur
Zurück zum Zitat Abe N, Zadrozny B, Langford J (2006) Outlier detection by active learning. In: Proceedings of the 12th ACM International Conference on Knowledge Discovery and Data Mining (SIGKDD), Philadelphia, pp 504–509. doi:10.1145/1150402.1150459 Abe N, Zadrozny B, Langford J (2006) Outlier detection by active learning. In: Proceedings of the 12th ACM International Conference on Knowledge Discovery and Data Mining (SIGKDD), Philadelphia, pp 504–509. doi:10.​1145/​1150402.​1150459
Zurück zum Zitat Achtert E, Kriegel HP, Schubert E, Zimek A (2013) Interactive data mining with 3D-parallel-coordinate-trees. In: Proceedings of the ACM international conference on management of data (SIGMOD), New York, pp 1009–1012. doi:10.1145/2463676.2463696 Achtert E, Kriegel HP, Schubert E, Zimek A (2013) Interactive data mining with 3D-parallel-coordinate-trees. In: Proceedings of the ACM international conference on management of data (SIGMOD), New York, pp 1009–1012. doi:10.​1145/​2463676.​2463696
Zurück zum Zitat Angiulli F, Pizzuti C (2002) Fast outlier detection in high dimensional spaces. In: Proceedings of the 6th European Conference on Principles of Data Mining and Knowledge Discovery (PKDD), Helsinki, pp 15–26. doi:10.1007/3-540-45681-3_2 Angiulli F, Pizzuti C (2002) Fast outlier detection in high dimensional spaces. In: Proceedings of the 6th European Conference on Principles of Data Mining and Knowledge Discovery (PKDD), Helsinki, pp 15–26. doi:10.​1007/​3-540-45681-3_​2
Zurück zum Zitat Barnett V, Lewis T (1994) Outliers in statistical data, 3rd edn. Wiley, New YorkMATH Barnett V, Lewis T (1994) Outliers in statistical data, 3rd edn. Wiley, New YorkMATH
Zurück zum Zitat Breunig MM, Kriegel HP, Ng R, Sander J (2000) LOF: identifying density-based local outliers. In: Proceedings of the ACM international conference on management of data (SIGMOD), Dallas, pp 93–104. doi:10.1145/342009.335388 Breunig MM, Kriegel HP, Ng R, Sander J (2000) LOF: identifying density-based local outliers. In: Proceedings of the ACM international conference on management of data (SIGMOD), Dallas, pp 93–104. doi:10.​1145/​342009.​335388
Zurück zum Zitat Dang XH, Micenková B, Assent I, Ng R (2013) Outlier detection with space transformation and spectral analysis. In: Proceedings ofthe 13th SIAM international conference on data mining (SDM), Austin, pp 225–233 Dang XH, Micenková B, Assent I, Ng R (2013) Outlier detection with space transformation and spectral analysis. In: Proceedings ofthe 13th SIAM international conference on data mining (SDM), Austin, pp 225–233
Zurück zum Zitat Dang XH, Assent I, Ng RT, Zimek A, Schubert E (2014) Discriminative features for identifying and interpreting outliers. In: Proceedings of the 30th International Conference on Data Engineering (ICDE), Chicago, pp 88–99. doi:10.1109/ICDE.2014.6816642 Dang XH, Assent I, Ng RT, Zimek A, Schubert E (2014) Discriminative features for identifying and interpreting outliers. In: Proceedings of the 30th International Conference on Data Engineering (ICDE), Chicago, pp 88–99. doi:10.​1109/​ICDE.​2014.​6816642
Zurück zum Zitat Davis J, Goadrich M (2006) The relationship between precision-recall and ROC curves. In: Proceedings of the 23rd international conference on machine learning (ICML), Pittsburgh, pp 233–240 Davis J, Goadrich M (2006) The relationship between precision-recall and ROC curves. In: Proceedings of the 23rd international conference on machine learning (ICML), Pittsburgh, pp 233–240
Zurück zum Zitat de Vries T, Chawla S, Houle ME (2010) Finding local anomalies in very high dimensional space. In: Proceedings of the 10th IEEE International Conference on Data Mining (ICDM), Sydney, pp 128–137. doi:10.1109/ICDM.2010.151 de Vries T, Chawla S, Houle ME (2010) Finding local anomalies in very high dimensional space. In: Proceedings of the 10th IEEE International Conference on Data Mining (ICDM), Sydney, pp 128–137. doi:10.​1109/​ICDM.​2010.​151
Zurück zum Zitat Demšar J (2006) Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res 7:1–30MathSciNetMATH Demšar J (2006) Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res 7:1–30MathSciNetMATH
Zurück zum Zitat Emmott AF, Das S, Dietterich T, Fern A, Wong WK (2013) Systematic construction of anomaly detection benchmarks from real data. In: Workshop on outlier detection and description, held in conjunction with the 19th ACM SIGKDD international conference on knowledge discovery and data mining, Chicago, pp 16–21 Emmott AF, Das S, Dietterich T, Fern A, Wong WK (2013) Systematic construction of anomaly detection benchmarks from real data. In: Workshop on outlier detection and description, held in conjunction with the 19th ACM SIGKDD international conference on knowledge discovery and data mining, Chicago, pp 16–21
Zurück zum Zitat Färber I, Günnemann S, Kriegel HP, Kröger P, Müller E, Schubert E, Seidl T, Zimek A (2010) On using class-labels in evaluation of clusterings. In: MultiClust: 1st international workshop on discovering, summarizing and using multiple clusterings held in conjunction with KDD 2010, Washington, DC Färber I, Günnemann S, Kriegel HP, Kröger P, Müller E, Schubert E, Seidl T, Zimek A (2010) On using class-labels in evaluation of clusterings. In: MultiClust: 1st international workshop on discovering, summarizing and using multiple clusterings held in conjunction with KDD 2010, Washington, DC
Zurück zum Zitat Gao J, Tan PN (2006) Converting output scores from outlier detection algorithms into probability estimates. In: Proceedings of the 6th IEEE international conference on data mining (ICDM), Hong Kong, pp 212–221. doi:10.1109/ICDM.2006.43 Gao J, Tan PN (2006) Converting output scores from outlier detection algorithms into probability estimates. In: Proceedings of the 6th IEEE international conference on data mining (ICDM), Hong Kong, pp 212–221. doi:10.​1109/​ICDM.​2006.​43
Zurück zum Zitat Hanley JA, McNeil BJ (1982) The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143:29–36CrossRef Hanley JA, McNeil BJ (1982) The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143:29–36CrossRef
Zurück zum Zitat Hautamäki V, Kärkkäinen I, Fränti P (2004) Outlier detection using k-nearest neighbor graph. In: Proceedings of the 17th international conference on pattern recognition (ICPR), Cambridge, pp 430–433. doi:10.1109/ICPR.2004.1334558 Hautamäki V, Kärkkäinen I, Fränti P (2004) Outlier detection using k-nearest neighbor graph. In: Proceedings of the 17th international conference on pattern recognition (ICPR), Cambridge, pp 430–433. doi:10.​1109/​ICPR.​2004.​1334558
Zurück zum Zitat Houle ME, Kriegel HP, Kröger P, Schubert E, Zimek A (2010) Can shared-neighbor distances defeat the curse of dimensionality? In: Proceedings of the 22nd international conference on scientific and statistical database management (SSDBM), Heidelberg, pp 482–500. doi:10.1007/978-3-642-13818-8_34 Houle ME, Kriegel HP, Kröger P, Schubert E, Zimek A (2010) Can shared-neighbor distances defeat the curse of dimensionality? In: Proceedings of the 22nd international conference on scientific and statistical database management (SSDBM), Heidelberg, pp 482–500. doi:10.​1007/​978-3-642-13818-8_​34
Zurück zum Zitat Jin W, Tung AKH, Han J, Wang W (2006) Ranking outliers using symmetric neighborhood relationship. In: Proceedings of the 10th Pacific-Asia conference on knowledge discovery and data mining (PAKDD), Singapore, pp 577–593. doi:10.1007/11731139_68 Jin W, Tung AKH, Han J, Wang W (2006) Ranking outliers using symmetric neighborhood relationship. In: Proceedings of the 10th Pacific-Asia conference on knowledge discovery and data mining (PAKDD), Singapore, pp 577–593. doi:10.​1007/​11731139_​68
Zurück zum Zitat Keller F, Müller E, Böhm K (2012) HiCS: high contrast subspaces for density-based outlier ranking. In: Proceedings of the 28th international conference on data engineering (ICDE), Washington, DC, pp 1037–1048. doi:10.1109/ICDE.2012.88 Keller F, Müller E, Böhm K (2012) HiCS: high contrast subspaces for density-based outlier ranking. In: Proceedings of the 28th international conference on data engineering (ICDE), Washington, DC, pp 1037–1048. doi:10.​1109/​ICDE.​2012.​88
Zurück zum Zitat Knorr EM, Ng RT (1997) A unified notion of outliers: properties and computation. In: Proceedings of the 3rd ACM international conference on knowledge discovery and data mining (KDD), Newport Beach, pp 219–222. doi:10.1145/782010.782021 Knorr EM, Ng RT (1997) A unified notion of outliers: properties and computation. In: Proceedings of the 3rd ACM international conference on knowledge discovery and data mining (KDD), Newport Beach, pp 219–222. doi:10.​1145/​782010.​782021
Zurück zum Zitat Knorr EM, Ng RT (1998) Algorithms for mining distance-based outliers in large datasets. In: Proceedings of the 24th international conference on very large data bases (VLDB), New York, pp 392–403 Knorr EM, Ng RT (1998) Algorithms for mining distance-based outliers in large datasets. In: Proceedings of the 24th international conference on very large data bases (VLDB), New York, pp 392–403
Zurück zum Zitat Kriegel HP, Schubert M, Zimek A (2008) Angle-based outlier detection in high-dimensional data. In: Proceedings of the 14th ACM international conference on knowledge discovery and data mining (SIGKDD), Las Vegas, pp 444–452. doi:10.1145/1401890.1401946 Kriegel HP, Schubert M, Zimek A (2008) Angle-based outlier detection in high-dimensional data. In: Proceedings of the 14th ACM international conference on knowledge discovery and data mining (SIGKDD), Las Vegas, pp 444–452. doi:10.​1145/​1401890.​1401946
Zurück zum Zitat Kriegel HP, Kröger P, Schubert E, Zimek A (2009a) LoOP: local outlier probabilities. In: Proceedings of the 18th ACM conference on information and knowledge management (CIKM), Hong Kong, pp 1649–1652. doi:10.1145/1645953.1646195 Kriegel HP, Kröger P, Schubert E, Zimek A (2009a) LoOP: local outlier probabilities. In: Proceedings of the 18th ACM conference on information and knowledge management (CIKM), Hong Kong, pp 1649–1652. doi:10.​1145/​1645953.​1646195
Zurück zum Zitat Kriegel HP, Kröger P, Zimek A (2009b) Clustering high dimensional data: a survey on subspace clustering, pattern-based clustering, and correlation clustering. ACM Trans Knowl Discov Data 3(1):1–58. doi:10.1145/1497577.1497578 CrossRef Kriegel HP, Kröger P, Zimek A (2009b) Clustering high dimensional data: a survey on subspace clustering, pattern-based clustering, and correlation clustering. ACM Trans Knowl Discov Data 3(1):1–58. doi:10.​1145/​1497577.​1497578 CrossRef
Zurück zum Zitat Kriegel HP, Kröger P, Schubert E, Zimek A (2011a) Interpreting and unifying outlier scores. In: Proceedings of the 11th SIAM international conference on data mining (SDM), Mesa, pp 13–24. doi:10.1137/1.9781611972818.2 Kriegel HP, Kröger P, Schubert E, Zimek A (2011a) Interpreting and unifying outlier scores. In: Proceedings of the 11th SIAM international conference on data mining (SDM), Mesa, pp 13–24. doi:10.​1137/​1.​9781611972818.​2
Zurück zum Zitat Kriegel HP, Schubert E, Zimek A (2011b) Evaluation of multiple clustering solutions. In: 2nd MultiClust Workshop: Discovering, Summarizing and Using Multiple Clusterings Held in Conjunction with ECML PKDD 2011, Athens, Greece, pp 55–66 Kriegel HP, Schubert E, Zimek A (2011b) Evaluation of multiple clustering solutions. In: 2nd MultiClust Workshop: Discovering, Summarizing and Using Multiple Clusterings Held in Conjunction with ECML PKDD 2011, Athens, Greece, pp 55–66
Zurück zum Zitat Kriegel HP, Schubert E, Zimek A (2015) The (black) art of runtime evaluation: Are we comparing algorithms or implementations? submitted Kriegel HP, Schubert E, Zimek A (2015) The (black) art of runtime evaluation: Are we comparing algorithms or implementations? submitted
Zurück zum Zitat Latecki LJ, Lazarevic A, Pokrajac D (2007) Outlier detection with kernel density functions. In: Proceedings of the 5th international conference on machine learning and data mining in pattern recognition (MLDM), Leipzig, pp 61–75. doi:10.1007/978-3-540-73499-4_6 Latecki LJ, Lazarevic A, Pokrajac D (2007) Outlier detection with kernel density functions. In: Proceedings of the 5th international conference on machine learning and data mining in pattern recognition (MLDM), Leipzig, pp 61–75. doi:10.​1007/​978-3-540-73499-4_​6
Zurück zum Zitat Lazarevic A, Kumar V (2005) Feature bagging for outlier detection. In: Proceedings of the 11th ACM international conference on knowledge discovery and data mining (SIGKDD), Chicago, pp 157–166. doi:10.1145/1081870.1081891 Lazarevic A, Kumar V (2005) Feature bagging for outlier detection. In: Proceedings of the 11th ACM international conference on knowledge discovery and data mining (SIGKDD), Chicago, pp 157–166. doi:10.​1145/​1081870.​1081891
Zurück zum Zitat Liu FT, Ting KM, Zhou ZH (2012) Isolation-based anomaly detection. ACM Trans Knowl Discov Data 6(1):31–39CrossRef Liu FT, Ting KM, Zhou ZH (2012) Isolation-based anomaly detection. ACM Trans Knowl Discov Data 6(1):31–39CrossRef
Zurück zum Zitat Marques HO, Campello RJGB, Zimek A, Sander J (2015) On the internal evaluation of unsupervised outlier detection. In: Proceedings of the 27th international conference on scientific and statistical database management (SSDBM), San Diego, pp 7:1–12. doi:10.1145/2791347.2791352 Marques HO, Campello RJGB, Zimek A, Sander J (2015) On the internal evaluation of unsupervised outlier detection. In: Proceedings of the 27th international conference on scientific and statistical database management (SSDBM), San Diego, pp 7:1–12. doi:10.​1145/​2791347.​2791352
Zurück zum Zitat Micenková B, van Beusekom J, Shafait F (2012) Stamp verification for automated document authentication. In: 5th International workshop on computational forensics Micenková B, van Beusekom J, Shafait F (2012) Stamp verification for automated document authentication. In: 5th International workshop on computational forensics
Zurück zum Zitat Müller E, Schiffer M, Seidl T (2011) Statistical selection of relevant subspace projections for outlier ranking. In: Proceedings of the 27th international conference on data engineering (ICDE), Hannover, pp 434–445. doi:10.1109/ICDE.2011.5767916 Müller E, Schiffer M, Seidl T (2011) Statistical selection of relevant subspace projections for outlier ranking. In: Proceedings of the 27th international conference on data engineering (ICDE), Hannover, pp 434–445. doi:10.​1109/​ICDE.​2011.​5767916
Zurück zum Zitat Müller E, Assent I, Iglesias P, Mülle Y, Böhm K (2012) Outlier ranking via subspace analysis in multiple views of the data. In: Proceedings of the 12th IEEE international conference on data mining (ICDM), Brussels, pp 529–538. doi:10.1109/ICDM.2012.112 Müller E, Assent I, Iglesias P, Mülle Y, Böhm K (2012) Outlier ranking via subspace analysis in multiple views of the data. In: Proceedings of the 12th IEEE international conference on data mining (ICDM), Brussels, pp 529–538. doi:10.​1109/​ICDM.​2012.​112
Zurück zum Zitat Nemenyi P (1963) Distribution-free multiple comparisons. PhD thesis, New Jersey Nemenyi P (1963) Distribution-free multiple comparisons. PhD thesis, New Jersey
Zurück zum Zitat Nguyen HV, Gopalkrishnan V (2010) Feature extraction for outlier detection in high-dimensional spaces. J Mach Learn Res Proc Track 10:66–75 Nguyen HV, Gopalkrishnan V (2010) Feature extraction for outlier detection in high-dimensional spaces. J Mach Learn Res Proc Track 10:66–75
Zurück zum Zitat Nguyen HV, Ang HH, Gopalkrishnan V (2010) Mining outliers with ensemble of heterogeneous detectors on random subspaces. In: Proceedings of the 15th international conference on database systems for advanced applications (DASFAA), Tsukuba, pp 368–383. doi:10.1007/978-3-642-12026-8_29 Nguyen HV, Ang HH, Gopalkrishnan V (2010) Mining outliers with ensemble of heterogeneous detectors on random subspaces. In: Proceedings of the 15th international conference on database systems for advanced applications (DASFAA), Tsukuba, pp 368–383. doi:10.​1007/​978-3-642-12026-8_​29
Zurück zum Zitat Orair GH, Teixeira C, Wang Y, Meira W Jr, Parthasarathy S (2010) Distance-based outlier detection: consolidation and renewed bearing. Proc VLDB Endow 3(2):1469–1480CrossRef Orair GH, Teixeira C, Wang Y, Meira W Jr, Parthasarathy S (2010) Distance-based outlier detection: consolidation and renewed bearing. Proc VLDB Endow 3(2):1469–1480CrossRef
Zurück zum Zitat Ramaswamy S, Rastogi R, Shim K (2000) Efficient algorithms for mining outliers from large data sets. In: Proceedings of the ACM international conference on management of data (SIGMOD), Dallas, pp 427–438. doi:10.1145/342009.335437 Ramaswamy S, Rastogi R, Shim K (2000) Efficient algorithms for mining outliers from large data sets. In: Proceedings of the ACM international conference on management of data (SIGMOD), Dallas, pp 427–438. doi:10.​1145/​342009.​335437
Zurück zum Zitat Schubert E, Wojdanowski R, Zimek A, Kriegel HP (2012) On evaluation of outlier rankings and outlier scores. In: Proceedings of the 12th SIAM international conference on data mining (SDM), Anaheim, pp 1047–1058. doi:10.1137/1.9781611972825.90 Schubert E, Wojdanowski R, Zimek A, Kriegel HP (2012) On evaluation of outlier rankings and outlier scores. In: Proceedings of the 12th SIAM international conference on data mining (SDM), Anaheim, pp 1047–1058. doi:10.​1137/​1.​9781611972825.​90
Zurück zum Zitat Schubert E, Zimek A, Kriegel HP (2014a) Generalized outlier detection with flexible kernel density estimates. In: Proceedings of the 14th SIAM International Conference on Data Mining (SDM), Philadelphia, pp 542–550. doi:10.1137/1.9781611973440.63 Schubert E, Zimek A, Kriegel HP (2014a) Generalized outlier detection with flexible kernel density estimates. In: Proceedings of the 14th SIAM International Conference on Data Mining (SDM), Philadelphia, pp 542–550. doi:10.​1137/​1.​9781611973440.​63
Zurück zum Zitat Schubert E, Koos A, Emrich T, Züfle A, Schmid KA, Zimek A (2015a) A framework for clustering uncertain data. Proc VLDB Endow 8(12):1976–1979CrossRef Schubert E, Koos A, Emrich T, Züfle A, Schmid KA, Zimek A (2015a) A framework for clustering uncertain data. Proc VLDB Endow 8(12):1976–1979CrossRef
Zurück zum Zitat Schubert E, Zimek A, Kriegel HP (2015b) Fast and scalable outlier detection with approximate nearest neighbor ensembles. In: Proceedings of the 20th international conference on database systems for advanced applications (DASFAA), Hanoi, Vietnam, pp 19–36. doi:10.1007/978-3-319-18123-3_2 Schubert E, Zimek A, Kriegel HP (2015b) Fast and scalable outlier detection with approximate nearest neighbor ensembles. In: Proceedings of the 20th international conference on database systems for advanced applications (DASFAA), Hanoi, Vietnam, pp 19–36. doi:10.​1007/​978-3-319-18123-3_​2
Zurück zum Zitat Tang J, Chen Z, Fu AWC, Cheung DW (2002) Enhancing effectiveness of outlier detections for low density patterns. In: Proceedings of the 6th Pacific-Asia conference on knowledge discovery and data mining (PAKDD), Taipei, pp 535–548. doi:10.1007/3-540-47887-6_53 Tang J, Chen Z, Fu AWC, Cheung DW (2002) Enhancing effectiveness of outlier detections for low density patterns. In: Proceedings of the 6th Pacific-Asia conference on knowledge discovery and data mining (PAKDD), Taipei, pp 535–548. doi:10.​1007/​3-540-47887-6_​53
Zurück zum Zitat von Luxburg U, Williamson RC, Guyon I (2012) Clustering: science or art? JMLR Workshop Conf Proc 27:65–79 von Luxburg U, Williamson RC, Guyon I (2012) Clustering: science or art? JMLR Workshop Conf Proc 27:65–79
Zurück zum Zitat Wang Y, Parthasarathy S, Tatikonda S (2011) Locality sensitive outlier detection: a ranking driven approach. In: Proceedings of the 27th international conference on data engineering (ICDE), Hannover, pp 410–421. doi:10.1109/ICDE.2011.5767852 Wang Y, Parthasarathy S, Tatikonda S (2011) Locality sensitive outlier detection: a ranking driven approach. In: Proceedings of the 27th international conference on data engineering (ICDE), Hannover, pp 410–421. doi:10.​1109/​ICDE.​2011.​5767852
Zurück zum Zitat Yang J, Zhong N, Yao Y, Wang J (2008) Local peculiarity factor and its application in outlier detection. In: Proceedings of the 14th ACM international conference on knowledge discovery and data mining (SIGKDD), Las Vegas, pp 776–784. doi:10.1145/1401890.1401983 Yang J, Zhong N, Yao Y, Wang J (2008) Local peculiarity factor and its application in outlier detection. In: Proceedings of the 14th ACM international conference on knowledge discovery and data mining (SIGKDD), Las Vegas, pp 776–784. doi:10.​1145/​1401890.​1401983
Zurück zum Zitat Zhang K, Hutter M, Jin H (2009) A new local distance-based outlier detection approach for scattered real-world data. In: Proceedings of the 13th Pacific-Asia conference on knowledge discovery and data mining (PAKDD), Bangkok, pp 813–822. doi:10.1007/978-3-642-01307-2_84 Zhang K, Hutter M, Jin H (2009) A new local distance-based outlier detection approach for scattered real-world data. In: Proceedings of the 13th Pacific-Asia conference on knowledge discovery and data mining (PAKDD), Bangkok, pp 813–822. doi:10.​1007/​978-3-642-01307-2_​84
Zurück zum Zitat Zimek A, Campello RJGB, Sander J (2013a) Ensembles for unsupervised outlier detection: challenges and research questions. ACM SIGKDD Explor 15(1):11–22CrossRef Zimek A, Campello RJGB, Sander J (2013a) Ensembles for unsupervised outlier detection: challenges and research questions. ACM SIGKDD Explor 15(1):11–22CrossRef
Zurück zum Zitat Zimek A, Gaudet M, Campello RJGB, Sander J (2013b) Subsampling for efficient and effective unsupervised outlier detection ensembles. In: Proceedings of the 19th ACM international conference on knowledge discovery and data mining (SIGKDD), Chicago, pp 428–436. doi:10.1145/2487575.2487676 Zimek A, Gaudet M, Campello RJGB, Sander J (2013b) Subsampling for efficient and effective unsupervised outlier detection ensembles. In: Proceedings of the 19th ACM international conference on knowledge discovery and data mining (SIGKDD), Chicago, pp 428–436. doi:10.​1145/​2487575.​2487676
Metadaten
Titel
On the evaluation of unsupervised outlier detection: measures, datasets, and an empirical study
verfasst von
Guilherme O. Campos
Arthur Zimek
Jörg Sander
Ricardo J. G. B. Campello
Barbora Micenková
Erich Schubert
Ira Assent
Michael E. Houle
Publikationsdatum
01.07.2016
Verlag
Springer US
Erschienen in
Data Mining and Knowledge Discovery / Ausgabe 4/2016
Print ISSN: 1384-5810
Elektronische ISSN: 1573-756X
DOI
https://doi.org/10.1007/s10618-015-0444-8

Weitere Artikel der Ausgabe 4/2016

Data Mining and Knowledge Discovery 4/2016 Zur Ausgabe

Premium Partner