Skip to main content
Erschienen in: Cluster Computing 2/2016

01.06.2016

A fast approach to identify trending articles in hot topics from XML based big bibliographic datasets

verfasst von: K. P. Swaraj, D. Manjula

Erschienen in: Cluster Computing | Ausgabe 2/2016

Einloggen

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

Abstract

Nowadays XML based big bibliographic datasets are common in different domains which provide meta data about articles published in that domain. They have well defined tags which give details of the year, title, authors, abstract, keywords, the type of article, the venue of publishing the article and other such specific details about each article. A lot of statistics can be extracted from this dataset. Most of the time the tag pertaining to domain sub topic information associated with the article will be absent in the dataset as it is not an article attribute. Hence for such statistics articles must be mapped to its associated sub domain. This paper investigates this problem and proposes a fast approach to find trending articles and hot topics from XML based big bibliographic datasets. The proposed framework uses domain ontology to first classify articles into its sub topics. Fast detection of hot topics, trending keywords and articles is achieved using novel Map Reduce algorithms implemented on a hadoop distributed framework. Performance comparison demonstrates that it outperforms its non-Map Reduce counterpart in quickly sorting out the trending keywords and titles in a particular hot topic from XML based bibliographic dataset.

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 102.000 Bücher
  • über 537 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 390 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe




 

Jetzt Wissensvorsprung sichern!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 340 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko




Jetzt Wissensvorsprung sichern!

Literatur
1.
Zurück zum Zitat Ley M.: The DBLP computer science bibliography: evolution, research issues, perspectives. In: Proceedings of the 9th International Symposium on String Processing and Information Retrieval, pp. 1–10, Springer, London (2002) Ley M.: The DBLP computer science bibliography: evolution, research issues, perspectives. In: Proceedings of the 9th International Symposium on String Processing and Information Retrieval, pp. 1–10, Springer, London (2002)
2.
Zurück zum Zitat Alwahaishi, S., Martinovič, J., Snášel, V.: Analysis of the DBLP publication classification using concept lattices. In: Digital Enterprise and Information Systems Communications in Computer and Information Science, vol. 194, pp. 99–108 (2011) Alwahaishi, S., Martinovič, J., Snášel, V.: Analysis of the DBLP publication classification using concept lattices. In: Digital Enterprise and Information Systems Communications in Computer and Information Science, vol. 194, pp. 99–108 (2011)
3.
Zurück zum Zitat Biryukov, M., Dong, C.: Analysis of Computer Science Communities Based on DBLP Research and Advanced Technology for Digital Libraries. Lecture Notes in Computer Science, Vol. 6273, pp. 228–235. Springer, Berlin (2010) Biryukov, M., Dong, C.: Analysis of Computer Science Communities Based on DBLP Research and Advanced Technology for Digital Libraries. Lecture Notes in Computer Science, Vol. 6273, pp. 228–235. Springer, Berlin (2010)
4.
Zurück zum Zitat Minks, S., Martinovic, J., Drazdilova, P., Slaninova, K.: Author cooperation based on terms of article titles from DBLP. In: Proceedings of the Third International Conference on Intelligent Human Computer Interaction (IHCI 2011), Prague, Czech Republic, pp. 281–290. Springer, Berlin (2011) Minks, S., Martinovic, J., Drazdilova, P., Slaninova, K.: Author cooperation based on terms of article titles from DBLP. In: Proceedings of the Third International Conference on Intelligent Human Computer Interaction (IHCI 2011), Prague, Czech Republic, pp. 281–290. Springer, Berlin (2011)
5.
Zurück zum Zitat Obadi, G., Drazdilova, P., Hlavacek, L., Martinovic, J., Snasel, V. : A tolerance rough set based overlapping clustering for the DBLP Data. In: Proceedings of the International Conference on Web Intelligence and Intelligent Agent Technology, pp. 57–60. IEEE (2010) Obadi, G., Drazdilova, P., Hlavacek, L., Martinovic, J., Snasel, V. : A tolerance rough set based overlapping clustering for the DBLP Data. In: Proceedings of the International Conference on Web Intelligence and Intelligent Agent Technology, pp. 57–60. IEEE (2010)
6.
Zurück zum Zitat Wartena, C., Brussee, R.: Topic detection by clustering keywords. In: Proceedings of the 19th International Conference on Database and Expert Systems Applications, pp. 54–58. IEEE Computer Society, Washington, DC (2008) Wartena, C., Brussee, R.: Topic detection by clustering keywords. In: Proceedings of the 19th International Conference on Database and Expert Systems Applications, pp. 54–58. IEEE Computer Society, Washington, DC (2008)
7.
Zurück zum Zitat Griffiths, T.I., Steyvers, M.: Finding scientific topics. Proc. Natl. Acad. Sci. USA 101, 5228–5235 (2004)CrossRef Griffiths, T.I., Steyvers, M.: Finding scientific topics. Proc. Natl. Acad. Sci. USA 101, 5228–5235 (2004)CrossRef
8.
Zurück zum Zitat Rathore, A.S., Devshri, R.: Performance of LDA and DCT models. J. Inf. Sci. 40(3), 281–292 (2014)CrossRef Rathore, A.S., Devshri, R.: Performance of LDA and DCT models. J. Inf. Sci. 40(3), 281–292 (2014)CrossRef
9.
Zurück zum Zitat Wang, X., McCallum, A.: Topics over time: a non-Markov continuous-time model for topical trends. In: Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 424–433. ACM, New York (2006) Wang, X., McCallum, A.: Topics over time: a non-Markov continuous-time model for topical trends. In: Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 424–433. ACM, New York (2006)
10.
Zurück zum Zitat Krishna, S.M., Bhavani, S.D.: An efficient approach for text clustering based on frequent itemsets. Eur. J. Sci. Res. 42(3), 399–410 (2010) Krishna, S.M., Bhavani, S.D.: An efficient approach for text clustering based on frequent itemsets. Eur. J. Sci. Res. 42(3), 399–410 (2010)
11.
Zurück zum Zitat Agarwal, R., Srikant, R.: Fast algorithms for mining association rules in large databases. In: Proceedings of the 20th International Conference on Very Large Data Bases, pp. 487–499. Morgan Kaufmann Publishers Inc, San Francisco (1994) Agarwal, R., Srikant, R.: Fast algorithms for mining association rules in large databases. In: Proceedings of the 20th International Conference on Very Large Data Bases, pp. 487–499. Morgan Kaufmann Publishers Inc, San Francisco (1994)
12.
Zurück zum Zitat Beil, F., Ester, M., Xu, X.: Frequent term-based text clustering. In: Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 436–442. ACM, New York (2002) Beil, F., Ester, M., Xu, X.: Frequent term-based text clustering. In: Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 436–442. ACM, New York (2002)
13.
Zurück zum Zitat Abe, H., Tsumoto, S.: Evaluating a temporal pattern detection method for finding research keys in bibliographical data. In: Transactions on Rough Sets XIV. Lecture Notes in Computer Science, vol. 6600, pp. 1–17 (2011) Abe, H., Tsumoto, S.: Evaluating a temporal pattern detection method for finding research keys in bibliographical data. In: Transactions on Rough Sets XIV. Lecture Notes in Computer Science, vol. 6600, pp. 1–17 (2011)
14.
Zurück zum Zitat Decker, S. L., Aleman-Meza, B., Cameron, D., Arpinar, I. B.: Detection of Bursty and Emerging Trends towards Identification of Researchers at the Early Stage of Trends. (Tech. Rep. No. 11148065665). University of Georgia, Computer Science Department (2007) Decker, S. L., Aleman-Meza, B., Cameron, D., Arpinar, I. B.: Detection of Bursty and Emerging Trends towards Identification of Researchers at the Early Stage of Trends. (Tech. Rep. No. 11148065665). University of Georgia, Computer Science Department (2007)
15.
Zurück zum Zitat Jun, S.: A Technology forecasting method using text mining and visual apriori algorithm. Appl. Math. Inf. Sci 8, 35–40 (2014)CrossRef Jun, S.: A Technology forecasting method using text mining and visual apriori algorithm. Appl. Math. Inf. Sci 8, 35–40 (2014)CrossRef
16.
Zurück zum Zitat Ma, J., Xu, W., Sun, Y., Turban, E., Wang, S., Liu, O.: An ontology-based text-mining method to cluster proposals for research project selection. IEEE Trans. Syst. Man Cybern. A 42(3), 784–790 (2012)CrossRef Ma, J., Xu, W., Sun, Y., Turban, E., Wang, S., Liu, O.: An ontology-based text-mining method to cluster proposals for research project selection. IEEE Trans. Syst. Man Cybern. A 42(3), 784–790 (2012)CrossRef
17.
Zurück zum Zitat Punnarut, R., Sriharee, G.A.: A researcher expertise search system using ontology-based data mining. In: Proceedings of the Seventh Asia-Pacific Conference on Conceptual Modelling, vol. 110, pp 71–78. Australian Computer Society, Inc., Darling Hurst (2010) Punnarut, R., Sriharee, G.A.: A researcher expertise search system using ontology-based data mining. In: Proceedings of the Seventh Asia-Pacific Conference on Conceptual Modelling, vol. 110, pp 71–78. Australian Computer Society, Inc., Darling Hurst (2010)
18.
Zurück zum Zitat Rajpathak, D.G.: An ontology based text mining system for knowledge discovery from the diagnosis data in the automotive domain. Comput. Ind. 64(5), 565–580 (2013)CrossRef Rajpathak, D.G.: An ontology based text mining system for knowledge discovery from the diagnosis data in the automotive domain. Comput. Ind. 64(5), 565–580 (2013)CrossRef
19.
Zurück zum Zitat Chen, L.-C., Kuo, P.-J., Liao, I.-E.: Ontology-based library recommender system using MapReduce. Clust. Comput. 18, 113–121 (2015)CrossRef Chen, L.-C., Kuo, P.-J., Liao, I.-E.: Ontology-based library recommender system using MapReduce. Clust. Comput. 18, 113–121 (2015)CrossRef
20.
Zurück zum Zitat Han, J.-S., Kim, G.-J.: A method of intelligent recommendation using task ontology. Clust. Comput. 17, 827–833 (2014)CrossRef Han, J.-S., Kim, G.-J.: A method of intelligent recommendation using task ontology. Clust. Comput. 17, 827–833 (2014)CrossRef
21.
Zurück zum Zitat Shubhankar, K., Singh, A.P., Pudi, V.: An efficient algorithm for topic ranking and modeling topic evolution. In: Database and Expert Systems Applications. Lecture Notes in Computer Science, vol. 6860, pp. 320–330. Springer (2011) Shubhankar, K., Singh, A.P., Pudi, V.: An efficient algorithm for topic ranking and modeling topic evolution. In: Database and Expert Systems Applications. Lecture Notes in Computer Science, vol. 6860, pp. 320–330. Springer (2011)
22.
Zurück zum Zitat Shubhankar, K., Singh, A. P., Pudi, V.: A Frequent keyword-set based algorithm for topic modeling and clustering of research papers. In: Proceedings of the 3rd Conference on Data Mining and Optimization (DMO), pp 96–102. IEEE, Selangor (2011) Shubhankar, K., Singh, A. P., Pudi, V.: A Frequent keyword-set based algorithm for topic modeling and clustering of research papers. In: Proceedings of the 3rd Conference on Data Mining and Optimization (DMO), pp 96–102. IEEE, Selangor (2011)
23.
Zurück zum Zitat Pan, Y., Lu, W., Zhang, Y., Chiu, K.: A static load-balancing scheme for parallel XML parsing on multicore CPUs. In: Proceedings of Seventh IEEE International Symposium on Cluster Computing and the Grid, CCGRID, Rio De Janeiro, pp. 351–362 (2007) Pan, Y., Lu, W., Zhang, Y., Chiu, K.: A static load-balancing scheme for parallel XML parsing on multicore CPUs. In: Proceedings of Seventh IEEE International Symposium on Cluster Computing and the Grid, CCGRID, Rio De Janeiro, pp. 351–362 (2007)
24.
Zurück zum Zitat Chen, R., Liao, H.: ParaParse: A parallel method for XML parsing. In: Proceedings of IEEE 3rd International Conference on Communication Software and Networks (ICCSN), pp. 81–85 (2011) Chen, R., Liao, H.: ParaParse: A parallel method for XML parsing. In: Proceedings of IEEE 3rd International Conference on Communication Software and Networks (ICCSN), pp. 81–85 (2011)
25.
Zurück zum Zitat Fen, Z., Yabin, X., Yanping, L.: Research on internet hot topic detection based on MapReduce architecture. In: Proceedings of the 4th International Conference on Intelligent Human-Machine Systems and Cybernetics, vol. 01, pp 81–84. IEEE Computer Society, Washington, DC (2012) Fen, Z., Yabin, X., Yanping, L.: Research on internet hot topic detection based on MapReduce architecture. In: Proceedings of the 4th International Conference on Intelligent Human-Machine Systems and Cybernetics, vol. 01, pp 81–84. IEEE Computer Society, Washington, DC (2012)
26.
Zurück zum Zitat Han, L., Ong, H.Y.: Parallel data intensive applications using MapReduce: a data mining case study in biomedical sciences. Clust. Comput. 18, 403–418 (2015) Han, L., Ong, H.Y.: Parallel data intensive applications using MapReduce: a data mining case study in biomedical sciences. Clust. Comput. 18, 403–418 (2015)
27.
Zurück zum Zitat Chim, H., Deng, X.: Efficient phrase-based document similarity for clustering. IEEE Trans. Knowl. Data Eng. 20(9), 1217–1229 (2008)CrossRef Chim, H., Deng, X.: Efficient phrase-based document similarity for clustering. IEEE Trans. Knowl. Data Eng. 20(9), 1217–1229 (2008)CrossRef
Metadaten
Titel
A fast approach to identify trending articles in hot topics from XML based big bibliographic datasets
verfasst von
K. P. Swaraj
D. Manjula
Publikationsdatum
01.06.2016
Verlag
Springer US
Erschienen in
Cluster Computing / Ausgabe 2/2016
Print ISSN: 1386-7857
Elektronische ISSN: 1573-7543
DOI
https://doi.org/10.1007/s10586-016-0561-1

Weitere Artikel der Ausgabe 2/2016

Cluster Computing 2/2016 Zur Ausgabe

Premium Partner