Skip to main content
Erschienen in:
Buchtitelbild

2023 | OriginalPaper | Buchkapitel

Threshold Text Classification with Kullback–Leibler Divergence Approach

verfasst von : Hiep Xuan Huynh, Cang Anh Phan, Tu Cam Thi Tran, Hai Thanh Nguyen, Dinh Quoc Truong

Erschienen in: Machine Learning and Mechanics Based Soft Computing Applications

Verlag: Springer Nature Singapore

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

Abstract

Text classification based on thresholds belongs to the supervised learning method which assigns text material to predefined classes or categories based on different thresholds with divergence approach. These categories are identified by a set of documents trained by an automated algorithm. This work presents an approach of text classification using an automatic keyword extraction algorithm based on the Kullback–Leibler divergence approach. The proposed method is evaluated on 2000 documents in Vietnamese, covering ten topics, collected from various e-journals and news portal Web sites including vietnamnet.vn, vnexpress.net, and so on to generate a completely new set of keywords. Such keywords, then, are leveraged to categorize the topic of new text documents. The obtained results verifying the practicality of our approach are feasible as well as outperform the state-of-the-art method.

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 102.000 Bücher
  • über 537 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 390 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe




 

Jetzt Wissensvorsprung sichern!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 340 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko




Jetzt Wissensvorsprung sichern!

Literatur
1.
Zurück zum Zitat Liu, F., Pennell, D., Liu, F., & Liu, Y. (2009). Unsupervised approaches for automatic keyword extraction using meeting transcripts. In Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the ACL, NAACL’09, (June 2009) (pp. 620–628), Boulder, Colorado. Liu, F., Pennell, D., Liu, F., & Liu, Y. (2009). Unsupervised approaches for automatic keyword extraction using meeting transcripts. In Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the ACL, NAACL’09, (June 2009) (pp. 620–628), Boulder, Colorado.
2.
Zurück zum Zitat Nguyen, C. T., Nguyen, T. K., Phan, X. H., Nguyen, L. M., & Ha, Q. T. (2006). Vietnamese word segmentation with CRFs and SVMs: an investigation. In Proceedings of the 20th Pacific Asia Conference on Language, Information and Computation (PACLIC 2006). Nguyen, C. T., Nguyen, T. K., Phan, X. H., Nguyen, L. M., & Ha, Q. T. (2006). Vietnamese word segmentation with CRFs and SVMs: an investigation. In Proceedings of the 20th Pacific Asia Conference on Language, Information and Computation (PACLIC 2006).
3.
Zurück zum Zitat Matsuo, Y., & Ishizuka, M. (2004). Keyword extraction from a single document using word co-occurrence statistical information. International Journal on AI Tools, 13(1), 157–169.CrossRef Matsuo, Y., & Ishizuka, M. (2004). Keyword extraction from a single document using word co-occurrence statistical information. International Journal on AI Tools, 13(1), 157–169.CrossRef
4.
Zurück zum Zitat Zaïane, O. R., & Antonie, M.-L. (2002). Classifying text documents by associating terms with text categories. Australian Computer Science and Communications, 24(2), 215–222. Zaïane, O. R., & Antonie, M.-L. (2002). Classifying text documents by associating terms with text categories. Australian Computer Science and Communications, 24(2), 215–222.
5.
Zurück zum Zitat Truong, Q. D., Huynh, H. X., & Nguyen, C. N. (2016). An abstract-based approach for text classification. In Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering (Vol 168, pp. 237–245). Springer. https://doi.org/10.1007/978-3-319-46909-6_22 Truong, Q. D., Huynh, H. X., & Nguyen, C. N. (2016). An abstract-based approach for text classification. In Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering (Vol 168, pp. 237–245). Springer. https://​doi.​org/​10.​1007/​978-3-319-46909-6_​22
7.
Zurück zum Zitat Han, E.-H. (Sam), Karypis, G., & Kumar, V. (2001). Text categorization using weight adjusted k-nearest neighbor classification. Springer Han, E.-H. (Sam), Karypis, G., & Kumar, V. (2001). Text categorization using weight adjusted k-nearest neighbor classification. Springer
8.
Zurück zum Zitat Dinh, Q. T., Le, H. P., Nguyen, T. M. H., Nguyen, C. T., & Mathias Rossignol, et al. (2008). Word segmentation of Vietnamese texts: A comparison of approaches. In 6th International Conference on Language Resources and Evaluation-LREC 2008, (May 2008), Marrakech, Morocco. Dinh, Q. T., Le, H. P., Nguyen, T. M. H., Nguyen, C. T., & Mathias Rossignol, et al. (2008). Word segmentation of Vietnamese texts: A comparison of approaches. In 6th International Conference on Language Resources and Evaluation-LREC 2008, (May 2008), Marrakech, Morocco.
9.
Zurück zum Zitat Lai, K. P., Ho, J. C. S., & Lam, W. (2020). Cross-domain sentiment classification using topic attention and dual-task adversarial training. In Artificial Neural Networks and Machine Learning—ICANN 2020. ICANN 2020. Lecture Notes in Computer Science (Vol. 12397). Springer. https://doi.org/10.1007/978-3-030-61616-8_46 Lai, K. P., Ho, J. C. S., & Lam, W. (2020). Cross-domain sentiment classification using topic attention and dual-task adversarial training. In Artificial Neural Networks and Machine Learning—ICANN 2020. ICANN 2020. Lecture Notes in Computer Science (Vol. 12397). Springer. https://​doi.​org/​10.​1007/​978-3-030-61616-8_​46
11.
Zurück zum Zitat Aggarwal, A. G. (2018). A multi-attribute online advertising budget allocation under uncertain preferences. Ingeniería Solidaria, 14(25), 1–10.MathSciNetCrossRef Aggarwal, A. G. (2018). A multi-attribute online advertising budget allocation under uncertain preferences. Ingeniería Solidaria, 14(25), 1–10.MathSciNetCrossRef
12.
Zurück zum Zitat Terrance, A. R., Shrivastava, S., Kumari, A., & Sivanandam, L. (2018). Competitive analysis of retail websites through search engine marketing. Ingeniería Solidaria, 14(25), 1–14.CrossRef Terrance, A. R., Shrivastava, S., Kumari, A., & Sivanandam, L. (2018). Competitive analysis of retail websites through search engine marketing. Ingeniería Solidaria, 14(25), 1–14.CrossRef
13.
Zurück zum Zitat Lopez-Inga, M. E., & Guerrero-Huaranga, R. M. (2018). Cloud business intelligence and analytics model for SMES in the retail sector in Peru/Modelo de inteligencia de negocios y analitica en la nube para pymes del sector retail en Peru/Modelo de inteligencia de negocios e analitica em nuvem para pmes do setor varejista no Peru. Revista Ingenieria Solidaria, 14(24). Lopez-Inga, M. E., & Guerrero-Huaranga, R. M. (2018). Cloud business intelligence and analytics model for SMES in the retail sector in Peru/Modelo de inteligencia de negocios y analitica en la nube para pymes del sector retail en Peru/Modelo de inteligencia de negocios e analitica em nuvem para pmes do setor varejista no Peru. Revista Ingenieria Solidaria, 14(24).
14.
Zurück zum Zitat Gupta, M., Solanki, V. K., & Singh, V. K. (2017). A novel framework to use association rule mining for classification of traffic accident severity. Ingeniería solidaria, 13(21), 37–44.CrossRef Gupta, M., Solanki, V. K., & Singh, V. K. (2017). A novel framework to use association rule mining for classification of traffic accident severity. Ingeniería solidaria, 13(21), 37–44.CrossRef
21.
Zurück zum Zitat Wallach, H. M. (2006). Topic modeling: Beyond bag-of-words. In Proceedings of the 23rd International Conference on Machine Learning, ICML (pp. 977–984), New York, NY, USA. ACM Wallach, H. M. (2006). Topic modeling: Beyond bag-of-words. In Proceedings of the 23rd International Conference on Machine Learning, ICML (pp. 977–984), New York, NY, USA. ACM
22.
Zurück zum Zitat Thi Tran, T. C., Huynh, H. X., Tran, P. Q., & Truong, D. Q. (2019). Text classification based on keywords with different thresholds. In ACM International Conference Proceeding Series. Thi Tran, T. C., Huynh, H. X., Tran, P. Q., & Truong, D. Q. (2019). Text classification based on keywords with different thresholds. In ACM International Conference Proceeding Series.
Metadaten
Titel
Threshold Text Classification with Kullback–Leibler Divergence Approach
verfasst von
Hiep Xuan Huynh
Cang Anh Phan
Tu Cam Thi Tran
Hai Thanh Nguyen
Dinh Quoc Truong
Copyright-Jahr
2023
Verlag
Springer Nature Singapore
DOI
https://doi.org/10.1007/978-981-19-6450-3_2

Premium Partner