Skip to main content
Erschienen in: Automatic Documentation and Mathematical Linguistics 1/2024

01.02.2024 | NATURAL LANGUAGE PROCESSING

Topic Modeling for Mining Opinion Aspects from a Customer Feedback Corpus

verfasst von: O. I. Babina

Erschienen in: Automatic Documentation and Mathematical Linguistics | Ausgabe 1/2024

Einloggen, um Zugang zu erhalten

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

Abstract

The paper introduces a methodology for extracting opinion aspects from textual content by identifying the customer-evaluated parameters regarding a given object. These parameters form the foundation for shaping the customer’s attitudes toward the product or service. The proposed approach leverages topic modeling tools to delineate classes of vocabulary exhibiting semantics aligned with the parameters influencing the customer’s opinion about the object. Our study specifically explores the application of the BERTopic model as a topic modeling tool to address this challenge. The outlined methodology encompasses several sequential steps, including the preprocessing of textual data involving the removal of stopwords, conversion to lowercase characters, and lemmatization. Additionally, special consideration is given to the distinct lexical manifestations of opinion aspects, obtained as a result of the extraction of nominal, verbal, and adjectival single- and multicomponent phrases from the corpus. Subsequently, the corpus sentences are represented as vectors in a feature space expressed by the extracted words and phrases. The final step involves the application of topic modeling using the BERTopic model on the customer review corpus, utilizing the vector representations of corpus sentences. The experimental inquiry is conducted on a domain-specific Russian-language corpus comprising customer feedback on airline services gathered from customer review websites. The resultant topic distribution is then juxtaposed against a manually constructed conceptual model of the domain. The comparative analysis reveals that the automatic topic distribution aligns with the conceptual structure of the domain, demonstrating a precision of 0.955 and a recall of 0.875. These findings affirm the efficacy of employing the BERTopic model to address the problem of the corpus-based mining of opinion aspects.
Literatur
7.
Zurück zum Zitat Semina, T.A., Sentiment analysis: Modern approaches and existing problems, Sotsial’nye Gumanitarnye Nauki. Otechestvennaya Zarubezhnaya Literatura. Ser. 6: Yazykoznanie. Referativnyi Zh., 2020, no. 4, pp. 47–63. Semina, T.A., Sentiment analysis: Modern approaches and existing problems, Sotsial’nye Gumanitarnye Nauki. Otechestvennaya Zarubezhnaya Literatura. Ser. 6: Yazykoznanie. Referativnyi Zh., 2020, no. 4, pp. 47–63.
11.
13.
Zurück zum Zitat Lark, J., Morin, E., and Saldarriaga, S.P., A comparative study of target-based and entity-based opinion extraction, Computational Linguistics and Intelligent Text Processing. CICLing 2017, Gelbukh, A., Ed., Lecture Notes in Computer Science, vol. 10762, Cham: Springer, 2017, pp. 211–223. https://doi.org/10.1007/978-3-319-77116-8_16CrossRef Lark, J., Morin, E., and Saldarriaga, S.P., A comparative study of target-based and entity-based opinion extraction, Computational Linguistics and Intelligent Text Processing. CICLing 2017, Gelbukh, A., Ed., Lecture Notes in Computer Science, vol. 10762, Cham: Springer, 2017, pp. 211–223. https://​doi.​org/​10.​1007/​978-3-319-77116-8_​16CrossRef
14.
Zurück zum Zitat Xu, R., Lin, H., Liao, M., Han, X., Xu, J., Tan, W., Sun, Y., and Sun, L., ECO v1: Towards event-centric opinion mining, findings of the, Findings of the Association for Computational Linguistics: ACL 2022, Dublin, 2022, Muresan, S., Nakov, P., and Villvicencio, A., Eds., Association for Computational Linguistics, 2022, pp. 2743–2753. https://doi.org/10.18653/v1/2022.findings-acl.216 Xu, R., Lin, H., Liao, M., Han, X., Xu, J., Tan, W., Sun, Y., and Sun, L., ECO v1: Towards event-centric opinion mining, findings of the, Findings of the Association for Computational Linguistics: ACL 2022, Dublin, 2022, Muresan, S., Nakov, P., and Villvicencio, A., Eds., Association for Computational Linguistics, 2022, pp. 2743–2753. https://​doi.​org/​10.​18653/​v1/​2022.​findings-acl.​216
19.
Zurück zum Zitat Arora, P., Bakliwal, A., and Varma, V., Hindi subjective lexicon generation using WordNet graph traversal, Int. J. Comput. Linguist. Appl., 2012, vol. 3, no. 1, pp. 25–39. Arora, P., Bakliwal, A., and Varma, V., Hindi subjective lexicon generation using WordNet graph traversal, Int. J. Comput. Linguist. Appl., 2012, vol. 3, no. 1, pp. 25–39.
21.
Zurück zum Zitat Loukachevitch, N. and Levchik, A., Creating a general Russian sentiment lexicon, Proc. Tenth Int. Conf. on Language Resources and Evaluation (LREC’16), Portorož, Slovenia, 2016, Calzolari, N. et al., Eds., European Language Resources Association, 2016, pp. 1171–1176. https://aclanthology.org/L16-1186. Loukachevitch, N. and Levchik, A., Creating a general Russian sentiment lexicon, Proc. Tenth Int. Conf. on Language Resources and Evaluation (LREC’16), Portorož, Slovenia, 2016, Calzolari, N. et al., Eds., European Language Resources Association, 2016, pp. 1171–1176. https://​aclanthology.​org/​L16-1186.​
22.
Zurück zum Zitat Koltsova, O., Alexeeva, S., and Kolcov, S., An opinion word lexicon and a training dataset for Russian sentiment analysis of social media, Komp’yuternaya lingvistika i intellektual’nye tekhnologii: po materialam ezhegodnoi mezhdunarodnoi konferentsii Dialog-2016 (Computational Linguistics and Intellectual Technologies: Proc. Int. Conf. Dialogue 2016), Moscow, 2016, Moscow: Izd-vo Ros. Gos. Gumanit. Univ., 2016, pp. 277–287. Koltsova, O., Alexeeva, S., and Kolcov, S., An opinion word lexicon and a training dataset for Russian sentiment analysis of social media, Komp’yuternaya lingvistika i intellektual’nye tekhnologii: po materialam ezhegodnoi mezhdunarodnoi konferentsii Dialog-2016 (Computational Linguistics and Intellectual Technologies: Proc. Int. Conf. Dialogue 2016), Moscow, 2016, Moscow: Izd-vo Ros. Gos. Gumanit. Univ., 2016, pp. 277–287.
23.
Zurück zum Zitat Kan, D., Rule-based approach to sentiment analysis at ROMIP 2011: Contest on sentiment analysis at the International Conference Dialogue-2011, 2012. https:// www.dialog-21.ru/media/1393/138.pdf. Kan, D., Rule-based approach to sentiment analysis at ROMIP 2011: Contest on sentiment analysis at the International Conference Dialogue-2011, 2012. https:// www.dialog-21.ru/media/1393/138.pdf.
27.
Zurück zum Zitat Agarwal, A., Xie, B., Vovsha, I., Rambow, O., and Passonneau, R., Sentiment analysis of twitter data, Proc. Workshop on Language in Social Media (LSM 2011), Portland, Ore., 2011, Nagarajan, M. and Gamon, M., Eds., Association for Computational Linguistics, 2011, pp. 30–38. https://aclanthology.org/W11-0705. Agarwal, A., Xie, B., Vovsha, I., Rambow, O., and Passonneau, R., Sentiment analysis of twitter data, Proc. Workshop on Language in Social Media (LSM 2011), Portland, Ore., 2011, Nagarajan, M. and Gamon, M., Eds., Association for Computational Linguistics, 2011, pp. 30–38. https://​aclanthology.​org/​W11-0705.​
28.
Zurück zum Zitat Turney, P.D., Thumbs up or thumbs down?: Semantic orientation applied to unsupervised classification of reviews, Proc. 40th Annu. Meeting on Association for Computational Linguistics, Philadelphia, 2002, Isabelle, P., Charniak, E., and Lin, D., Eds., Association for Computational Linguistics, 2002, pp. 417–424. https://doi.org/10.3115/1073083.1073153 Turney, P.D., Thumbs up or thumbs down?: Semantic orientation applied to unsupervised classification of reviews, Proc. 40th Annu. Meeting on Association for Computational Linguistics, Philadelphia, 2002, Isabelle, P., Charniak, E., and Lin, D., Eds., Association for Computational Linguistics, 2002, pp. 417–424. https://​doi.​org/​10.​3115/​1073083.​1073153
29.
Zurück zum Zitat Zhang, L. and Liu, B., Aspect and entity extraction for opinion mining, data mining and knowledge discovery for big data, Data Mining and Knowledge Discovery for Big Data. Studies in Big Data, Chu, W.W., Ed., Studies in Big Data, vol. 1, Berlin: Springer, 2014, pp. 1–40. https://doi.org/10.1007/978-3-642-40837-3_1 Zhang, L. and Liu, B., Aspect and entity extraction for opinion mining, data mining and knowledge discovery for big data, Data Mining and Knowledge Discovery for Big Data. Studies in Big Data, Chu, W.W., Ed., Studies in Big Data, vol. 1, Berlin: Springer, 2014, pp. 1–40. https://​doi.​org/​10.​1007/​978-3-642-40837-3_​1
30.
Zurück zum Zitat Roi, D.A. and Efremova, N.E., Methods for extracting aspectual terms from opinions, Nov. Inf. Tekhnol. Avtomatizirovannykh Sistemakh, 2018, no. 21, pp. 212–216. Roi, D.A. and Efremova, N.E., Methods for extracting aspectual terms from opinions, Nov. Inf. Tekhnol. Avtomatizirovannykh Sistemakh, 2018, no. 21, pp. 212–216.
31.
Zurück zum Zitat Golubev, A. and Loukachevitch, N., Improving results on Russian sentiment datasets, Artificial Intelligence and Natural Language, Filchenkov, A., Kauttonen, J., and Pivovarova, L., Eds., Communications in Computer and Information Science, Cham: Springer, 2020, pp. 109–121. https://doi.org/10.1007/978-3-030-59082-6_8 Golubev, A. and Loukachevitch, N., Improving results on Russian sentiment datasets, Artificial Intelligence and Natural Language, Filchenkov, A., Kauttonen, J., and Pivovarova, L., Eds., Communications in Computer and Information Science, Cham: Springer, 2020, pp. 109–121. https://​doi.​org/​10.​1007/​978-3-030-59082-6_​8
36.
Zurück zum Zitat Blei, D.M., Ng, A.Y., and Jordan, M.I., Latent Dirichlet allocation, J. Mach. Learn. Res., 2003, vol. 3, no. 2, pp. 993–1022. Blei, D.M., Ng, A.Y., and Jordan, M.I., Latent Dirichlet allocation, J. Mach. Learn. Res., 2003, vol. 3, no. 2, pp. 993–1022.
49.
Zurück zum Zitat Reimers, N. and Gurevych, I., Sentence-BERT: Sentence embeddings using Siamese BERT-networks, Proc. 2019 Conf. on Empirical Methods in Natural Language Processing, Hong Kong, 2019, Inui, K., Jiang, J., Ng, V., and Wan, X., Eds., Association for Computational Linguistics, 2019, pp. 3982–3992. https://doi.org/10.18653/v1/D19-1410 Reimers, N. and Gurevych, I., Sentence-BERT: Sentence embeddings using Siamese BERT-networks, Proc. 2019 Conf. on Empirical Methods in Natural Language Processing, Hong Kong, 2019, Inui, K., Jiang, J., Ng, V., and Wan, X., Eds., Association for Computational Linguistics, 2019, pp. 3982–3992. https://​doi.​org/​10.​18653/​v1/​D19-1410
52.
54.
Zurück zum Zitat Gerasimenko, N., Chernyavskiy, A., Nikiforova, M., Ianina, A., and Vorontsov, K., Incremental topic modeling for scientific trend topics extraction, Komp’yuternaya lingvistika i intellektual’nye tekhnologii: Po materialam ezhegodnoi mezhdunarodnoi konferentsii Dialog-2023 (Computational Linguistics and Intellectual Technologies: Proc. Int. Conf. Dialogue 2023), Moscow, 2023, Moscow: 2023, pp. 88–103. https://www. dialog-21.ru/media/5893/gerasimenkonplusetal012.pdf. Gerasimenko, N., Chernyavskiy, A., Nikiforova, M., Ianina, A., and Vorontsov, K., Incremental topic modeling for scientific trend topics extraction, Komp’yuternaya lingvistika i intellektual’nye tekhnologii: Po materialam ezhegodnoi mezhdunarodnoi konferentsii Dialog-2023 (Computational Linguistics and Intellectual Technologies: Proc. Int. Conf. Dialogue 2023), Moscow, 2023, Moscow: 2023, pp. 88–103. https://​www.​ dialog-21.ru/media/5893/gerasimenkonplusetal012.pdf.
55.
Zurück zum Zitat Udupa, A., Adarsh, K.N., Aravinda, A., Godihal, N.H., and Kayarvizhy, N., An exploratory analysis of GSDMM and BERTopic on short text topic modelling, Fourth Int. Conf. on Cognitive Computing and Information Processing (CCIP-2022), Bengaluru, India, 2022, IEEE, 2022, pp. 1–9. https://doi.org/10.1109/CCIP57447.2022.10058687 Udupa, A., Adarsh, K.N., Aravinda, A., Godihal, N.H., and Kayarvizhy, N., An exploratory analysis of GSDMM and BERTopic on short text topic modelling, Fourth Int. Conf. on Cognitive Computing and Information Processing (CCIP-2022), Bengaluru, India, 2022, IEEE, 2022, pp. 1–9. https://​doi.​org/​10.​1109/​CCIP57447.​2022.​10058687
57.
Zurück zum Zitat Hu, M. and Liu, B., Mining opinion features in customer reviews, Proc. 19th Natl. Conf. on Artificial Intelligence, San Jose, Calif., 2004, Cohn, A.G., Ed., AAAI Press, 2004, pp. 755–760. Hu, M. and Liu, B., Mining opinion features in customer reviews, Proc. 19th Natl. Conf. on Artificial Intelligence, San Jose, Calif., 2004, Cohn, A.G., Ed., AAAI Press, 2004, pp. 755–760.
58.
Zurück zum Zitat Yi, J., Nasukawa, T., Bunescu, R., and Niblack, W., Sentiment analyzer: Extracting sentiments about a given topic using natural language processing techniques, Proc. IEEE Int. Conf. on Data Mining (ICDM), Melbourne, Fla., IEEE, 2003, pp. 427–434. https://doi.org/10.1109/ICDM.2003.1250949 Yi, J., Nasukawa, T., Bunescu, R., and Niblack, W., Sentiment analyzer: Extracting sentiments about a given topic using natural language processing techniques, Proc. IEEE Int. Conf. on Data Mining (ICDM), Melbourne, Fla., IEEE, 2003, pp. 427–434. https://​doi.​org/​10.​1109/​ICDM.​2003.​1250949
59.
Zurück zum Zitat Sheremetyeva, S.O., Extraction of multicomponent terms and keywords from multilingual patent documentation, Nauchn.-Tekhn. Inform., Ser. 2. Protsessy Sist., 2019, no. 4, pp. 25–33. Sheremetyeva, S.O., Extraction of multicomponent terms and keywords from multilingual patent documentation, Nauchn.-Tekhn. Inform., Ser. 2. Protsessy Sist., 2019, no. 4, pp. 25–33.
60.
Zurück zum Zitat Korobov, M., Morphological analyzer and generator for Russian and Ukrainian languages, Analysis of Images, Social Networks and Texts, Khachay, M., Konstantinova, N., Panchenko, A., Ignatov, D., and Labunets, V., Eds., Communications in Computer and Information Science, vol. 542, Cham: Springer, 2015, pp. 320–332. https://doi.org/10.1007/978-3-319-26123-2_31 Korobov, M., Morphological analyzer and generator for Russian and Ukrainian languages, Analysis of Images, Social Networks and Texts, Khachay, M., Konstantinova, N., Panchenko, A., Ignatov, D., and Labunets, V., Eds., Communications in Computer and Information Science, vol. 542, Cham: Springer, 2015, pp. 320–332. https://​doi.​org/​10.​1007/​978-3-319-26123-2_​31
Metadaten
Titel
Topic Modeling for Mining Opinion Aspects from a Customer Feedback Corpus
verfasst von
O. I. Babina
Publikationsdatum
01.02.2024
Verlag
Pleiades Publishing
Erschienen in
Automatic Documentation and Mathematical Linguistics / Ausgabe 1/2024
Print ISSN: 0005-1055
Elektronische ISSN: 1934-8371
DOI
https://doi.org/10.3103/S0005105524010060

Weitere Artikel der Ausgabe 1/2024

Automatic Documentation and Mathematical Linguistics 1/2024 Zur Ausgabe

Premium Partner