Skip to main content

2022 | OriginalPaper | Buchkapitel

Sentence Classification to Detect Tables for Helping Extraction of Regulatory Interactions in Bacteria

verfasst von : Dante Sepúlveda, Joel Rodríguez-Herrera, Alfredo Varela-Vega, Axel Zagal Norman, Carlos-Francisco Méndez-Cruz

Erschienen in: Computational Intelligence Methods for Bioinformatics and Biostatistics

Verlag: Springer International Publishing

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

Abstract

The biomedical knowledge about transcriptional regulation in bacteria is rapidly published in scientific articles, so keeping biological databases up to date by manual curation is rather than impossible. Despite the efforts in biomedical text mining, there are still challenges in extracting regulatory interactions (RIs) between transcription factors and genes from text documents. One of them is produced by text extraction from PDF files. We have observed that the extraction of RIs from text lines that comes from tables of the original PDF article produces false positives. Here, we address the problem of automatically separating this text lines from those that are regular sentences by using automatic classification. Our best model was a Support Vector Classifier trained with n-grams of characters of tags of parts of speech, numbers, symbols, punctuation, brackets, and hyphens. Despite a significant imbalanced data, our classifier archived a positive class F1-score of 0.87. Our best classifier will be coupled eventually to a preprocessing pipeline for the automatic generation of transcriptional regulatory networks of bacteria by discarding text lines that comes from tables of the original PDF.

Sie haben noch keine Lizenz? Dann Informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 102.000 Bücher
  • über 537 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt Wissensvorsprung sichern!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 390 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe




 

Jetzt Wissensvorsprung sichern!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 67.000 Bücher
  • über 340 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko




Jetzt Wissensvorsprung sichern!

Literatur
1.
Zurück zum Zitat Alpaydin, E.: Introduction to Machine Learning. MIT Press, Cambridge (2020) Alpaydin, E.: Introduction to Machine Learning. MIT Press, Cambridge (2020)
2.
Zurück zum Zitat Angeli, G., Johnson Premkumar, M.J., Manning, C.D.: Leveraging linguistic structure for open domain information extraction. In: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, vol. 1, pp. 344–354. Association for Computational Linguistics, Beijing (2015). https://doi.org/10.3115/v1/P15-1034 Angeli, G., Johnson Premkumar, M.J., Manning, C.D.: Leveraging linguistic structure for open domain information extraction. In: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, vol. 1, pp. 344–354. Association for Computational Linguistics, Beijing (2015). https://​doi.​org/​10.​3115/​v1/​P15-1034
3.
Zurück zum Zitat Bekkar, M., Djemaa, H.K., Alitouche, T.A.: Evaluation measures for models assessment over imbalanced data sets. J. Inf. Eng. Appl. 3(10), 27–39 (2013) Bekkar, M., Djemaa, H.K., Alitouche, T.A.: Evaluation measures for models assessment over imbalanced data sets. J. Inf. Eng. Appl. 3(10), 27–39 (2013)
4.
Zurück zum Zitat Bergstra, J., Bengio, Y.: Random search for hyper-parameter optimization. J. Mach. Learn. Res. 13(2), 281–305 (2012) Bergstra, J., Bengio, Y.: Random search for hyper-parameter optimization. J. Mach. Learn. Res. 13(2), 281–305 (2012)
5.
Zurück zum Zitat Bishop, C.M.: Pattern Recognition and Machine Learning, p. 738. Springer, NY (2006) Bishop, C.M.: Pattern Recognition and Machine Learning, p. 738. Springer, NY (2006)
11.
Zurück zum Zitat Escorcia-Rodríguez, J.M., Tauch, A., Freyre-González, J.A.: Abasy Atlas v2.2: the most comprehensive and up-to-date inventory of meta-curated, historical, bacterial regulatory networks, their completeness and system-level characterization. Comput. Struct. Biotechnol. J. 18, 1228–1237 (2020). https://doi.org/10.1016/j.csbj.2020.05.015 Escorcia-Rodríguez, J.M., Tauch, A., Freyre-González, J.A.: Abasy Atlas v2.2: the most comprehensive and up-to-date inventory of meta-curated, historical, bacterial regulatory networks, their completeness and system-level characterization. Comput. Struct. Biotechnol. J. 18, 1228–1237 (2020). https://​doi.​org/​10.​1016/​j.​csbj.​2020.​05.​015
12.
Zurück zum Zitat Fàbrega, A., Vila, J.: Salmonella enterica serovar Typhimurium skills to succeed in the host: virulence and regulation. Clin. Microbiol. Rev. 26(2), 308–341 (2013)CrossRefPubMedPubMedCentral Fàbrega, A., Vila, J.: Salmonella enterica serovar Typhimurium skills to succeed in the host: virulence and regulation. Clin. Microbiol. Rev. 26(2), 308–341 (2013)CrossRefPubMedPubMedCentral
14.
Zurück zum Zitat Ferrario, A., Nagelin, M.: The art of natural language processing: classical, modern and contemporary approaches to text document classification. Modern and Contemporary Approaches to Text Document Classification (March 1, 2020) (2020) Ferrario, A., Nagelin, M.: The art of natural language processing: classical, modern and contemporary approaches to text document classification. Modern and Contemporary Approaches to Text Document Classification (March 1, 2020) (2020)
15.
Zurück zum Zitat Jeni, L., Cohn, J., De la Torre, F.: Facing imbalanced data – recommendations for the use of performance metrics. In: Proceedings - 2013 Humaine Association Conference on Affective Computing and Intelligent Interaction, ACII 2013, vol. 2013, pp. 245–251 (2013). https://doi.org/10.1109/ACII.2013.47 Jeni, L., Cohn, J., De la Torre, F.: Facing imbalanced data – recommendations for the use of performance metrics. In: Proceedings - 2013 Humaine Association Conference on Affective Computing and Intelligent Interaction, ACII 2013, vol. 2013, pp. 245–251 (2013). https://​doi.​org/​10.​1109/​ACII.​2013.​47
17.
Zurück zum Zitat Konheim, A.G.: Cryptography, a Primer. Wiley, Chichester (1981) Konheim, A.G.: Cryptography, a Primer. Wiley, Chichester (1981)
18.
Zurück zum Zitat Kubat, M., Matwin, S., et al.: Addressing the curse of imbalanced training sets: one-sided selection. In: Icml, vol. 97, p. 179. Citeseer (1997) Kubat, M., Matwin, S., et al.: Addressing the curse of imbalanced training sets: one-sided selection. In: Icml, vol. 97, p. 179. Citeseer (1997)
19.
Zurück zum Zitat Lemaître, G., Nogueira, F., Aridas, C.K.: Imbalanced-learn: a Python toolbox to tackle the curse of imbalanced datasets in machine learning. J. Mach. Learn. Res. 18(17), 1–5 (2017) Lemaître, G., Nogueira, F., Aridas, C.K.: Imbalanced-learn: a Python toolbox to tackle the curse of imbalanced datasets in machine learning. J. Mach. Learn. Res. 18(17), 1–5 (2017)
20.
Zurück zum Zitat Liu, Y., Bai, K., Mitra, P., Giles, C.L.: TableSeer: automatic table metadata extraction and searching in digital libraries. In: Proceedings of the 7th ACM/IEEE-CS Joint Conference on Digital Libraries, pp. 91–100 (2007) Liu, Y., Bai, K., Mitra, P., Giles, C.L.: TableSeer: automatic table metadata extraction and searching in digital libraries. In: Proceedings of the 7th ACM/IEEE-CS Joint Conference on Digital Libraries, pp. 91–100 (2007)
21.
Zurück zum Zitat Lusa, L., et al.: Joint use of over-and under-sampling techniques and cross-validation for the development and assessment of prediction models. BMC Bioinform. 16(1), 1–10 (2015) Lusa, L., et al.: Joint use of over-and under-sampling techniques and cross-validation for the development and assessment of prediction models. BMC Bioinform. 16(1), 1–10 (2015)
22.
Zurück zum Zitat Marcus, M.P., Marcinkiewicz, M.A., Santorini, B.: Building a large annotated corpus of English: the Penn treebank. Comput. Linguist. 19(2), 313–330 (1993) Marcus, M.P., Marcinkiewicz, M.A., Santorini, B.: Building a large annotated corpus of English: the Penn treebank. Comput. Linguist. 19(2), 313–330 (1993)
25.
Zurück zum Zitat Pedregosa, F., et al.: Scikit-learn: machine learning in Python. J. Mach. Learn. Res. 12(85), 2825–2830 (2011) Pedregosa, F., et al.: Scikit-learn: machine learning in Python. J. Mach. Learn. Res. 12(85), 2825–2830 (2011)
26.
Zurück zum Zitat Pinto, D., McCallum, A., Wei, X., Croft, W.B.: Table extraction using conditional random fields. In: Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, pp. 235–242 (2003) Pinto, D., McCallum, A., Wei, X., Croft, W.B.: Table extraction using conditional random fields. In: Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, pp. 235–242 (2003)
27.
Zurück zum Zitat Qi, P., Zhang, Y., Zhang, Y., Bolton, J., Manning, C.D.: Stanza: a Python natural language processing toolkit for many human languages. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (2020) Qi, P., Zhang, Y., Zhang, Y., Bolton, J., Manning, C.D.: Stanza: a Python natural language processing toolkit for many human languages. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (2020)
29.
Zurück zum Zitat Sparck Jones, K.: A statistical interpretation of term specificity and its application in retrieval. J. Doc. 28(1), 11–21 (1972)CrossRef Sparck Jones, K.: A statistical interpretation of term specificity and its application in retrieval. J. Doc. 28(1), 11–21 (1972)CrossRef
33.
Zurück zum Zitat Yoon, H., Lim, S., Heu, S., Choi, S., Ryu, S.: Proteome analysis of Salmonella enterica serovar Typhimurium fis mutant. FEMS Microbiol. Lett. 226(2), 391–396 (2003)CrossRefPubMed Yoon, H., Lim, S., Heu, S., Choi, S., Ryu, S.: Proteome analysis of Salmonella enterica serovar Typhimurium fis mutant. FEMS Microbiol. Lett. 226(2), 391–396 (2003)CrossRefPubMed
34.
Zurück zum Zitat Zhai, Z., et al.: ChemTables: a dataset for semantic classification on tables in chemical patents. J. Cheminformatics 13(1), 97 (2021)CrossRef Zhai, Z., et al.: ChemTables: a dataset for semantic classification on tables in chemical patents. J. Cheminformatics 13(1), 97 (2021)CrossRef
Metadaten
Titel
Sentence Classification to Detect Tables for Helping Extraction of Regulatory Interactions in Bacteria
verfasst von
Dante Sepúlveda
Joel Rodríguez-Herrera
Alfredo Varela-Vega
Axel Zagal Norman
Carlos-Francisco Méndez-Cruz
Copyright-Jahr
2022
DOI
https://doi.org/10.1007/978-3-031-20837-9_12

Premium Partner