Skip to main content

Tipp

Weitere Kapitel dieses Buchs durch Wischen aufrufen

2020 | OriginalPaper | Buchkapitel

Modelling Patient Sequences for Rare Disease Detection with Semi-supervised Generative Adversarial Nets

verfasst von : Kezi Yu, Yunlong Wang, Yong Cai

Erschienen in: Advanced Analytics and Learning on Temporal Data

Verlag: Springer International Publishing

share
TEILEN

Abstract

Rare diseases affect 350 million patients worldwide, but they are commonly delayed in diagnosis or misdiagnosed. The problem of detecting rare disease faces two main challenges: the first being extreme imbalance of data and the second being finding the appropriate features. In this paper, we propose to address the problems by using semi-supervised generative adversarial networks (GANs) to deal with the data imbalance issue and recurrent neural networks (RNNs) to directly model patient sequences. We experimented with detecting patients with a particular rare disease (exocrine pancreatic insufficiency, EPI). The dataset includes 1.8 million patients with 29,149 patients being positive, from a large longitudinal study using 7 years medical claims. Our model achieved 0.56 PR-AUC and outperformed benchmark models in terms of precision and recall.

Sie möchten Zugang zu diesem Inhalt erhalten? Dann informieren Sie sich jetzt über unsere Produkte:

Springer Professional "Wirtschaft+Technik"

Online-Abonnement

Mit Springer Professional "Wirtschaft+Technik" erhalten Sie Zugriff auf:

  • über 69.000 Bücher
  • über 500 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Maschinenbau + Werkstoffe
  • Versicherung + Risiko

Jetzt 90 Tage mit der neuen Mini-Lizenz testen!

Springer Professional "Technik"

Online-Abonnement

Mit Springer Professional "Technik" erhalten Sie Zugriff auf:

  • über 50.000 Bücher
  • über 380 Zeitschriften

aus folgenden Fachgebieten:

  • Automobil + Motoren
  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Elektrotechnik + Elektronik
  • Energie + Nachhaltigkeit
  • Maschinenbau + Werkstoffe



 


Jetzt 90 Tage mit der neuen Mini-Lizenz testen!

Springer Professional "Wirtschaft"

Online-Abonnement

Mit Springer Professional "Wirtschaft" erhalten Sie Zugriff auf:

  • über 58.000 Bücher
  • über 300 Zeitschriften

aus folgenden Fachgebieten:

  • Bauwesen + Immobilien
  • Business IT + Informatik
  • Finance + Banking
  • Management + Führung
  • Marketing + Vertrieb
  • Versicherung + Risiko





Jetzt 90 Tage mit der neuen Mini-Lizenz testen!

Literatur
1.
Zurück zum Zitat Bai, T., Zhang, S., Egleston, B.L., Vucetic, S.: Interpretable representation learning for healthcare via capturing disease progression through time. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. pp. 43–51. ACM (2018) Bai, T., Zhang, S., Egleston, B.L., Vucetic, S.: Interpretable representation learning for healthcare via capturing disease progression through time. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. pp. 43–51. ACM (2018)
2.
Zurück zum Zitat Boat, T.F., Field, M.J., et al.: Rare Diseases and Orphan Products: Accelerating Research and Development. National Academies Press, Washington, DC (2011) Boat, T.F., Field, M.J., et al.: Rare Diseases and Orphan Products: Accelerating Research and Development. National Academies Press, Washington, DC (2011)
3.
Zurück zum Zitat Cameron, M.J., Horst, M., Lawhorne, L.W., Lichtenberg, P.A.: Evaluation of academic detailing for primary care physician dementia education. Am. J. Alzheimer’s Dis. Other Dement.® 25(4), 333–339 (2010) CrossRef Cameron, M.J., Horst, M., Lawhorne, L.W., Lichtenberg, P.A.: Evaluation of academic detailing for primary care physician dementia education. Am. J. Alzheimer’s Dis. Other Dement.® 25(4), 333–339 (2010) CrossRef
4.
Zurück zum Zitat Che, Z., Purushotham, S., Khemani, R.G., Liu, Y.: Interpretable deep models for ICU outcome prediction. In: AMIA Annual Symposium Proceedings. AMIA Symposium, vol. 2016, pp. 371–380 (2016) Che, Z., Purushotham, S., Khemani, R.G., Liu, Y.: Interpretable deep models for ICU outcome prediction. In: AMIA Annual Symposium Proceedings. AMIA Symposium, vol. 2016, pp. 371–380 (2016)
5.
Zurück zum Zitat Ching, T., et al.: Opportunities and obstacles for deep learning in biology and medicine. J. R. Soc. Interface 15(141), 20170387 (2018) CrossRef Ching, T., et al.: Opportunities and obstacles for deep learning in biology and medicine. J. R. Soc. Interface 15(141), 20170387 (2018) CrossRef
6.
Zurück zum Zitat Cho, K., Van Merriënboer, B., Bahdanau, D., Bengio, Y.: On the properties of neural machine translation: encoder-decoder approaches. arXiv preprint arXiv:​1409.​1259 (2014) Cho, K., Van Merriënboer, B., Bahdanau, D., Bengio, Y.: On the properties of neural machine translation: encoder-decoder approaches. arXiv preprint arXiv:​1409.​1259 (2014)
7.
Zurück zum Zitat Choi, E., Bahadori, M.T., Schuetz, A., Stewart, W.F., Sun, J.: Doctor AI: predicting clinical events via recurrent neural networks. In: Machine Learning for Healthcare Conference, pp. 301–318 (2016) Choi, E., Bahadori, M.T., Schuetz, A., Stewart, W.F., Sun, J.: Doctor AI: predicting clinical events via recurrent neural networks. In: Machine Learning for Healthcare Conference, pp. 301–318 (2016)
8.
Zurück zum Zitat Choi, E., Bahadori, M.T., Sun, J., Kulas, J., Schuetz, A., Stewart, W.: Retain: an interpretable predictive model for healthcare using reverse time attention mechanism. In: Advances in Neural Information Processing Systems, pp. 3504–3512 (2016) Choi, E., Bahadori, M.T., Sun, J., Kulas, J., Schuetz, A., Stewart, W.: Retain: an interpretable predictive model for healthcare using reverse time attention mechanism. In: Advances in Neural Information Processing Systems, pp. 3504–3512 (2016)
9.
Zurück zum Zitat Choi, E., Schuetz, A., Stewart, W.F., Sun, J.: Using recurrent neural network models for early detection of heart failure onset. J. Am. Med. Inform. Assoc. 24(2), 361–370 (2016) Choi, E., Schuetz, A., Stewart, W.F., Sun, J.: Using recurrent neural network models for early detection of heart failure onset. J. Am. Med. Inform. Assoc. 24(2), 361–370 (2016)
10.
Zurück zum Zitat Chung, J., Gulcehre, C., Cho, K., Bengio, Y.: Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:​1412.​3555 (2014) Chung, J., Gulcehre, C., Cho, K., Bengio, Y.: Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:​1412.​3555 (2014)
11.
Zurück zum Zitat Dai, Z., Yang, Z., Yang, F., Cohen, W.W., Salakhutdinov, R.R.: Good semi-supervised learning that requires a bad GAN. In: Advances in Neural Information Processing Systems, pp. 6510–6520 (2017) Dai, Z., Yang, Z., Yang, F., Cohen, W.W., Salakhutdinov, R.R.: Good semi-supervised learning that requires a bad GAN. In: Advances in Neural Information Processing Systems, pp. 6510–6520 (2017)
12.
Zurück zum Zitat Ghassemi, M., Naumann, T., Schulam, P., Beam, A.L., Ranganath, R.: Opportunities in machine learning for healthcare. arXiv preprint arXiv:​1806.​00388 (2018) Ghassemi, M., Naumann, T., Schulam, P., Beam, A.L., Ranganath, R.: Opportunities in machine learning for healthcare. arXiv preprint arXiv:​1806.​00388 (2018)
13.
Zurück zum Zitat Goodfellow, I., et al.: Generative adversarial nets. In: Advances in Neural Information Processing Systems, pp. 2672–2680 (2014) Goodfellow, I., et al.: Generative adversarial nets. In: Advances in Neural Information Processing Systems, pp. 2672–2680 (2014)
15.
Zurück zum Zitat Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural Comput. 9(8), 1735–1780 (1997) CrossRef Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural Comput. 9(8), 1735–1780 (1997) CrossRef
17.
Zurück zum Zitat Kaplan, W., Wirtz, V., Mantel, A., Béatrice, P.: Priority medicines for Europe and the world update 2013 report. Methodology 2(7), 99–102 (2013) Kaplan, W., Wirtz, V., Mantel, A., Béatrice, P.: Priority medicines for Europe and the world update 2013 report. Methodology 2(7), 99–102 (2013)
19.
Zurück zum Zitat Li, W., Wang, Y., Cai, Y., Arnold, C., Zhao, E., Yuan, Y.: Semi-supervised rare disease detection using generative adversarial network. arXiv preprint arXiv:​1812.​00547 (2018) Li, W., Wang, Y., Cai, Y., Arnold, C., Zhao, E., Yuan, Y.: Semi-supervised rare disease detection using generative adversarial network. arXiv preprint arXiv:​1812.​00547 (2018)
20.
Zurück zum Zitat van der Maaten, L., Hinton, G.: Visualizing data using t-SNE. J. Mach. Learn. Res. 9(Nov), 2579–2605 (2008) MATH van der Maaten, L., Hinton, G.: Visualizing data using t-SNE. J. Mach. Learn. Res. 9(Nov), 2579–2605 (2008) MATH
21.
Zurück zum Zitat Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed representations of words and phrases and their compositionality. In: Advances in Neural Information Processing Systems, pp. 3111–3119 (2013) Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed representations of words and phrases and their compositionality. In: Advances in Neural Information Processing Systems, pp. 3111–3119 (2013)
22.
Zurück zum Zitat Miotto, R., Li, L., Kidd, B.A., Dudley, J.T.: Deep patient: an unsupervised representation to predict the future of patients from the electronic health records. Sci. Rep. 6, 26094 (2016) CrossRef Miotto, R., Li, L., Kidd, B.A., Dudley, J.T.: Deep patient: an unsupervised representation to predict the future of patients from the electronic health records. Sci. Rep. 6, 26094 (2016) CrossRef
23.
Zurück zum Zitat Obermeyer, Z., Emanuel, E.J.: Predicting the future–big data, machine learning, and clinical medicine. New Engl. J. Med. 375(13), 1216 (2016) CrossRef Obermeyer, Z., Emanuel, E.J.: Predicting the future–big data, machine learning, and clinical medicine. New Engl. J. Med. 375(13), 1216 (2016) CrossRef
24.
Zurück zum Zitat Pennington, J., Socher, R., Manning, C.: Glove: global vectors for word representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532–1543 (2014) Pennington, J., Socher, R., Manning, C.: Glove: global vectors for word representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532–1543 (2014)
25.
Zurück zum Zitat Purves, R.D.: Optimum numerical integration methods for estimation of area-under-the-curve (AUC) and area-under-the-moment-curve (AUMC). J. Pharmacokinet. Biopharm. 20(3), 211–226 (1992) CrossRef Purves, R.D.: Optimum numerical integration methods for estimation of area-under-the-curve (AUC) and area-under-the-moment-curve (AUMC). J. Pharmacokinet. Biopharm. 20(3), 211–226 (1992) CrossRef
26.
Zurück zum Zitat Rajkomar, A., et al.: Scalable and accurate deep learning with electronic health records. NPJ Dig. Med. 1(1), 18 (2018) CrossRef Rajkomar, A., et al.: Scalable and accurate deep learning with electronic health records. NPJ Dig. Med. 1(1), 18 (2018) CrossRef
27.
Zurück zum Zitat Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X.: Improved techniques for training GANs. In: Advances in Neural Information Processing Systems, pp. 2234–2242 (2016) Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., Chen, X.: Improved techniques for training GANs. In: Advances in Neural Information Processing Systems, pp. 2234–2242 (2016)
28.
Zurück zum Zitat Salimans, T., Kingma, D.P.: Weight normalization: a simple reparameterization to accelerate training of deep neural networks. In: Advances in Neural Information Processing Systems, pp. 901–909 (2016) Salimans, T., Kingma, D.P.: Weight normalization: a simple reparameterization to accelerate training of deep neural networks. In: Advances in Neural Information Processing Systems, pp. 901–909 (2016)
29.
Zurück zum Zitat Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: Dropout: a simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 15(1), 1929–1958 (2014) MathSciNetMATH Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: Dropout: a simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 15(1), 1929–1958 (2014) MathSciNetMATH
30.
Zurück zum Zitat Xiao, C., Choi, E., Sun, J.: Opportunities and challenges in developing deep learning models using electronic health records data: a systematic review. J. Am. Med. Inform. Assoc. 25(10), 1419–1428 (2018) CrossRef Xiao, C., Choi, E., Sun, J.: Opportunities and challenges in developing deep learning models using electronic health records data: a systematic review. J. Am. Med. Inform. Assoc. 25(10), 1419–1428 (2018) CrossRef
Metadaten
Titel
Modelling Patient Sequences for Rare Disease Detection with Semi-supervised Generative Adversarial Nets
verfasst von
Kezi Yu
Yunlong Wang
Yong Cai
Copyright-Jahr
2020
DOI
https://doi.org/10.1007/978-3-030-39098-3_11

Premium Partner