Skip to main content
Erschienen in:
Buchtitelbild

2003 | OriginalPaper | Buchkapitel

Combating the Sparse Data Problem of Language Modelling

verfasst von : Frederick Jelinek

Erschienen in: Text, Speech and Dialogue

Verlag: Springer Berlin Heidelberg

Aktivieren Sie unsere intelligente Suche, um passende Fachinhalte oder Patente zu finden.

search-config
loading …

The talk will concern several ideas that combat the sparse data problem of language modeling. All alleviate it, neither solves it. These ideas are: equivalence classification of histories, positional clustering (different cluster systems for different n-gram positions), use of linguistic classes (e.g., Wordnet), class constraints in maximum entropy estimation, random forests, and neural network classification. An interesting problem that must be faced is as follows: words that are sparse and need to be classified do not have sufficient statistics to indicate their appropriate class membership.

Metadaten
Titel
Combating the Sparse Data Problem of Language Modelling
verfasst von
Frederick Jelinek
Copyright-Jahr
2003
Verlag
Springer Berlin Heidelberg
DOI
https://doi.org/10.1007/978-3-540-39398-6_1

Premium Partner