To Memorize or to Predict: Prominence Labeling in Conversational Speech

Nenkova, Ani; Brenier, Jason; Kothari, Anubha; Calhoun, Sasha; Whitton, Laura; Beaver, David; Jurafsky, Dan

To Memorize or to Predict: Prominence Labeling in Conversational Speech

Files

2007_to_memorize_or_to_predict__prominence_labeling_in_conversational_speech.pdf (85.04 KB)

Penn collection

Departmental Papers (CIS)

Subject

Computer Sciences

Permalink

https://repository.upenn.edu/handle/20.500.14332/6800

View all metadata

Author

Nenkova, Ani

Brenier, Jason

Kothari, Anubha

Calhoun, Sasha

Whitton, Laura

Beaver, David

Jurafsky, Dan

Abstract

The immense prosodic variation of natural conversational speech makes it challenging to predict which words are prosodically prominent in this genre. In this paper, we examine a new feature, accent ratio, which captures how likely it is that a word will be realized as prominent or not. We compare this feature with traditional accent-prediction features (based on part of speech and N-grams) as well as with several linguistically motivated and manually labeled information structure features, such as whether a word is given, new, or contrastive. Our results show that the linguistic features do not lead to significant improvements, while accent ratio alone can yield prediction performance almost as good as the combination of any other subset of features. Moreover, this feature is useful even across genres; an accent-ratio classifier trained only on conversational speech predicts prominence with high accuracy in broadcast news. Our results suggest that carefully chosen lexicalized features can outperform less fine-grained features.

Date of presentation

2007-04-01

Conference name

Departmental Papers (CIS)

Conference dates

2023-05-17T07:17:23.000

Comments

Nenkova, A., Brenier, J., Kothari, A., Calhoun, S., Whitton, L., Beaver, D., & Jurafsky, D., To Memorize or to Predict: Prominence Labeling in Conversational Speech, Human Language Technology Conference of the North American Chapter of the association of Computational Linguistics, April 2007. http://www.aclweb.org/anthology/N07-1002

Collection

Presentations