Using Multiple Time Scales in the Framework of Multi-Stream Speech Recognition

Hagen, Astrid; Bourlard, Hervé

doi:10.21437/ICSLP.2000-87

Hagen, Astrid; Bourlard, Hervé

2000

Download

Formats

Format
BibTeX
MARC
MARCXML
DublinCore
EndNote
NLM
RefWorks
RIS

Files

Abstract

In this paper, we present a new approach to incorporating multiple time scale information as independent streams in multi-stream processing. To illustrate the procedure, we take two different sets of multiple time scale features. In the first system, these are features extracted over variable sized windows of three and five times the original window size. In the second system, we take as separate input streams the commonly used difference features, i.e. the first and second order derivatives of the instantaneous features. In the same way, any other kinds of multiple time scale features could be employed. The approach is embedded in the recently introduced ``full combination'' approach to multi-stream processing in which, the phoneme probabilities from all possible combinations of streams are combined in a weighted sum. As an extension of this approach we have found that replacing the sum of probabilities by their product, in the same ``all wise'' context, can result in higher robustness. Capturing different information in each stream, and with the longer time scale features being more robust to noise, the multiple time scale multi-stream system gained a significant performance improvement in both clean speech and in real-environmental noise.

Details

Title Using Multiple Time Scales in the Framework of Multi-Stream Speech Recognition

Author(s) Hagen, Astrid ; Bourlard, Hervé

Published in 6th International Conference on Spoken Language Processing (ICSLP 2000)

Volume 1

Pages 349-352

Conference ICSLP

Date 2000

Keywords

multiple time scales; difference features; multi-stream; hagen; full combination; morris; bourlard; speech; HMM/ANN-Hybrid

Note IDIAP-RR 00-22

DOI https://doi.org/10.21437/ICSLP.2000-87

Additional link URL; Related documents

Laboratories LIDIAP

Record Appears in Scientific production and competences > STI - School of Engineering > IEM - Institut d'Electricité et de Microtechnique > LIDIAP - L'IDIAP Laboratory
Scientific production and competences > Euler Center for Signal Processing
Conference Papers
Work produced at EPFL
Published

Record creation date 2006-03-10

Actions

Preview

Select file: