Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct
      Speech Translation

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases they are often presented with continuous audio requiring automatic (and sub-optimal) segmentation. After comparing existing techniques (VAD-based, fixed-length and hybrid segmentation methods), in this paper we propose enhanced hybrid solutions to produce better results without sacrificing latency. Through experiments on different domains and language pairs, we show that our methods outperform all the other techniques, reducing by at least 30% the gap between the traditional VAD-based approach and optimal manual segmentation.

Related collections

Author and article information

Journal

Publication date Created: 23 April 2021

Article

ArXiV ID: 2104.11710

SO-VID: 9386f11b-1de6-4f22-a928-9bda9b112430

License:

http://creativecommons.org/licenses/by-nc-sa/4.0/

History

Custom metadata

Categories cs.SD cs.CL eess.AS

ScienceOpen disciplines: Theoretical computer science,Electrical engineering,Graphics & Multimedia design

Data availability:

ScienceOpen disciplines: Theoretical computer science, Electrical engineering, Graphics & Multimedia design

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

Read this article at

Abstract

Related collections

Fandom activity

Author and article information

Journal

Article

History

Custom metadata

Comments

Comment on this article

Similar content 153