A review of social background profiling of speakers from speech accents

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Social background profiling of speakers is heavily used in areas, such as, speech forensics, and tuning speech recognition for accuracy improvement. This article provides a survey of recent research in speaker background profiling in terms of accent classification and analyses the datasets, speech features, and classification models used for the classification tasks. The aim is to provide a comprehensive overview of recent research related to speaker background profiling and to present a comparative analysis of the achieved performance measures. Comprehensive descriptions of the datasets, speech features, and classification models used in recent research for accent classification have been presented, with a comparative analysis made on the performance measures of the different methods. This analysis provides insights into the strengths and weaknesses of the different methods for accent classification. Subsequently, research gaps have been identified, which serve as a useful resource for researchers looking to advance the field.

Related collections

Most cited references 67

Record: found
Abstract: found
Article: not found

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, Niki Parmar … (2017)

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data. 15 pages, 5 figures

0 comments Cited 2123 times – based on 0 reviews      Review now

Bookmark

Record: found
Abstract: found
Article: not found

Generative adversarial networks

Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza … (2020)

Generative adversarial networks are a kind of artificial intelligence algorithm designed to solve the generative modeling problem. The goal of a generative model is to study a collection of training examples and learn the probability distribution that generated them. Generative Adversarial Networks (GANs) are then able to generate more examples from the estimated probability distribution. Generative models based on deep learning are common, but GANs are among the most successful generative models (especially in terms of their ability to generate realistic high-resolution images). GANs have been successfully applied to a wide variety of tasks (mostly in research settings) but continue to present unique challenges and research opportunities because they are based on game theory while most other approaches to generative modeling are based on optimization.

0 comments Cited 629 times – based on 0 reviews      Review now

Bookmark

Record: found
Abstract: not found
Article: not found

Front-End Factor Analysis for Speaker Verification

Najim Dehak, Patrick J. Kenny, Réda Dehak … (2011)

0 comments Cited 322 times – based on 0 reviews      Review now

Bookmark

All references

Author and article information

Contributors

Mohammad Ali Humayun

Journal

Journal ID (nlm-ta): PeerJ Comput Sci

Journal ID (iso-abbrev): PeerJ Comput Sci

Journal ID (publisher-id): peerj-cs

Title: PeerJ Computer Science

Publisher: PeerJ Inc. (San Diego, USA )

ISSN (Electronic): 2376-5992

Publication date (Electronic): 16 April 2024

Publication date Collection: 2024

Volume: 10

Electronic Location Identifier: e1984

Affiliations

[1 ]Department of Computer Science, Information Technology University , Lahore, Pakistan

[2 ]Department of Computer and Information Sciences, Universiti Teknologi PETRONAS , Seri Iskandar, Malaysia

[3 ]Faculty of Integrated Technologies, Universiti Brunei Darussalam , Jalan Tungku Link, Brunei

Article

Publisher ID: cs-1984

DOI: 10.7717/peerj-cs.1984

PMC ID: 11042007

PubMed ID: 38660189

SO-VID: 507added-996b-411b-94c4-83e33882b529

License:

This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, reproduction and adaptation in any medium and for any purpose provided that it is properly attributed. For attribution, the original author(s), title, publication source (PeerJ Computer Science) and either DOI or URL of the article must be cited.

History

Date received : 29 January 2024

Date accepted : 18 March 2024

Funding

The authors received no funding for this work.

A review of social background profiling of speakers from speech accents

Read this article at

Abstract

Related collections

Special Issue: Social pedagogical work with children, youth and their families with refugee and migrant background in Europe

Most cited references 67

Attention Is All You Need

Generative adversarial networks

Front-End Factor Analysis for Speaker Verification

Author and article information

Contributors

Journal

Affiliations

Article

History

Funding

Categories

Comments

Comment on this article

Similar content 697

Most referenced authors 547