Related Experiment Video
Updated: Mar 27, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Machine learning based sample extraction for automatic speech recognition using dialectal Assamese speech
Swapna Agarwalla1, Kandarpa Kumar Sarma1
1Department of Electronics and Communication Engineering, Gauhati University, Guwahati, 781014, Assam, India.
Machine learning techniques enhance Automatic Speaker Recognition (ASR) for Assamese speech by extracting relevant data from big data. Recurrent Neural Networks with composite features show superior performance in recognizing variations.
Area of Science:
- Speech Processing and Recognition
- Machine Learning Applications
- Human-Computer Interaction
Background:
- Automatic Speaker Recognition (ASR) is evolving with Human-Computer Interaction (HCI), incorporating big data and the Internet of Things (IoT).
- Learning-based techniques are gaining traction in ASR due to their ability to mimic biological behavior for improved modeling and processing.
- Current ASR methods are integrating big data and IoT concepts, necessitating advanced machine learning approaches.
Purpose of the Study:
- To investigate machine learning (ML) approaches for extracting relevant samples from big data for ASR.
- To apply soft computing techniques for ASR of Assamese speech, addressing dialectal variations.
- To evaluate the performance of various ML models, including Artificial Neural Networks (ANNs) and Deep Neural Networks (DNNs), for speaker recognition tasks.
Main Methods:
- Utilized ML techniques, including feedforward (FF) and deep neural networks (DNNs), for sample extraction and feature engineering.
- Employed Multi-Layer Perceptrons (MLPs) with raw speech, extracted features, and frequency domain inputs for class information learning.
- Applied Recurrent Neural Networks (RNNs) and Fully Focused Time Delay Neural Networks (FFTDNNs) with spectral and prosodic features for recognition.
Main Results:
- Proposed ML-based sentence extraction techniques and a composite feature set with RNN as a classifier outperformed other methods.
- ANNs in feedforward form were used as feature extractors, enabling performance evaluation and comparison.
- Experimental results demonstrated that leveraging big data samples significantly enhanced ASR system learning.
Conclusions:
- ANN-based sample and feature extraction techniques are efficient for integrating ML into big data aspects of ASR systems.
- The developed ML approaches are effective for recognizing speaker, dialect, gender, and mood variations in Assamese speech.
- The study highlights the potential of big data and advanced ML techniques for advancing ASR capabilities.
Related Concept Videos
Extraction: Advanced Methods
Sample Handling
Samples should be transported carefully from collection points to the laboratory. They should be properly sealed and clearly labeled to prevent cross-contamination. To preserve the sample integrity, optimal temperature conditions during transport are essential. This could involve using...
Sampling Methods: Sample Types
Solid samples include a variety of substances, such as sediments from water bodies, soil, metals, and biological tissues. Two standard methods for extracting sediments from water bodies are grab sampling and piston coring. Grab sampling involves using a device to collect a discrete sediment sample from the bottom of a water body with minimal disturbance. Grab samples do not always represent the entire area due to...
Sampling Methods: Overview
In analytical chemistry, the choice of...
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Sampling Continuous Time Signal
In the...

