Related Experiment Video
Updated: Jan 18, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
LoRA-INT8 Whisper: A Low-Cost Cantonese Speech Recognition Framework for Edge Devices
Lusheng Zhang1,2, Shie Wu1,2, Zhongxun Wang1,2
1School of Physics and Electronic Information, Yantai University, Yantai 264005, China.
This study introduces an efficient Cantonese automatic speech recognition (ASR) system using LoRA fine-tuning and INT8 quantization. The method significantly reduces model size and improves inference speed for low-resource ASR applications.
Area of Science:
- Speech Recognition
- Artificial Intelligence
- Machine Learning
Background:
- Low-resource automatic speech recognition (ASR) for Cantonese faces challenges with data scarcity, large models, and slow inference.
- Edge deployment of ASR systems requires efficient models that balance accuracy, speed, and size.
Purpose of the Study:
- To develop a cost-effective Cantonese ASR system for low-resource and edge-deployment settings.
- To address bottlenecks in data scarcity, model size, and inference speed.
Main Methods:
- Parameter-efficient fine-tuning of Whisper-tiny using LoRA (Low-Rank Adaptation) with rank = 8 on the Common Voice zh-HK dataset.
- INT8 quantization of the fine-tuned model using dynamic quantization in ONNX Runtime to create a compact checkpoint.
- Evaluation of the quantized model's performance on CPU and GPU, measuring character error rate (CER) and real-time factor (RTF).
Main Results:
- LoRA fine-tuning reduced CER from 49.5% to 11.1%, approaching full fine-tuning performance while significantly decreasing training costs.
- The INT8 quantized model achieved a 60 MB size, with RTF = 0.20 on CPU (5x real-time) and RTF = 0.06 on GPU, demonstrating substantial speed improvements.
- Ablation studies confirmed the LoRA-INT8 configuration provides an optimal trade-off between accuracy, speed, and model size.
Conclusions:
- The proposed LoRA-INT8 Cantonese ASR system offers a practical solution for resource-constrained environments.
- The system achieves significant improvements in efficiency and performance, paving the way for real-time ASR on edge devices.
- Future work aims to further enhance performance and energy efficiency for practical IoT applications.
More Related Videos
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Related Concept Videos
Components of Language
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Language and Cognition