NVIDIA Cuts Saudi Arabic Speech Errors Using SDAIA Dataset

2 Min Read

NVIDIA has used Saudi Arabia’s Saudi Audio Dataset for Arabic (SADA) to fine-tune its open-weight Nemotron 3.5 ASR speech recognition model for Najdi and Hijazi dialects.

According to NVIDIA, the model’s word error rate on Saudi Arabic fell from 55% to about 30% after training on SADA audio. Across the dataset’s full multi-dialect test set, word error rate dropped from 58.84% to 35.61%, while character-level error fell from 35.40% to 15.97%.

Developed by the Saudi Data and Artificial Intelligence Authority (SDAIA) with the Saudi Broadcasting Authority, SADA contains around 667 hours of transcribed audio spanning more than 10 Saudi dialects. It includes over 600 hours from 57 television programmes and more than 125,000 categorised audio clips, released publicly on Kaggle.

NVIDIA said the dialect improvements did not come at the expense of English or Modern Standard Arabic performance. The fine-tuning process took about 4.5 hours using two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.

The enhanced model supports around 40 languages and dialects, with latency starting at 80 milliseconds. Potential uses include voice assistants, conversational agents, live subtitling and speech transcription. NVIDIA has also published the workflow and tools for adapting the approach to other languages and dialects.

Source: Middle East AI News

Share This Article