How The Open ASR Leaderboard Is Embracing Its First Global South Language
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How The Open ASR Leaderboard Is Embracing Its First Global South Language on ThorstenMeyerAI.com

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

TL;DR

The Open ASR Leaderboard has introduced Hindi and Indian English evaluation sets, making Hindi the first Indic and Global South language on the platform. This development aims to improve model assessment across diverse populations and address biases in speech recognition.

The Open ASR Leaderboard on Hugging Face has added two new evaluation sets — Monsoon en-IN for Indian English and Monsoon hi-IN for Hindi — making Hindi the first Indic and Global South language included in the platform’s speech recognition benchmarking’s multilingual tab. This marks a significant milestone in broadening the scope of speech recognition benchmarking beyond European languages, with implications for model development and fairness.

The new sets are designed with diversity in mind: they include speech from 4,888 speakers across India, recorded in natural, unscripted conversations using their own devices. The datasets feature extensive metadata, including age, gender, occupation, income, device type, and geographic location, enabling detailed analysis of model performance across different populations. For more on the importance of diverse datasets, see the original analysis on speech recognition evaluation.

Each language set is split into public and private segments to facilitate unbiased benchmarking. The Indian English sets contain approximately 11.2 hours of audio, while the Hindi sets include around 5 hours. The audio is drawn from a variety of environments and speaker backgrounds, with the aim of capturing real-world speech variability. Notably, the Hindi transcripts include a lattice of accepted spellings due to orthographic variations, a novel approach compared to the normalisation used for English.

The announcement emphasizes that traditional single-number error metrics, like Word Error Rate (WER), can mask disparities in model performance across different demographics. This development is discussed in detail in the original analysis. By incorporating detailed speaker attributes, the leaderboard aims to promote the development of more equitable ASR systems. The new datasets are now available for self-scoring, with the full private test sets to be released later.

At a glance
reportWhen: announced March 2024
The developmentThe Open ASR Leaderboard has expanded to include Hindi and Indian English evaluation sets, reflecting a significant step toward inclusivity and comprehensive benchmarking for multilingual speech recognition.
At a glance
announcementWhen: announced now; sets released publicly w…
The developmentVoice Arena and Hugging Face have added Hindi and Indian English evaluation sets — Monsoon hi-IN and Monsoon en-IN — to the Open ASR Leaderboard, making Hindi the first Global South language it covers.

Inclusion of Hindi and Indian English Transforms Benchmarking

This expansion represents a major step toward more inclusive and representative speech recognition benchmarks. Hindi, spoken by over half a billion people, has historically been underrepresented in global speech datasets and evaluation benchmarks. Its inclusion on the Hugging Face platform signals a shift toward recognizing the linguistic diversity of the Global South and addressing biases in AI models.

Furthermore, the detailed metadata allows researchers and developers to analyze model performance disparities related to geography, age, gender, and socioeconomic status. This could lead to more targeted improvements, reducing disparities that have been documented in prior research, such as racial and gender biases in commercial ASR systems.

Overall, this move could influence the direction of future ASR research, encouraging the creation of models that perform well across diverse accents, dialects, and socioeconomic backgrounds, and fostering greater trust in speech technology for underserved populations.

Amazon

speech recognition microphone for Indian languages

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Previous Limitations in ASR Benchmarking for Global South Languages

Until now, the Open ASR Leaderboard primarily focused on European languages, with benchmarks largely based on scripted or controlled speech datasets. While some datasets included diverse accents within those languages, there was little emphasis on languages from the Global South, such as Hindi or other Indic languages.

The lack of representation has contributed to performance gaps and biases in commercial speech recognition systems, which often perform poorly for speakers with accents or dialects not well represented in training data. Prior research, including studies on racial and gender disparities, has highlighted these issues but lacked comprehensive benchmarking tools for non-Western languages.

The introduction of Hindi and Indian English datasets addresses these gaps by providing real-world, diverse speech data from a broad demographic, enabling more accurate evaluation and development of inclusive models.

“Adding Hindi and Indian English to the Open ASR Leaderboard marks a significant step toward inclusive benchmarking, reflecting the linguistic diversity of the Global South.”

— Thorsten Meyer, Hugging Face

Amazon

AI speech recognition software for Hindi

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Dataset Impact and Model Performance

It remains unclear how current state-of-the-art models perform on the new Hindi and Indian English datasets, as baseline results have not yet been published. The stability of rankings given the relatively small size of the Hindi set (around 5 hours) is also uncertain, raising questions about the robustness of benchmarking in this context.

Additionally, the effectiveness of the lattice approach for Hindi, which accounts for orthographic variation, has not been directly compared to traditional normalisation methods. Whether the new datasets will influence model improvements across all demographic attributes remains to be seen.

Finally, it is not yet known whether leaderboard participants will disaggregate their results by the detailed speaker attributes now available, which could provide deeper insights into model fairness and bias.

Amazon

voice recognition devices for multilingual use

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Steps for Benchmarking and Model Development

The datasets are now accessible for public self-scoring, and the private test sets are expected to be released later in 2024. Researchers and developers will likely evaluate existing models on these new datasets, providing initial baseline results.

Further work is anticipated to explore the impact of detailed speaker attributes on model performance, as well as testing the effectiveness of the Hindi lattice approach in reducing error rates related to orthographic variation.

Long-term, the inclusion of Hindi and Indian English may inspire the development of more inclusive ASR systems tailored to diverse linguistic and socio-economic contexts, ultimately fostering broader adoption and trust in speech technology across the Global South.

Amazon

speech-to-text app for Indian English

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is the inclusion of Hindi on the Open ASR Leaderboard significant?

It marks the first time a language from the Global South, spoken by over half a billion people, is included, addressing a major gap in benchmarking and encouraging more equitable ASR development.

What makes the datasets for Hindi and Indian English different from previous benchmarks?

They feature diverse, unscripted speech from speakers across India, with extensive metadata and real-world recording conditions, capturing linguistic and demographic variability.

How might this development affect future speech recognition models?

It could lead to models that perform better across different populations, reducing biases related to accent, region, and socio-economic status, and promote more inclusive AI systems.

Are baseline performance results available for these new datasets?

No, the announcement did not include baseline results, so current model performance on Hindi and Indian English remains to be evaluated.

Will the detailed speaker attributes influence how models are evaluated?

It is expected that researchers will analyze performance across different demographic groups, which could reveal disparities and guide targeted improvements.

Primary source: Hugging Face · via ThorstenMeyerAI.com

GRILLING SEASON

Grilling season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is building a cloud business aimed at selling surplus AI computing capacity, expanding beyond its social media roots. Details are still emerging.

OpenAI Cuts Off Cursor: The Developers Are The Collateral

OpenAI plans to cut off its models from Cursor by November 12 due to ownership transfer to SpaceX, impacting developers relying on the tool.

Ox Alpha

OpenRouter announces Ox Alpha, its new AI model, marking a significant step in open-source AI development. Details are still emerging.

The Menu: What Ten Answers Reveal

An analysis of ten jurisdictions’ approaches to automation, income, and AI, revealing patterns and challenges in managing the post-labor transition.