Shobhit Banga, Large Scale Human Evaluations – Rethinking Evaluations for Today’s Voice Models

Shobhit Banga, Large Scale Human Evaluations – Rethinking Evaluations for Today’s Voice Models

🎙 Shobhit Banga 👥 4K 📅 September 1, 2026 ⏱ 80 min 👁 0 📄 expert opinion 🧭 2026-09-01
Available in: English (current) Français

Keywords

human evaluationASR benchmarkTTS leaderboardvoice modelsdata collection

Summary

Shobhit Banga, founder of a non-profit that shares inspiring stories, argues that large-scale human evaluation is now feasible and essential for voice models. He presents three projects: Voice of India, an ASR benchmark covering 15 languages with 36,000 speakers, designed to be representative across geography, age, gender, vocabulary, devices, acoustic environments, speech type, speed, and multiple valid transcripts. He details the process of collecting spontaneous conversations via a mobile app, and the human-in-the-loop transcription validation that accounts for multiple valid spellings. Results show significant differences in model performance when considering these factors, revealing hidden weaknesses. The project took 90 days and cost $500,000. He then discusses a TTS leaderboard that evaluates models across languages and domains, using pairwise human comparisons. He emphasizes that human evaluation is no longer too slow or expensive, and encourages researchers to embrace it. The talk includes Q&A and practical advice for replication.

147 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical challenges and benefits of large-scale human evaluation for voice models. The speaker’s argument is compelling, supported by concrete examples from his projects, such as the impact of multiple valid transcripts on word error rates and the discovery of model-specific weaknesses. He effectively demonstrates that human evaluation can be done at scale within reasonable time and cost, challenging the common resistance to human-in-the-loop approaches. The argumentation is persuasive, though it relies heavily on anecdotal evidence and the speaker’s own experience rather than systematic comparison with alternative methods.

Scientific Rigor, Source Quality, Title Accuracy

The talk is based on the speaker’s own projects and experience, which lends authenticity but limits scientific rigor. No external sources are cited, and the methodology is described at a high level without full transparency. The title accurately reflects the content, which is a case for large-scale human evaluation. The speaker’s claims about cost and time are plausible but not independently verified. The talk would benefit from more detailed methodological disclosure and peer-reviewed validation.

183 words

Title / Content Match

The title accurately reflects the content, which focuses on the case for large-scale human evaluations in voice model assessment.

Quality & Reliability

7/10

The speaker is a practitioner with direct experience in large-scale human evaluation, but the talk is largely anecdotal and lacks peer-reviewed validation or detailed methodological transparency. The claims about cost and time are plausible but not independently verified.

Key Moments

Cited Sources

  • Voice of India benchmark — Mentioned as an ASR benchmark for Indian languages, presented at Interspeech 2026.
  • TTS leaderboard — Mentioned as a project for evaluating TTS models across languages and domains.

Concurring Sources

  • Interspeech 2026 — Conference where the Voice of India benchmark was presented.

Contribution & Novelties

The talk provides a practical blueprint for conducting large-scale human evaluations of voice models, demonstrating that it is feasible in terms of time and cost. It highlights the importance of considering multiple valid transcripts, which is often overlooked in ASR benchmarks. The speaker’s experience offers actionable insights for researchers and practitioners.

Pour aller plus loin :

93 words

Radar Profile

The radar profile shows high scores in quantity of information and quality of information, reflecting the talk's rich content and practical insights. The technical level is moderate, as the talk is accessible to a broad audience. The global reliability is moderate, due to the anecdotal nature of the evidence.

Reliability 6/10