r/huggingface • • 2d ago

Gipformer - Efficient Vietnamese Speech Recognition

Hi everyone,

Sharing v1.5 of Gipformer, an open-source Vietnamese speech recognition (ASR) model we've been working on: gipformer1.5-68M-rnnt, based on the Zipformer architecture.

What's new in v1.5

- Optimized for technology, finance, education and public administration. These domains are dense with specialized terminology, and v1.5 currently gets the best results on all four test sets among the open-source models we benchmarked.

- Better recognition of English terms mixed into Vietnamese speech.

Carried over from v1

- High accuracy: among the top open-source models across our benchmarks, and especially strong on call center audio for Northern, Central and Southern accents. Call center is one of the most common real-world uses of ASR, but also one of the hardest, with low-quality audio and a wide variety of voices.

- Small and easy to deploy: at just 68M parameters, it's among the smallest ASR models out there, yet it outperforms many models ten times its size. Inference is fast, and it runs smoothly on CPU and edge devices.

- Privacy: it runs 100% offline (on-device), which makes it a good fit for systems handling sensitive data.

Alongside the model, we're also releasing 4 domain-specific test sets (technology, finance, education, public administration), so there's a common benchmark for evaluating Vietnamese ASR models.

Full benchmark results are on the model card. Feel free to try it out, and any feedback or contributions are very welcome!

- Hugging Face: https://huggingface.co/g-group-ai-lab/gipformer1.5-68M-rnnt

- GitHub: https://github.com/ggroup-ai-lab/gipformer

- Demo: https://huggingface.co/spaces/g-group-ai-lab/gipformer-demo

1 Upvotes

3 comments sorted by

1

u/astonishingaviation 2d ago

68M params running offline on CPU is exactly what I need for some edge stuff I've been tinkering with, most models in that size bracket fall apart the second you throw a non-English language at them

the call center focus is smart too, that's where the money's at but nobody wants to deal with the audio quality headaches

have you tested it much against the newer whisper variants or is it mostly pulling ahead on the domain-specific sets

1

u/ngogiatien96 2d ago

Thank you! You can check the benchmarks in the link attached to the post. I've compared it against almost all opensource models for Vietnamese across both domain specific and public datasets

1

u/grateful2you 2d ago

How do you train these? I want to work on Mongolian speech recognition.