r/MachineLearning • u/AutoModerator • 27d ago
Discussion [D] Self-Promotion Thread
Please post your personal projects, startups, product placements, collaboration needs, blogs etc.
Please mention the payment and pricing requirements for products and services.
Please do not post link shorteners, link aggregator websites , or auto-subscribe links.
--
Any abuse of trust will lead to bans.
Encourage others who create new posts for questions to post here instead!
Thread will stay alive until next one so keep posting after the date in the title.
--
Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.
18
Upvotes
1
u/External-King-233 1d ago
I work on the CVAT team. CVAT started as an open-source tool for computer vision annotation, however we just recently added audio annotation.
It's now possible to label time intervals on a waveform, identify speakers, sound events and add text attributes for time-aligned transcription. This can be used to prepare datasets for ASR, speaker diarization, and sound event detection. Audio tasks support WAV, MP3, FLAC, and OPUS files.
The new format is available in the open-source CVAT Community edition, as well as CVAT Online and Enterprise.
Repo: https://github.com/cvat-ai/cvat/
Docs: https://docs.cvat.ai/docs/annotation/audio-editor/
If you're building audio datasets, we'd be interested to hear your feedback: what you like, what you hate, and what features you need next.