I have +7 year of experience working with deep learning applied to speech recognition:
- Speech to text,
- Diarization,
- Voice Activity Detection,
- Sound Event Detection,
- Denoising,
- Audio Signal Processing,
- Emotion
- Voice Agents...
in different languages.
I have been working with SOTA Automatic Speech Recognition APIs and frameworks: Whisper, Kaldi, Vosk, MMS, DeepSpeech, speechbrain and wav2vec2. I have been working to fine-tuned models to improve WER and speed inference on multiple language.
Hugging Face: https://huggingface.co/deepdml
Github: https://github.com/djpg... Read more