I will build a voice ai agent using elevenlabs, stt, and llm in python


About this gig
Deploy an intelligent Voice AI agent that can listen, reason, and speak in real time.
I build custom conversational Voice AI architectures in Python, combining low-latency audio streaming with LLMs and speech synthesis.
What I deliver:
- End-to-end Voice Pipelines: Voice Activity Detection (VAD) + Whisper (STT) + LLM Reasoning + ElevenLabs (TTS)
- Ultra-low Latency Audio Routing via asynchronous Linux pipes
- Telegram Voice Calls integration via Pyrogram
- Cloud GPU Deployment (RunPod container workers, parallel queue processing, dynamic scaling)
- Modular Python backend with isolated audio capture and synthesis
Stack: Python, asyncio, ElevenLabs, Whisper, RunPod GPU, Pyrogram.
Please contact me before ordering to review API credentials and voice latency requirements.
Get to know Moddie
Fullstack web developer
- FromUkraine
- Member sinceOct 2024
- Avg. response time1 hour
Languages
English, Ukrainian, Russian
My Portfolio
FAQ
How do you achieve low latency in speech synthesis?
I stream audio chunks via asynchronous Linux pipes directly into the player without saving temporary audio files to disk, reducing latency below 600ms
Can this voice agent make phone or Telegram calls?
Yes, I can integrate the agent with Telegram voice calls via Pyrogram or external telephony APIs (Twilio/LiveKit)

