Qureos

Find The RightJob.

Senior Voice AI Engineer

Senior Voice AI Engineer
Platform: Convogent
Level: Mid–Senior (5+ years | 2+ years building voice AI in production)
Location: Bangalore Office
About the Role
You will own the distributed, real-time systems that orchestrate parallel voice calls at massive scale: the core platform behind every conversation Convogent runs. The job is keeping thousands of simultaneous calls fast, reliable, and natural, owning how STT/ASR, LLM, TTS, RAG, and tool-calling work together, and where every millisecond and every point of quality goes across the pipeline.

What you’ll do
  • Own voice AI at scale. Keep the pipeline fast and reliable at enterprise concurrency, thousands of simultaneous calls, by lifting calls-per-vCPU through extending Pipecat or rewriting hot paths in Go.
  • Own end-to-end latency. Keep the conversation under ~800ms by decomposing the budget across STT, LLM, TTS, RAG, tool-calling, turn-detection, and network hops.
  • Diagnose the bad call. When a call does not feel right, locate the fault in the pipeline fast by asking the right questions, reading the right signals, and owning the fix.
  • Make quality measurable. Build voice-to-voice evals that gate releases and catch STT/LLM/TTS provider drift before customers do.
  • Own provider tradeoffs. Decide how STT / LLM / TTS / RAG / tool-calling choices balance voice quality against latency, per conversation flow.
  • Build adjacent products. Extend the platform into live assist, agent assist, and what comes next, on the same pipeline foundations.
  • Set the bar. Review, mentor, and define engineering standards for a growing team.

What we’re looking for
  • 5+ years building production products and systems, with at least 2 of those years in real-time voice, operating it in production, not just shipping a demo.
  • Deep understanding of the full voice pipeline, every stage end to end and how each one affects the next and shapes quality and latency.
  • Strong in Go/Python, comfortable across the whole pipeline. Production-grade, performance-critical Go is especially valued: you have rewritten hot-path components in Go to lift throughput under real load.
  • A track record of running voice AI at scale, scaling Pipecat (or a comparable real-time voice framework) to high concurrency in production.
  • Provider-tradeoff fluency: how STT / LLM / TTS / RAG / tool-calling choices play off each other on quality and latency.
  • An eval-driven approach to quality. You measure voice quality with real signals, not vibes, and have built or run evals on a real-time voice or LLM system.
  • Fluency with agentic coding tools and a habit of treating quality as a measured signal.

What we’re NOT looking for
  • A prompt engineer / "AI app" builder who has only called LLM APIs and never operated a real-time pipeline under load.
  • Someone who has only used managed platforms (Vapi/Retell/Bland) and never scaled the layer beneath them.
  • A pure web-backend CRUD engineer with no real-time/streaming/voice exposure.
  • An ML researcher who wants to train/fine-tune models. We integrate providers; we don't build models.

© 2026 Qureos. All rights reserved.