Agentic VOICE AI developer (Senior and Junior Positions)

MuralWorksAI Private Limited , Pune · muralworksai.com · Full-time employment · Programming

Build the voice layer of the AI workforce.

MuralWorksAI is hiring a Voice AI Developer who wants to build systems that actually talk to people, understand them, reason in real time, and get things done.

This is not a “connect an LLM to Twilio” role.

You’ll be building production-grade conversational voice agents—the kind that need to handle interruptions, accents, silence, background noise, ambiguous requests, changing context, tool calls, failures, and real humans behaving like humans.

If you’re obsessed with making AI conversations feel fast, intelligent, natural, and reliable, this is your playground.

What you’ll own

You’ll work across the full Voice AI stack:

  • Speech-to-Text (ASR) — streaming transcription, accuracy, latency, punctuation, accents and noisy environments
  • LLMs & agent intelligence — prompting, structured outputs, tool calling, memory, reasoning and agent orchestration
  • Text-to-Speech (TTS) — natural speech, prosody, latency, voice selection and conversational expressiveness
  • Telephony — SIP, WebRTC, telephony APIs, call routing, recordings, transfers and real-time media streams
  • Real-time systems — streaming audio, WebSockets, event-driven architectures and low-latency inference
  • Agent tools — APIs, databases, CRMs, calendars, internal systems and external actions
  • Conversation state — context management, interruptions, retries, fallbacks and human handoffs
  • Evaluation & observability — call quality, latency, task completion, hallucination rates, containment, failure modes and cost
  • Multilingual / code-switched conversations — especially the messy, beautiful way real people actually speak

You won't just build demos.

You'll take voice agents from prototype → production → scale.

What you’ll do

Build ridiculously good voice agents

Design and ship conversational agents that can:

  • Understand a caller in real time
  • Respond naturally without awkward delays
  • Interrupt and be interrupted gracefully
  • Maintain context across long conversations
  • Call tools and APIs while talking
  • Recover when things go wrong
  • Know when they don't know
  • Escalate to humans at the right moment
  • Complete actual business workflows

Push latency down

In voice AI, 500ms matters.

You’ll profile the entire conversation loop—from caller speech to ASR → orchestration → LLM → TTS → audio playback—and continuously hunt for unnecessary latency.

Make conversations feel human

Not “human-like” because the marketing page says so.

Actually human.

You'll work on:

  • Turn-taking
  • Barge-in detection
  • Silence handling
  • Filler words
  • Prosody
  • Speaking speed
  • Sentence interruption
  • Confirmation strategies
  • Contextual responses
  • Error recovery
  • Conversation flow

Build the evaluation engine

A voice agent isn't finished when it works once.

You'll help create automated and human evaluation systems that answer:

Did the agent actually accomplish what it was supposed to accomplish?

You'll build test suites, simulations, regression tests, production monitoring and quality metrics around real conversations.

Experiment relentlessly

  • Try new models.
  • Break them.
  • Benchmark them.
  • Replace them.
  • Tune them.
  • Ship the better version.

We want someone who treats the Voice AI stack as an evolving engineering system—not a collection of APIs.

What we're looking for

Strong software engineering fundamentals.

You should be comfortable building production systems in at least one of:

  • Python
  • TypeScript / JavaScript
  • Go

And understand:

  • APIs
  • Async programming
  • WebSockets
  • REST
  • Databases
  • Queues / event-driven systems
  • Cloud infrastructure
  • Git
  • Testing and observability

And you should understand Voice AI beyond the buzzwords.

You don't need to be an academic speech researcher, but you should understand the fundamentals of:

ASR → LLM / Agent → TTS → Telephony → Real-time streaming

You should be comfortable debugging problems across that entire pipeline.

Strongly preferred

Experience with one or more of:

  • OpenAI Realtime / speech models
  • ElevenLabs
  • Deepgram
  • Cartesia
  • Gemini / Google speech stack
  • Azure Speech
  • AWS Transcribe / Polly
  • Twilio
  • Telnyx
  • Plivo
  • SIP / SIP trunks
  • WebRTC
  • LiveKit
  • Pipecat
  • LangChain / LangGraph
  • Temporal
  • Redis
  • PostgreSQL
  • Kafka / RabbitMQ
  • Docker
  • Kubernetes
  • AWS / GCP / Azure

Bonus points if you've built:

  • Voice agents
  • Call-center automation
  • IVR systems
  • Conversational AI
  • Speech recognition systems
  • Real-time audio systems
  • Contact-center integrations
  • Multilingual conversational systems
  • Agentic workflows
  • LLM evaluation infrastructure

The kind of person we're looking for

You might be a great fit if:

  • You hear a 1.5-second delay and immediately want to know why.
  • You don't trust a demo until you've tried to break it.
  • You care about the 1% of conversations where everything goes wrong.
  • You'd rather ship a working system than spend three weeks debating architecture diagrams.
  • You understand that production AI is 30% model and 70% engineering, evaluation, edge cases and relentless iteration.
  • You can look at a failed conversation and figure out whether the problem was ASR, prompt design, context management, tool execution, latency, TTS, telephony—or the underlying product logic.

Most importantly:

You want to build something people actually use

What you'll get to work on

This role has unusually high ownership.

You'll have the opportunity to work on:

Real-time AI
Build systems where milliseconds matter.

Agentic AI
Give voice agents the ability to reason, use tools and take actions.

Speech technology
Work at the intersection of audio, language models and human conversation.

Production infrastructure
Design systems that need to survive real call volume—not just a hackathon demo.

AI evaluation
Figure out how to objectively measure whether an AI conversation is actually good.

Next-generation human-computer interaction
Help define what it feels like to interact with software by talking to it.

This IS the role if…

You want to wake up and think:

“How do we make an AI have a genuinely great conversation with a human?”

…and then spend the day actually solving it.

We care more about what you've built than the number of years you've been doing it.

How to apply

Send us:

  1. Your GitHub / portfolio
  2. 1–3 things you've built that you're genuinely proud of 

If you've built a voice agent, give us a number we can call. We'd rather talk to something you've built than read 500 words about how passionate you are about AI.

We're not looking for someone to maintain yesterday's AI.
We're looking for someone who wants to build tomorrow's voice interface.

Apply for this position

Login with Google or GitHub to see instructions on how to apply. Your identity will not be revealed to the employer.

It is NOT OK for recruiters, HR consultants, and other intermediaries to contact this employer