A multi-turn prompt attack convinced one of the biggest providers’ voice agents to issue a $150 discount code.
Tarush Agarwal is the co-founder and CEO of Cekura.ai, which builds testing and verification infrastructure for voice agents. Before founding Cekura, he studied computer science at IIT Bombay and worked on low-latency quantitative trading systems in London and Chicago, where teams optimized performance at roughly seven to nine nanoseconds. Cekura entered Y Combinator after pivoting from a legal voice-agent product, later raised a $2.5 million seed round, and now works with more than 200 customers while running millions of simulations.
Voice agents can perform well in controlled tests and still fail during real conversations. Interruptions, background noise, mixed languages, transcription errors, emotional callers, and multi-turn manipulation can expose problems that never appear in a text-based evaluation.
Tarush’s most counterintuitive claim is that better models do not automatically produce better voice agents. Around half of Cekura’s customers still use GPT-4.1 because newer reasoning-heavy models can introduce delays that do not work in live calls. Production performance depends on the full system, including latency, transcription, turn detection, interruption handling, speech quality, instruction following, and the infrastructure connecting each component.
In Today’s Episode We Discuss
- 00:01Introducing Tarush Agarwal and Cekura.ai
- 00:27From IIT Bombay to quantitative trading
- 02:45Founder life versus low-latency engineering
- 04:26Building voice agents for personal injury law firms
- 06:27Pivoting during the first week of Y Combinator
- 07:11Early growth, the $2.5 million seed round, and customer focus
- 09:49The current state of voice AI
- 12:49The metrics that determine voice-agent quality
- 15:17Compliance, healthcare, and high-stakes conversations
- 18:02How multi-turn prompt attacks exploit voice agents
- 19:17The quiet problem with how companies run evals
- 22:33Why testing voice agents through text is insufficient
- 24:06Cascading systems versus speech-to-speech models
- 25:36Building realistic simulation environments
- 27:12What changed in voice AI over two years
- 29:29Public benchmarks, latency gains, and accuracy limits
- 31:14Cekura’s long-term vision beyond voice
- 32:16Moving from founder-led sales to a dedicated GTM team
- 33:57The product metric Tarush watches every day
- 35:37Why voice AI could become larger than software
Cekura began after Tarush and his co-founders spent three hours after dinner manually calling their own legal voice agent. He explains why healthcare teams must simulate distressed patients, how multi-turn testing exposed the $150 discount exploit, and why his team sometimes shipped a bug fix before the customer reporting it had finished the call.
The episode returns to an old engineering principle: reliability begins when reality is allowed to break the system.
Pull Quotes
“Everyone talks about evals. I don’t think most people know how to do it correctly.”
“You need to build your own evals. You need to own your evals.”
Follow Tarush Agarwal on LinkedIn: LinkedIn
Follow Tarush Agarwal on X: X

