0:00
/

Ignite AI: The Future of Voice AI Testing and Self-Improving Agents with Tarush Agarwal | Ep284

Episode 284 of the Ignite Podcast

A voice agent can nail a demo and still fall apart on a real call. A caller interrupts mid-sentence. Someone mutters three digits of a phone number and trails off. Background noise garbles a word, the transcription gets it wrong, and the model confidently answers a question nobody asked.

That gap between “it worked when we tested it” and “it works when a stranger calls in angry” is what Tarush Agarwal has spent the last two years trying to close. He’s the co-founder and CEO of Cekura.ai, a platform that simulates, tests, and evaluates voice agents for more than 200 companies, running millions of conversations through systems before they ever touch a real customer.

Before Cekura, Tarush was optimizing quant trading systems in London and Chicago down to seven or nine nanoseconds of latency. Cekura itself started as an internal tool: he and his co-founders were building a voice agent for personal injury law firms, and instead of shipping the product, they ended up spending three hours every night after dinner personally calling their own agent looking for failures. They automated that process, showed it to a few friends building similar things, and the reaction told them what the real business was. They pivoted during the first week of Y Combinator.

In this episode, Tarush gets specific about why testing voice agents through text alone is a mistake, why roughly half of Cekura’s customers still run GPT-4.1 instead of newer models, and a piece of research that should worry anyone shipping a customer-support agent: a system-prompt attack that succeeds in one turn about 20 percent of the time succeeds across a multi-turn conversation roughly 92 percent of the time. In one test, his team talked a major provider’s support agent into handing out a $150 discount code just by repeating a false claim over several turns.

He also lays out what “good” actually looks like across different domains (a healthcare receptionist and a customer-support bot are not measured the same way), why the newest model is often the wrong choice for a live phone call, and where he thinks this goes once voice is solved: chat, avatars, and eventually physical AI.

Read about the full conversation here

Discussion about this video

User's avatar

Ready for more?