Testing Voice AI Agents at Scale
Brooke Hopkins · Founder · Coval
Voice AI agents must do more than respond correctly to a prompt. Brooke Hopkins, founder of Coval, joins The Tech Trek to explain how teams simulate conversations, evaluate production performance, and build trust in voice interfaces. Drawing on her experience leading evaluation infrastructure at Waymo, she also explores how AI is reshaping engineering workflows, technical hiring, and decisions as a founder.
Brooke compares testing voice AI with autonomous-vehicle simulation. From her work leading evaluation infrastructure at Waymo, she describes how complex systems can be tested across many possible conditions before users encounter failures.
The conversation explores why speech recognition alone did not make earlier voice assistants truly useful, how reasoning models changed multi-turn interactions, and why earning trust depends on systems consistently carrying out real tasks.
Brooke also discusses the changing role of voice versus visual interfaces, coordinating multiple coding agents, evaluating engineers in the AI era, and choosing when automation should give way to careful human review.
As a solo founder, Brooke describes using different AI tools across engineering, sales, marketing, and management while relying on a trusted team for important decisions.
Full transcript of this conversation.
Amir: On this episode of the show, I have with me Brooke Hopkins. She is founder at Coval, and we're gonna be talking about real-time models to support voice and chat agents, and we're gonna talk about how those are different from other models and some of the challenges, and the opportunity, the impact of voice and chat within the ecosystem.
And also, a lot of people have had bad experiences with voice solutions in the past, and how the current solutions are overcoming that, and I'm sure a couple other things as well. But Brooke, thanks for taking the time.
Brooke: Yeah, thanks so much for having me on. Super excited.
Amir: Absolutely. Okay, before we start, what do you guys do over there?
Brooke: So we build a platform for voice agents so that you can scale them to tens of millions of conversations. So think like everything from testing it before production to running hundreds of simulations to be able to see where things are going wrong with simulated users. And then once it's in production, understanding what you missed before, so next generation QA, but also product insights of what are your customers asking for, where are things going wrong.
And then human in the loop labeling. So, um, actually grounding all of these evals and how do humans actually judge them so that you can scale that human judgment to y- all of your evals across the board. So we work with companies like Perplexity or Deepgram to be able to do all their evals in-house and make sure that their agents are working as you expect.
Um, and so my background's from Waymo. My... I led our evaluation infrastructure team at Waymo that was responsible for all of our developer tools, so everything for launching and running simulations, which is very similar to voice agents. So how do you make sure that a self-driving car can go from point A to point B and simulate all those possible environments?
Um, but also architecturally, self-driving cars are very similar, where you have, um, you know, perception, what's happening in the world around me, planning, what should I do next, and then, uh, control, so how should I actually take that action? And voice agents are the same. So you have transcription, pl- um, reasoning, so what should I say next, and then actually saying that thing out loud.
Um, so we've helped hundreds of teams to deploy their voice agents at scale and yeah, super, super excited to dive in.
Amir: Absolutely. So yeah, I guess maybe a starting point. Wh- when we're talking about, you know, supporting voice, uh, and chat AI, I mean, obviously we're, we... There's a need for real time. And, and talk to us a little bit about how that looks different than maybe other types of solutions out there.
Brooke: Yes. Uh, so you mean like real time versus cascaded model or real time versus, like, say, chat?
Amir: Uh, m- maybe both actually in this case.
Brooke: Yeah, so kind of there's a variety of different agentic use cases that we're seeing, uh, emerge. There was, you know, early in LLMs, you just had a prompt and then it would produce some result, and that was the majority of LLM use cases.
The majority of agents that we're seeing now in production are conversational. So think, um, you know, going back and forth with your coding agent or going back and forth to create some financial report. Um, but text and voice are actually still pretty different from an evaluation perspective because you have to simulate all these possible, um, audio, background noise, interruptions, et cetera.
It's kind of like going from 2D to 3D. You're just adding a whole new dimension. And so I think of text-based evals as 2D evals and then, you know, voice is adding that third dimension, and then pretty soon you'll have, like, embodied robotics and that adds, you know, all sorts of dimensions
Amir: Absolutely.
Voice is not new. I mean, it's been around for a long time. People have traditionally not liked it. Uh, I guess how does, how does AI play a role in shifting that and changing that?
Brooke: Yeah. That's a great call-out because actually the voice models and the transcription models, if you look back at like 2016... I actually worked on Google Assistant in 2016, and in 2017 actually.
And if you look back then, like the voice is actually, um, certainly I wouldn't say it's made like orders of magnitude improvement, um, def- I definitely think there's been massive improvements in terms of model quality from the perspective of like real-time capabilities and shifting all sorts of things there.
But there, the thing that really led to the unlock was having a model, a reasoning model that was fast enough, cheap enough, and, um, good enough at reasoning through complex use cases to be able to carry on the conversation. So GPT, um, when GPT-4o came out in summer 2024, that was the first time that we really saw a model that was capable enough to have this multi-turn conversation, which is im- which is required for being able to have a conversation.
And so you can kind of think of it as we had like the eyes and the mouth, but we didn't have the brain, so we finally got the brain to be able to connect these, um, these pieces together.
Amir: Interesting.
what, what does that actually mean for people who will be using and interacting with these AIs?
Brooke: Yeah. I mean, I think that there's so much opportunity in voice because you have this... The reasoning capabilities of agents are just leaps and bounds ahead of what you can do with IVR. But there's definitely this user problem that you mentioned where people are so used to IVR that's totally, um, you know, incapable of doing the thing that you want it to do.
And so as soon as you get on the phone with what sounds like an automated agent, you say, "Human, human, human," or like, "Escalate," et cetera. Um, you assume that... You kind of assume incompetence. But I think we're gonna start to see that agents actually are able to take actions on your behalf, and voice is such a na- a natural way to do this.
So instead of, you know, changing your flight by, you know, you're on a trip and you pull out your laptop and you're going to the website to be able to see all the different flight options, um, in the same way that mobile kind of revolutionized this for airlines, I think we're gonna see another shift, especially in these places where you aren't already in front of a screen, to voice, where, like, you're on the road, you're a truck driver, healthcare.
Like, all these places where voice AI has taken off actually map to a lot of industries where people aren't necessarily in front of a computer. You're in a doctor's office, like visiting a patient. You're, like, um, on the road. You're, like, visiting your clients for HVAC replacements and whatnot. Um, so I'm not only excited there, but I think there's going to be all these new voice interfaces that completely replace, so it's not just calling in on the phone.
But I mean, Perplexity has an awesome comet browser where you can actually interact with your browser and dispatch agents. Especially, like, as a software engineer, my job has completely changed. I don't know if anyone has heard about this, but software engineering is totally different, um, where instead of doing one task at a time, you're, you know, dispatching your, like, a bunch of different agents on different parts of the task.
And so I've noticed as a software engineer, I actually have to break down the tasks in a totally different way than I did when I was a software engineer at Google. Um, it's kind of the... I remember when I switched from Gi- um, GitHub to Google's internal, uh, version of versioning. You couldn't stack branches in, in the way that I used to with GitHub.
And so I remember I had to change my work of, like, how do I think about my work in terms of pieces that I don't have to stack on top of each other, but pieces I can do in parallel? And in a lot of ways, I actually think that's come back, where you're, you have these agents that are going off in, like, lots of directions.
The same way I think that voice is the most natural interface to do that. Like, instead of clicking around, like, you could just have agents coming back to you and you're, like, feeding back yeah, like, "Yes," "No," "Proceed," like, "Change this thing," et cetera.
Amir: That's interesting. You know, as you were kind of going through that and talking about, you know, some of, uh, the evolution here that we're seeing, I, I guess when we're looking at voice and we're looking at, you know, you mentioned, um, you know, doing things that are a little bit more difficult in certain situations.
I actually find sometimes I go, instead of doing it on my phone, I'm gonna wait until I'm back in front of my laptop, just because I know it'll be easier. And I, I guess the notion of the phone itself and voice, I think I had, I had someone else on. I, I asked this question. I was... I'm fascinated by it. I'm still curious, is whether we need the screen on our phones because our fingers are what's operating the device.
Like, our voice isn't necessarily... hasn't been good enough. Obviously, for entertainment, like games or, uh, you know, watching a video or stuff like that, I get why a screen's necessary. But is there a world where if voice AI is that good that we can navigate without the need of traditional ways of navigating devices and interacting with other platforms?
Brooke: Yeah. I mean, I think that there's a really interesting, like, evolutionary piece of this, right? Where your brain actually develops the ability to see and, um, and it, like, speak, uh, much earlier than you were able to read. And so you're much faster at perceiving, like, visual information and then speaking and, like, definitely typing, right?
Like, typing is a very new skill, and we're very slow at it. Um, even if you're, like, an insane typer, typist. Um, and then, uh, in terms of, like, being able to speak, it's, like, much faster and more natural. We talk about this all the time at Coval as we move towards a more agentic, uh, platform and, you know, our users are primarily interfacing with Coval through cloud code or through our agent in the UI or through, like, any agentic harness.
Um, you know, like, agent as our piece of software, then, like, what is the role of the UI? And I still think that, like, processing lots of information is much easier for humans to do visually. Um, but controlling and, like, saying what needs to happen next, that is much easier to do via voice. So I can, like... You know, that's why when you're describing something complicated, you say like, "Let's just hop on a quick call right now, and then I can explain it to you," um, versus typing everything out.
So I definitely am really excited about that direction, and I think a lot of these wearable devices, um... We actually have, uh, a prev-- an ex-founder and also a previous, um... And then she, in a past life, was a previous lead of partnerships at Humane, and I think Humane was just, like, a little bit too early to this space where, like, having the Humane pin, it was like voice AI just, like, wasn't quite there to be able to control it.
Um, and now I think if you look at where the space is today, like, it's just matured so much, and being able to control, like, your browser, like, the surroundings, or, like, get information makes a, makes a lot of sense.
Amir: You know, uh, uh, we won't name names, but there are certain, uh, uh, assistants on devices that are, that are not good Right.
I think people have had that experience, and there's been this... A- a- aside from calling in for a customer service line, I think their personal interaction with these devices have not been smooth.
It's been rough. It's like, "Oh, this, this, uh, this, this, uh, assistant can't understand what I want. It can't actually give me the information I want." And then now we're talking about a world where we can actually ask and get this interaction to a level where, I don't know to how close to fidelity of I want this and it gives me exactly that back, but that certainly changes the landscape.
It's still... And again, you said humans. It still needs to get into devi- It need, it needs to be somewhere. I mean, obviously- Yeah ... I think we hear of, uh, OpenAI is working on something, but th- there's got to be solutions to bring that closer to the edge and actually then have a, a, a useful, actionable AI there.
Brooke: I think something that's really interesting to me is how users kind of develop taste and appreciation for quality as a technology matures. So for example, with mobile apps, like, people will say, "Oh, this company has a bad mobile app," or, "Has a good mobile app." Um, and I think everyone knows what it feels like.
It's like, it's really f- There's all these non-visual pieces of it that make it a good mobile app. So it's, like, not buggy, it's really visu- like responsive. Um, you know, it, like, has a lots of capabilities that they want to be able to do. And I think we've seen this with models as well of, like, companies that have a powerful AI versus not powerful AI.
Um, like what people can expect from a chat, a text box in different UI contexts. So if you have like a little help button at the bottom of like an airline website, you're going to expect different capabilities than like a text box in the center of like a AI native company. And users adapt their expectations based on that.
And so I, I actually think that users just, like, need to build trust with places where maybe there were previous contexts where the AI was not as capable. So a lot of this is just repetitions of, like, if they do take a chance on you to try and execute a task, like being able to execute that will solidify those pathways.
And I think that we see this already with like AI models of like, do people like, you know, ChatGPT versus Gemini, whether they use Claude versus all these different, um, models for different use cases. Obviously, I think it's still very early days. Um, but I think people... I think humans are very perceptive and very good at adapting their u- their users pa- their user patterns.
Amir: Absolutely. I guess a question you, you alluded to, um, you know, within engineering right now, uh, the use of agentic to develop Uh, you know, produce code. Uh, obviously you guys are, you know, on the forefront of, uh, developing AI, uh, tools. Uh, what, what does that look like for you guys? Are you... I mean, uh, I'm sure you're using it, but I guess to what extent, where are you on the maturity curve of leveraging agentic within the, the engineering space?
Brooke: Yeah. I mean, it's so fun 'cause I think... I mean, Co- so Coby and I are, um, head of engineering. He-- we worked together at Waymo, and he was... You know, so like we've w- you know, grown up in software pre-AI, and now I think to switch all of our workflows in AI has been so much fun, 'cause I feel like every single week we're, like, completely changing, like, what harness we're using, what tools we're using.
Like, everyone on our team, I think, is constantly experimenting and, like, posting. We have a, like, AI, um, like tips and tricks Slack channel that we have, kind of like posting what people have found that's working or not working. I think AI sovereignty is, like, a really interesting concept that a lot of people are talking about right now, and kind of like how do you start to have...
Like, when do you use localized models versus large models? And moving away from I just use Claude or I just use OpenAI, but actually I have this fleet of tools and I can use, you know, local, um, open source LLMs for a certain set. I can use one as like kind of the dispatcher. Um, you know, maybe I, I have like different sets of models for different sets of use cases.
Um, but then also, like, how do you change your workflow? I alluded to this earlier of, um, having lots of agents going at once and, like, training your brain to be able to th- break down the problem and then also, like, remember the context of what's happening. Um, like doing two totally different tasks is really hard, so I like to kind of take a similar task and break it down into, like, paralyzed...
It's like MapReduce, but for humans. It's like break it down into a bunch of parallel workflows that you're, like, still working on the same thing, but it's doing different pieces, and so you're kind of like checking in on different pieces of it
Amir: Absolutely. Uh, interesting, 'cause obviously you're drawing a distinction between, um, you know, your previous workflows, your current workflows.
W-when it comes to hiring, uh, I'm sure you were involved in hiring in the, in the old ways of doing things, uh, what are you seeing shift in... And the reason I ask is 'cause, um, you know, the span of a year, everyone's, you know, preferences of h-how interviews should be conducted in regards to AI has changed.
Um, are, are you guys experiencing similar? Have you adjusted your processes to reflect, uh, the nature of agentic?
Brooke: Yeah. I mean, I think I was never a fan of LeetCode style problems in general. Like, I love system design interviews or, like, coding interviews where there's lots of right answers, and really thinking through, like, how did you get there is oftentimes more important than, like, what the answer was.
Um, because I think, yeah, like, kind of you're, if you're looking for a very... There are very few times at work where I'm looking for a very specific answer. A lot of times you're like, "I actually just want, like, a, a good answer and a good reason for why you did that." Um, so but beyond that, I think certainly we have always been very pro-AI and using AI because I think today, as a software engineer, a lot of your skill set is do you use AI well?
Not just, like, can AI do this problem, but there's lots of ways to use AI well and not well, and then, like, integrating it into your workflow. We're constantly having this discussion on our team of, like, um, you know, what things should be automated versus, like, what things do you really need, like, human editing on, and being able to modulate that and have that taste.
Uh, some people say taste, but it's kind of that, like, that strategic thinking or that critical thinking of when, when should I step in and be like, "Wait, what am I actually trying to do here? What am I trying to say here?" I think this is especially... I found this really hard with AI, where it's very easy to kind of be lulled into an answer, where you're kind of like, "Well, I guess I could include that detail," if you're, like, writing a, um, if you're writing a blog post or you're doing, um, something more on the writing side.
And same on the coding side, where you're kind of like, "I guess I'll make five files for this," or, "I guess I'll, like, do it in this more complicated way and check for all these errors." Um, but you can end up kind of being like, "Wait, what am I actually trying to do here?" And that's the question that we've been talking about a lot as a team, is kind of like when do you zoom in versus when do you step back?
Amir: Absolutely. I, I guess may- maybe just one other question, 'cause, um, before we recorded, we were joking, you wear a lot of hats. Um, when it comes to AI within y- the day-to-day operations with all these different hats, uh, do you use AI differently depending on the hat you need to wear? I mean, you're a founder, so you've got everyone's gonna be aware you have to wear different hats, but do you use AI at different capacities depending on the hat?
Brooke: Definitely. Um, I think especially as a solo founder, you know, I'm doing engineering, sales, marketing, customer success, like hiring, so on any given day doing a bunch of different things. Um, I think definitely on the review side, I use a lot less AI because a lot of times when I'm, like, looking at what someone's doing, I actually wanna think deeply about, like, what are we doing here?
Like, what's, what's really the goal here? Um, and like, are we taking the right direction? So sometimes I actually find that, you know, as a founder, you end up using less AI in those workflows because you're really trying to figure out what to do, not, like, how to do it. Um, and so I found, like, the more we scale the team, the more I find myself in that.
That being said, I think AI is so helpful as a manager, and, like, allows you to have flat organizations for much longer because you can consume much more information. So I can just, like... I know, like, PRs that are going in. I know, like, all the different, like, sales materials. You know, conversations that we're having with lots of different partners.
Like, you can, um, kind of have a lay of the land much easier and a broader swath of information. But I, I think when it comes to that, like, critical thinking and review, I still find it-- The human brain, man, it's good. But on the engineering side, I find myself using, like, a totally different set of models and tools compared to, um, like creating sales materials or creating, um, like understanding usage and reports and whatnot.
Uh, still will use like Cloud Code or Codex for these things, but in like very different ways. Um, I'm trying to think of examples. Uh, even as, like, useful as just, um, kind of like how you consume the output. Like, I've been loving creating HTML files and how Cloud started creating HTML files and websites of, like, helping to explain, like, what's this customer's usage or, um, you know, like explaining the concepts of Coval via, like, animated graphics in a way that you never would've been able to before with, um, without AI.
Like, that slide deck would've taken me, you know, a week pl- I don't even know if I have the skill set to create animated graphics. Um, so that's been really cool.
Amir: Yeah, that's really cool. Uh, maybe final question
When it comes to being a solo founder W- who do you turn to for a sounding board?
Brooke: Yeah. I think this is true kind of if you're a solo founder or not, is like building a really strong founding team is so important for the company, and like building trust with your founding team that I trust you to make good decisions and we align on our values for this area, and then we kind of share a brain when it comes to, you know, how are we approaching customers?
How are we approaching marketing? How are we approaching engineering? Um, and so I think within the company, I have like, you know... We've also hired people that, you know, I've worked with for a long time. Um, so one of my best friends, Raya, was our founding engineer. Kobe and I worked together for years at Waymo.
So that certainly helps 'cause you have this trust. But I think, uh, even beyond that, like just building really strong leaders on the early team can be really foundational in that. And then also my husband is a founder as well, so that has... that's really great I think for this like super high level, um, kind of the f- like understanding the founder journey.
Um, it's awesome for us to kind of like support each other through that 'cause he's a founder of a different company.
Amir: Man, that, that's a whole episode of itself. Uh, y- darn it, I wish we had more time. I, I, I'm curious about how that works 'cause I don't see that a lot as, uh, to, uh, you know, h- husband and wife both having, uh, startups.
But we'll, we'll save that for a different day. Um, but I do wanna appreciate, uh, Brooke, your time. Thanks for coming on. Thanks for sharing with us. Um, if someone has a follow-up question or they wanna learn more, what are some good ways of getting in touch?
Brooke: Yeah, you can always email me at brooke@coval.dev, and you can also find me on LinkedIn.
I'm always happy to hear your thoughts. We also have been doing a lot of AMAs recently, as well as I post content there. I love to chat with people in DMs and kind of share, share what's working and what's not working in voice AI.
Amir: Absolutely. Definitely appreciate your time. Thanks, thanks for taking the time.
Brooke: Thanks so much, Amir.
Amir: All right. Absolutely. That's it for this episode. Be back again, different guest, different topic. Until then, two things. One, love it if you could share this episode with somebody else who might be interested in voice, uh, or chat AI,
Also, like, subscribe, comment. Let me know how the show's going for you. Until next time, thank you and goodbye.
