Interactly's CEO on what separates conversational AI that completes work from chatbots that deflect, with a 7-question buyer's checklist.

The demo has been great since 2017.
I mean that literally. I've sat through conversational AI demos for a decade, first at Cisco and TripAdvisor, then building Huddl, now on the vendor side of the table. The demo was impressive in 2017 and it is impressive today. Healthcare is where impressive demos go to die, and it's worth being precise about why. Because the technology finally is ready, and the organizations deploying it well are seeing real numbers, while others are buying the same disappointment they bought five years ago with a better voice on it.
The gap between those two outcomes isn't the AI. It's everything wrapped around it.
A patient call is almost never just a question. "Do you take Aetna?" is a scheduling call wearing a disguise. "How early should I get there?" is a confirmed visit. "I need my blood pressure medication" is a refill workflow with a protocol behind it, a prescriber in the loop, and a pharmacy at the end.
Most conversational AI in healthcare fails right here. It's built to answer, and healthcare calls are built to finish something. A chatbot that says "yes, we take Aetna" and hangs up has completed nothing. The patient still has to book, and half of them won't call back.
So the first question I'd ask any vendor, including us, is not "how natural does it sound?" It's this: when the call ends, what exists in my EHR that didn't exist before?
Generation one was phone trees and keyword bots. Press two for scheduling. These reduced call volume mostly by making patients give up, which the industry politely called deflection. Patients hate them for good reason.
Generation two was intent chatbots, roughly the 2018 to 2022 wave. They classified what you wanted and served an answer or a form. Fine for office hours. Useless the moment a conversation had a second step, and in healthcare every conversation that matters has a second step.
Generation three is LLM-driven agents, and they are genuinely different. They handle interruptions, accents, topic changes, the patient who says "next Tuesday" and then "actually, wait, does Thursday work?" For practical purposes, the conversational layer is solved.
Which is exactly why the conversation is no longer where the risk lives. The risk moved.
Put a raw language model on a phone line and it will happily improvise. In retail, improvisation costs a refund. In healthcare it's the one failure mode you cannot have. Not with scheduling rules, not with controlled-substance refill policies, not with a patient describing chest pain.
The unglamorous work that decides whether conversational AI belongs in a clinic looks like this.
Protocols, not vibes. The agent's next step should come from your clinical and business rules, not from the model's mood. Refills follow your refill protocol. Symptom mentions trigger your escalation rules, every time, with an audit trail.
Action boundaries enforced outside the model. The system should know, structurally, what it is never allowed to do, and that boundary can't live inside the model, because models can be talked out of things. We took this seriously enough to publish: our work on pre-action governance and phase-level evaluation appears at AAAI 2026, and our INSURE-Dial benchmark at EACL 2026. I'm biased, obviously. But "our safety approach is peer-reviewed" and "trust us" are different sentences, and buyers should notice the difference.
Documentation as a first-class outcome. If the interaction isn't written back into the EHR, you've created a shadow channel your compliance team will eventually find. They will not be smiling.
And escalation designed as a feature, not an apology. The question is whether the handoff carries context, or whether the patient starts over with a human who knows nothing about the last four minutes.
Numbers from our own deployments, because I can vouch for them.
Capital Nephrology, a specialty group, moved appointment reminders and confirmations to AI agents. No-shows fell from 15 percent to 6.6 percent, roughly $120K a year in visits that now happen.
PromiseCare, a health plan, runs member outreach across more than 300,000 members. Care-gap work that used to consume a call center for a week completes in about an hour, and preventive care visits rose 21.5 percent.
At platform scale, Luma Health runs 150,000+ bilingual patient conversations a month on our infrastructure. Across everything, we've completed over 3 million patient interactions in 15 languages.
Notice what's missing from that list: containment rate. Containment measures how often you avoided helping someone. Booked visits, closed gaps, completed refills. That's the scoreboard.
Seven questions sort this field quickly.
A vendor who answers all seven without flinching is worth a pilot. A vendor who redirects to how natural the voice sounds is selling generation three's demo with generation two's plumbing.
The interesting shift over the next few years isn't better voices. It's conversational AI becoming the front end of operational healthcare: eligibility, prior authorization status, referral coordination, post-discharge follow-up. The phone call becomes the API for the large share of healthcare that still runs on phone calls.
The organizations that win won't be the ones that saw the best demo. They'll be the ones that picked systems built to finish work, measured them on outcomes, and expanded one workflow at a time. In healthcare, conversational AI is worth exactly as much as the work that's complete when the conversation ends.
Nava Davuluri is the CEO and co-founder of Interactly.ai. Interactly's research on AI governance appears at AAAI 2026 and EACL 2026. Related reading: what a medical answering service costs and why most reminder systems don't move the no-show number.
Book a 30-minute walkthrough with our team. We'll load one of your real workflows into the sandbox.
Book a demo to discover the transformative impact interactly can have on your organisation.