AI Agent Reporting

From Metrics to Meaning: Rethinking AI Agent Reporting

By: Damon Lanphear, CTO, Artera

Reporting in patient communication has historically meant conventional quantitative metrics. We track how many patients scheduled, how many messages were sent, and how long before someone responded through dashboards our customers know quite well. But a number can only tell you how many times an event occurred or did not occur. Quantitative metrics seldom tell you why. When our goal is to drive deep continuous improvement in patient experience, we need to understand the “why.”

At Artera, we’ve been focused on how we can dive deep into the “why” behind our patient experiences, and then use that insight to drive improvements for our patients and customers.

The limits of quantitative data

Imagine you’re an operator at a clinic. You open your weekly report and see that the fill rate for your available appointment slots dropped week over week. Now what? How do you translate this metric trend into a root cause on which you can act?

Until recently, you might take the approach of sampling scheduling conversations. You’d pull a handful of conversations where a patient failed to schedule, read the transcripts, form a theory, and hope it generalized. It was slow, inherently biased, and depended entirely on whether you happened to look in the right place.

Sampling was required because we lacked the tools to analyze text transcripts at scale. It is not feasible for a few people to read the thousands of scheduling transactions and synthesize common root causes. For a single enterprise, that can be four to five thousand conversations every single day. No team can read all of that, let alone find the patterns hiding inside it. We reported what we could measure, and the “why” stayed just out of reach.

A change in methodology

A dashboard works beautifully when your workflows are deterministic: a fixed set of steps, a fixed set of things to count. When you’re running agents that handle the open-ended range of situations brought by patients, a dashboard can’t keep up. You can’t chart in advance for scenarios you haven’t imagined yet. We changed the methodology to take advantage of the fact that contemporary AI changed the economics of understanding. With the advancing capabilities of large language models (LLMs), we can analyze patient conversations qualitatively, at scale.

Instead of a person sampling transcripts, we have a sophisticated analysis platform driving LLM-based agents to read all of them, surface themes and ground its observations in the goals of the patient and the clinic. The result is a detailed narrative rather than an opaque set of quantitative metrics: scheduling increased because a policy changed on Thursday, and that change let a specific group of patients receive care at a higher rate. 

The narrative shows where you are winning appointments, where you are losing them, and which visit types represent the single biggest opportunity to offload more work. It closes the gap between noticing something and knowing what to do about it, and it does so faster than we have been able to before.

So it’s worth being concrete about what we actually report now. For a given customer, in a given week, we’ll surface where patients are successfully getting care and where they’re falling out, broken down by the providers they’re asking for and the reasons they’re calling. We’ll flag the visit types patients want but can’t currently book, which is often the single biggest opportunity to hand more work to the agent. We’ll show how a specific scheduling policy or standard of care is playing out once thousands of patients run into it. We take the next step to translate these narrative observations into recommendations for your specific improvements that we apply and validate through the next review cycle.

A conversation with your own data

The clearest proof of how much this changes things is that we now run our own operations this way. We point the same capability at our live product. One pipeline reads all of our AI agent conversations, automatically redacts PHI, and stores them so we can ask questions in plain language. When a customer went live with warm transfers, we could ask how it was going: whether patients were getting where they needed to be, whether the transfers were working, and where the friction was. No dashboard would anticipate those questions in advance. This is a conversation with your own data. It is the same shift we describe for providers, turned on ourselves: the questions worth asking are often the ones you did not know to ask until something happened.

Seeing your own operations clearly

Every clinic, often every physician, has its own scheduling preferences and standards of care. Those are choices, and we’re bound to honor them in the scheduling decisions our agents make. It’s one thing to reason about a policy one patient at a time. It’s another to see its effect across a whole population, when thousands of people arrive at your door every day with every kind of need.

Think about a rule-oriented system anywhere in healthcare. Say you’ve hurt your back. The policy is that you can’t see the specialist until you’ve waited several weeks, undergone physical therapy, and been re-evaluated. Throughout this process, you wait while in pain. When you finally get in, the specialist asks why you didn’t come sooner. The rule that was meant to help improve the standard of care created a delay that eroded the patient’s quality of life.

That kind of pattern is almost invisible one case at a time, except to the clinician who may see this pattern evolve individually. In aggregate, however, it’s obvious and it’s fixable. When a provider can see how a standard of care is actually playing out across their patient population, they may reach the same conclusion themselves. We’re taking the way a clinician already thinks about their practice and connecting it directly to what’s happening at their front door.

A collaboration, not a dashboard

Each report we provide is bespoke to the customer, because every practice is different: some are running outbound and reminder campaigns, some are scheduling complex procedures, and every patient population has its own character.

We’re not waiting for customers to ask the right question and pull the data. We’re bringing the insight to them, and getting to the substance of what matters faster. When a customer reads a narrative and sees an opportunity, we are ready to act on it quickly. In that kind of back-and-forth, we’ve seen key metrics improve by 10, 15, even 20 percent, produced by a cycle of understanding, judgment, and action repeating over time.

That’s the part I find most exciting. The best outcomes come from a blend of AI and human judgment: the technology surfaces what’s happening and why, and an experienced operator who knows their business applies the judgment to steer it. Our job is to make that loop faster, clearer, and more grounded in what patients actually experience.

Related Posts

Artera Co-Founder & CEO Guillaume de Zwirek recently joined host Sandy Vance on The AI at ViVE Podcast, published by...
Voice AI is no longer experimental in specialty patient access. Practices across the country are piloting it, deploying it, and...
By: Dan Goldsmith, Co-Founder & Partner, Proofpoint Capital Everyone wants to know the same thing right now: as AI becomes...
Connect with Us