Controlled Experiment · Retention · NPS · Healthcare

Post-consultation NPS experiment — killing an assumption before it became a feature

The CEO hypothesized that patients needed clinical follow-up to improve their post-consultation NPS. I designed a lightweight controlled experiment to test this assumption — at zero development cost. The data revealed the real problem was entirely different. This is the case that best demonstrates what good experimentation looks like.

Blue Medical Guatemala & Costa Rica 2022–2024 Program Manager NPS · Contact center data · QC feedback
$0
development cost — experiment ran operationally before any build
6%
of contacts actually needed clinical follow-up — 94% had operational issues
4 wks
to surface the real problem and redirect investment to the right solution
Confidentiality note: Patient counts and engagement rates are approximated to protect NDA obligations. All figures reflect real work on a live product across multiple clinics and specialties.
Context

At Blue Medical, NPS was more than a satisfaction score — it was a proxy for retention and downstream revenue. A patient who left a consultation dissatisfied was less likely to return, less likely to follow through on lab and prescription recommendations, and more likely to choose a competitor for their next visit. Improving NPS was directly linked to the cross-sell model that drove the company's revenue.

The business problem
  • Post-consultation NPS scores were below the benchmark the business wanted. Industry context: NPS above +50 is considered excellent in healthcare; above +70 is world-class (Zonka Feedback, 2026).
  • The CEO had a clear hypothesis about the cause — and was prepared to invest engineering resources in a solution based on that hypothesis.
  • The question was not whether to improve NPS. The question was whether the proposed solution was addressing the right root cause.
The CEO's hypothesis
  • Patients had unresolved clinical questions after their consultation — things they forgot to ask the doctor, or realized they needed to clarify after leaving.
  • Proposed solution: offer every post-consultation patient access to a follow-up call with their doctor or contact center support.
  • The hypothesis felt intuitive — and that's exactly why it needed to be tested before resources were committed.
🧠 Confirmation Bias 🧠 Narrative Fallacy
Psychology: Confirmation Bias (growth.design) — the tendency to seek evidence that confirms what we already believe. Intuitive hypotheses feel true, which makes them the most dangerous kind to leave untested.
The experiment
Experiment type: Operational (no-code)  ·  Primary metric: NPS score post-consultation  ·  Secondary metric: Reason for contact
Experiment design
  • Rather than build and ship the follow-up feature, I recommended running a lightweight operational test first — using existing contact center staff to simulate the experience without any engineering investment.
  • ~500 patients were selected post-consultation across multiple clinics, doctors, and specialties — ensuring variance in setting, physician type, and medical specialty to avoid selection bias.
  • The test group was offered the opportunity to speak with their doctor or receive contact center support after their visit.
  • The critical design decision: tracking the reason for contact, not just whether contact was made. This turned a binary metric into an actionable dataset.
Why pre-defining the measurement mattered
  • Confirmation Bias (growth.design): without pre-defined metrics and a structured reason-for-contact taxonomy, it would be easy to interpret any result as supporting the original hypothesis. Defining what "success" and "failure" look like before the experiment runs is non-negotiable.
  • Curse of Knowledge (growth.design): the team assumed patients' post-consultation needs were clinical, because that's the lens of a healthcare organization. Systematically capturing the actual reason for contact revealed what patients actually needed — which was completely different.
CRO principle: an experiment without a pre-defined measurement plan produces ambiguous results that can be interpreted to support any conclusion. The reason-for-contact taxonomy was the experiment's most important design element.
No lift
NPS did not improve consistently across clinics or specialties
~12%
of patients engaged with the follow-up offer
6% clinical
only 6% of contacts had a clinical question — 94% had operational issues
What the data actually showed
  • NPS did not improve consistently across clinics or specialties — the intervention had no meaningful effect on the metric it was designed to move.
  • Of the ~12% of patients who engaged with the follow-up offer, only 6% had a clinical question for their doctor. The remaining 94% had operational issues:
  • The majority had problems with insurance coordination — coverage questions, claim processing, authorization delays.
  • A significant portion had issues with lab results not yet delivered or communicated.
  • Others were waiting on medications that hadn't arrived after the consultation.
  • None of these were clinical follow-up needs. They were operational failures downstream from the consultation itself.
Psychology — why the scores were low
  • Halo Effect (growth.design): patients' overall NPS score was being dragged down by operational failures that happened after the consultation — insurance issues, missing lab results, late medications. A great clinical experience followed by a broken operational process produces a poor overall NPS. The consultation wasn't the problem — the post-visit infrastructure was.
  • Peak-End Rule (growth.design): people judge experiences by their emotional peak and how they end. If the consultation went well but the last interaction was a frustrating insurance call, that's what patients remembered when asked to rate their experience.
Research context: NPS in healthcare is driven more by operational smoothness than clinical care quality. Insurance navigation, lab result communication, and medication access are consistently top drivers of patient dissatisfaction (Zonka Feedback, 2026). NPS scores vary significantly by specialty and operational context — emergency settings score 15–20 points lower than elective care.
The recommendation
  • Do not build the doctor follow-up call feature. The data showed it would not move NPS in a meaningful way because it was solving the wrong problem — only 6% of patients actually needed what it offered.
  • Redirect investment to the three operational failures that were actually causing patient dissatisfaction: insurance coordination workflows, lab result delivery communication, and medication fulfillment status visibility.
  • The experiment cost zero in development and produced a directional insight that protected significant engineering time from being spent on a feature with a ~6% addressable use case.
Why this matters as a case study
  • This experiment cost nothing to run. The insight it produced was more valuable than many experiments that cost significantly more.
  • Presenting a null result honestly to the CEO — with data and a clear alternative recommendation — is a harder conversation than confirming a hypothesis. It requires trust in the process and confidence in the data.
  • This is what good experimentation looks like: cheap tests that prevent expensive wrong bets. The goal is not to prove ideas right — it's to find out quickly whether they're wrong before committing resources.
🧠 Sunk Cost Fallacy 🧠 Outcome Bias
CRO principle: the value of an experiment is not measured by whether the hypothesis was confirmed — it's measured by how much better the next decision is because of what you learned. A well-designed null result is a success (VWO, 2026; growth.design).
What I learned
1 — Test before you build, always
The most dangerous hypotheses are the ones that feel obviously true. The CEO's hypothesis was intuitive, well-intentioned, and backed by logical reasoning. It was also wrong. A 4-week operational test with no development cost surfaced this — a full feature build would have taken months and delivered no NPS improvement. The principle: the cheaper the test, the faster you learn, and the more hypotheses you can test before committing to solutions.
2 — NPS measures the whole journey, not just one touchpoint
Patient satisfaction is not determined by any single interaction — it's the aggregate of every touchpoint from booking to post-visit follow-up. Optimizing one step without understanding the full journey produces misleading signals. The Halo Effect and Peak-End Rule mean that operational failures at the end of the experience disproportionately determine the overall score. In healthcare, this means post-visit operational processes (insurance, labs, medications) are often bigger NPS drivers than the clinical consultation itself.
3 — The reason for contact is more valuable than whether contact was made
Measuring engagement rate alone would have produced an ambiguous result — "12% of patients engaged" could support either the hypothesis (patients wanted follow-up) or contradict it (88% didn't engage). Tracking the reason transformed the data from ambiguous to actionable. This design decision — adding one data field to the experiment — was the difference between a result you can act on and a result you have to interpret.
Sources
Growth.Design — 106 Cognitive Biases
growth.design/psychology
Principles cited: Confirmation Bias, Halo Effect, Peak-End Rule, Curse of Knowledge, Sunk Cost Fallacy, Outcome Bias. Applied to experiment design, result interpretation, and recommendation framing.
Zonka Feedback — NPS in Healthcare 2026
zonkafeedback.com/blog/nps-in-healthcare
Healthcare NPS benchmarks: +50 excellent, +70 world-class. NPS varies significantly by specialty — emergency settings score 15–20 points lower than elective care. Operational smoothness (insurance, labs, medications) is a primary driver of patient dissatisfaction beyond clinical care quality.
VWO — CRO Statistics 2026
vwo.com/conversion-rate-optimization/cro-statistics
CRO is an end-to-end process — from research to prioritizing ideas and implementing changes. Documenting experiment learnings, including null results, allows teams to track strategy development and avoid repeating ineffective approaches.
Relias — Healthcare NPS: Relational vs Transactional 2025
relias.com/blog/healthcare-net-promoter-score
14% improvement in NPS response rate when surveys are sent within one day of the visit. Patients express higher satisfaction when reporting on specific recent interactions — reinforcing that operational timing and context shape NPS scores significantly.