Edvard Osnaya ← Back to portfolio
Case Study 02 · Research, Service Design & Usability Testing

Rebuilding trust in a score no one believed

An internal customer-health scorecard was supposed to predict renewals. Instead, teams manually re-verified its data before every client meeting — and some had stopped sharing its reports entirely. I led the research and design workstream that exposed why, and validated the redesign that shipped.

Enterprise SaaS / CRM platform Customer-health scoring platform Internal B2B tool + external beta 35 research participants
RoleLead Designer — research & insights workstream
TeamStrategy lead · UX researcher · Designer
Duration~8-week research program
Methods1:1 interviews · Service blueprint · Expert audit · Usability testing
16
1:1 discovery interviews across three roles
19
Usability-test participants — internal & external
15
Expert audit findings, severity-scored
3×3
Design variants A/B/C tested on two features
Context

A score that drives renewal decisions

Inside one of the world's largest CRM platforms, a customer-health scorecard aggregates each account's product adoption, customer expertise, and technical health into a single score. Customer Success Managers, Account Executives, and Renewal Managers all lean on it — to prepare business reviews, forecast attrition risk, spot expansion opportunities, and build renewal strategy 90–120 days out.

The stakes are high: the score feeds forecasting, pricing conversations, and — for some roles — compensation. But by the time we started, the tool's core asset was gone. Not its data. Its credibility.

The Challenge

Data distrust had become the real product problem

Scoring gaps — on-premise deployments and newer AI product SKUs that the model simply didn't track — produced false "0" scores on healthy accounts. Users couldn't tell a genuine adoption gap from a tracking error, so they assumed the worst and worked around the tool.

Manual re-verification before every meeting

Users independently validated the data outside the platform before facing a client, because a "0" might just be a tracking error.

Reports withheld to protect credibility

Account Executives had stopped sharing the auto-generated reports altogether — system errors risked reflecting on their own professional competence.

A workflow scattered across five tools

To see unbundled product footprints and contract terms, users constantly context-switched between the scorecard, the CRM system of record, spreadsheets, and personal living documents.

Scores that moved for no visible reason

Back-end re-weighting could swing a score 20 points with zero customer action — leaving teams unable to explain a "win" or a "loss."

The biggest pain point is severe data distrust. I manually verify the data outside the platform before every client meeting — because a "0" could just be a tracking error.

— Senior Customer Success Manager, discovery participant
Approach

Four layers of evidence, built to survive scrutiny

Leadership had heard the complaints. What was missing was a body of evidence strong enough to redirect a roadmap — and specific enough for engineering to act on. I structured the program so each layer answered a different question: who are the users, where does the workflow break, what's broken in the UI, and does the fix actually work?

Layer 1 · Qualitative Discovery

Understand the three roles behind the score

I ran 16 in-depth 1:1 interviews (45-minute deep dives) across the three roles that live in the tool, and turned them into working personas built on real jobs — not job titles.

  • 9 Customer Success Managers — including lead, technical, and account-dedicated CSMs.
  • 4 Account Executives — core AEs and service specialists.
  • 3 Renewal Managers — senior and principal.
Layer 2 · Service Design

Map where the workflow actually breaks

Interviews surface complaints; a blueprint surfaces systems. I built an end-to-end service blueprint of the customer lifecycle across all three roles — evidence, actions, emotions, and the backstage tooling behind each step — which made the fragmentation impossible to argue with.

Layer 3 · Expert Audit

Evaluate the interface against usability heuristics

A screen-by-screen expert review of the dashboard produced 15 findings, each scored on a 0–4 severity scale, mapped to a usability heuristic, and paired with a concrete, low-effort recommendation and its expected impact.

Layer 4 · Validation

Test the redesign before it ships

Ahead of a major product release, I ran unmoderated remote usability testing with 19 participants — 8 internal operators and 11 external users new to the tool — using think-aloud protocol on interactive prototypes, plus A/B/C preference testing on two contested design decisions.

The Blueprint

Three phases, three roles, one shared picture

The blueprint traced the full account lifecycle — showing, at every step, which tool each role reached for, what they felt, and where the platform quietly handed the work back to them.

Phase 1

Baseline adoption & health monitoring

Routine health checks and manual tracking to cover visibility gaps — plus hunting for upsell signals in utilization metrics.

Phase 2

Mid-cycle value realization & business reviews

Preparing customer-facing reviews, reconciling adoption data against contracted licenses, and rewriting AI-generated decks by hand.

Phase 3

Renewal strategy & expansion

Modeling pricing uplifts in spreadsheets, hunting shelfware to protect the negotiation, and opening the renewal conversation 90 days out.

The blueprint became the artifact everyone pointed at — product, engineering, and leadership finally looking at the same map instead of at each other.

Findings

15 audit findings, scored and owned

Each finding got a severity rating, a heuristic, a recommendation, and a stated impact — so prioritization was a conversation about evidence, not opinion.

2
5
4
4
2 catastrophic — imperative to fix 5 major 4 minor 4 cosmetic
Sev 0

Unexplained score fluctuations

Pain pointScores moved from back-end re-weighting and new signals, not customer behavior — so teams couldn't explain wins or losses, and the scorecard lost credibility with clients.
RecommendationA "score insight" transparency drawer that separates customer-driven changes from system/algorithm updates, plus a proactive alert whenever the model changes.Visibility of system status
Sev 0

Unclear signal names

Pain pointSignal names were dense and undefined in-product, pushing users out of the platform to research terminology and fracturing the workflow.
RecommendationInline definitions on hover, in plain language, with the business value of each signal — keeping the user in the workflow.Recognition rather than recall
Sev 1

Bundled product data

Pain pointPaid add-ons were buried inside broader categories, making it impossible to tell whether a customer used the capabilities they were paying for.
RecommendationSurface every contracted product as its own line item, with a detail view that splits bundled contributions.Match with the real world
Sev 2

Signals buried in a hierarchy tree

Pain pointFinding one specific signal meant clicking through nested subcategories and waiting on page loads — with no guarantee it lived where users expected.
RecommendationAdd signal search, plus "recently viewed" and pinned signals.Flexibility & efficiency of use
Validation

Then we tested the fixes — and killed our favourite

Two design decisions were genuinely contested inside the team. Rather than let seniority settle them, we put all three variants of each in front of real users. The results were not what the team expected.

Option A

Side panel

Polarizing
Split by audience

Some external users liked keeping the dashboard visible behind it. Internal teams rejected it outright — it felt disjointed from their workflow.

Winner
Option B

Modal window

61.1%
overall approval · 85.7% of internal users

Captured and focused attention; the centralized call-to-action made next steps immediate and obvious.

Option C

Embedded list

Rejected
Both audiences

Keeping recommendations inline preserved context but produced visual clutter and cognitive overload.

What validated

91% found the new recommendations section easy to locate (internal users averaged 15 seconds). The critical-alerts banner scored 5/5 for visibility with 100% of internal users. And both audiences endorsed the shift from generic "product adoption" to measurable business objectives, and from monthly to weekly change tracking.

What we got wrong

The "explore feature" control scored only 66.7% discoverability — and worse, it opened an AI chat when users expected static learning material. Internal users called the repeated buttons clutter. Testing caught it before release, not after.

The most useful finding was the uncomfortable one: a subset of users disliked all three variants, calling the experience overwhelming. We reported that plainly. Option B shipped as a baseline — not as a finish line.

Outcomes & Impact

Trust became a design requirement

Reframed the problem
Leadership stopped treating this as a data-accuracy bug and started treating explainability and trust as product requirements — the transparency drawer, change notifications, and in-product signal definitions all trace back to the research.
Validated a release
Three beta features were tested and refined before a major product release, with the winning variants and a documented list of the friction still to solve.
One shared map
The service blueprint gave Product, Engineering, and the Customer Success org a single, agreed picture of the workflow — including every workaround the platform was quietly forcing.
Prioritized backlog
15 severity-scored findings with recommendations and impact statements became actionable, arguable-free input to the roadmap.
Reflection

What I took from it

The scorecard's problem was never really the score. It was that the product asked people to stake their professional credibility on a number it wouldn't explain. Once we reframed "accuracy" as "explainability," the roadmap almost wrote itself — and testing the beta features taught me that the finding worth protecting is the one nobody on the team wanted to hear.

The highlighted line is a placeholder in your voice — replace it with the lesson that feels most true to you.