How to check that a Customer Success tool really sorts its alerts
A tool that rates too many accounts high risk stops sorting anything, and I have never seen a vendor publish that share. Here is mine: 36 real diagnostics my engine had rated high, judged one by one in July 2026. Twenty-one were over-escalated. Here is the method, and the five questions that go with it.
Sommaire
Suleiman Mulla · Founder of Phano
Published on July 27, 2026
Key takeaways
- The method is four rules: judge the pile that is already escalated, write the conditions before judging, one verdict per case, publish the count. It works on any tool that assigns a risk level, not only on mine.
- A vendor should be able to hand you that distribution. Here is mine: out of 36 diagnostics the engine had itself rated high, 21 should have been medium or opportunity, that is 58.3% of the escalated pile. Seven dominant patterns cover 19 of those 21 cases.
- The corrected calibration is written, put through adversarial review, amended with seven fixes, and deliberately not deployed: it ships only after a before-and-after re-measurement shows that large accounts and near renewals stay on the list. A number that drops without measurement is not an improvement.
Alerts nobody believes anymore
A Customer Success Manager and an Account Manager learn fast to ignore a tool that cries wolf. They stop reading the severity level, they read the account name and make up their own mind. The tool keeps running, it no longer decides anything.
That severity level decides everything else: what rises to the top of the list, what fires an alert, what gets opened first on a Monday morning. In my engine it takes four values: critical, high, medium or opportunity.
An engine that over-escalates does not produce visibly false text. It produces plausible, well written diagnostics about accounts whose health score is fine. The symptom never shows up on a single diagnostic: it only exists in the distribution of a batch. No sales demo can reveal it, and a vendor can go years without seeing it.
Over-escalation is paid for twice: in hours spent on false positives, then in lost trust. The second bill is the one you never make back.
The question to ask before you sign
What share of my accounts does your tool rate high risk, and under what written conditions? I have never seen a Customer Success vendor publish that answer, and I have never seen a buyer demand it before signing. The two go together: the question does not get asked, so the measurement does not get made.
I do not think anyone is hiding it: nothing forces a vendor to run that measurement. Running it means judging your own output against written conditions, case by case, knowing in advance the result will look bad.
I build Phano, an engine that diagnoses every account and prepares every meeting, with no dashboard to open. So I am judge and party here, and I would rather say it upfront. What I put on the table is not a promise of accuracy, it is a measurement: it covers my own engine, it is published with its protocol, and you can hold it against me.
So what follows serves two purposes: auditing the tool you use today, and judging the one I sell.
The audit protocol, in four rules
The method does not depend on Phano: it works on any tool that puts a risk level on an account.
- Judge the pile that is already escalated. Take the diagnostics the tool itself rated at the highest level. The denominator is not the whole portfolio: it is what came up, so it is what eats your team's time.
- Write the conditions before you judge. Three conditions in my case: a dated and observed fact, a named metric with its value, no contradiction on the point the case rests on. Those are what settled every case.
- One verdict per case, out of three values. Justified when all three conditions hold; over-escalated when one is missing, meaning the diagnostic should have come out medium, sometimes opportunity; borderline when the case can be argued both ways.
- Count, then publish the count. The distribution is the result. A well chosen example proves nothing, in either direction.
Applied to Phano: 36 real diagnostics from one customer's portfolio, all of them already rated high by the engine, in July 2026. The judgment, then its review, came out of a setup built for the occasion, separate from the engine it judges. Judging the escalated pile does not give you the distribution of a portfolio, only the accuracy of what came out of it.
The weakness is obvious and I put it on the table: a language model judging the output of a language model. The real alternative exists, having those 36 cases settled by the Customer Success Manager and the Account Manager who run the accounts, or waiting for the renewals and comparing outcomes. I did neither. What follows is a model judging a model, and it should be read that way.
If you run this protocol yourself, you have an advantage I did not have: your teams know the accounts. An escalated pile the size of mine, 36 cases, settled in a meeting by the Customer Success Manager and the Account Manager who run those accounts, is enough to tell you whether your alerts still sort anything.
What the method gave on my own engine
Of the 36 diagnostics judged, 21 came out over-escalated, 12 justified, 3 borderline. That is 58.3% over-escalated, 33.3% justified, 8.3% borderline. I recounted the rate from the verdicts themselves: the count adds up, which says nothing about whether the verdicts are right.
Close to three escalated diagnostics in five should have come out medium, sometimes opportunity. None of those 21 accounts dropped out of the analysis: the diagnostic was there, the level attached to it was too high. At that rate it is not noise on a few difficult accounts: it is an escalation bias written into the very instructions that define severity. That kind of bias is fixable. You still have to see it, and to see it you have to count.
21 / 36
diagnostics over-escalated, out of the 36 the engine had itself rated high. This is the number I have never seen published anywhere else, because it is never measured. The denominator is not the portfolio: it is the pile that was already escalated, the one that lands at the top of the list of a Customer Success Manager or an Account Manager. Seven dominant patterns cover 19 of those 21 cases; two over-escalations stay outside that grid.
These seven patterns are not specific to Phano: they are the ways an alerting engine gets the level wrong. Look for them in yours, they carry the same signatures.
| Dominant over-escalation pattern | What the data says | Cases |
|---|---|---|
| Contradictory signals | One family of signals fires on silence, another shows active stakeholders and a contact reached recently. | 8 |
| Low confidence treated as sufficient | Assertive diagnostic built on a confidence of 0.25, 0.29, 0.44 and 0.29. | 4 |
| Data gap read as decline | Adoption not measured, emails not synced, last contact and ARR empty. Absence becomes the proof. | 3 |
| Timeline inconsistency | Last contact dated July 17, 2023 against 5 recent visits claimed. | 1 |
| Short silence on a healthy account | Silence of 22 days, health 68, Business Review already scheduled. | 1 |
| Remediation under way read as departure | Compensation already paid, health 70, counted as a worsening. | 1 |
| Renewal risk mistaken for opportunity | POC progressing, sponsor identified, health 70, tagged as risk. | 1 |
Three of those seven patterns are not about reasoning but about reading the data: contradictory signals ignored, an empty field taken as proof, an inconsistent timestamp. They account for 12 of the 19 cases covered. CRM data quality decides the quality of the alerts before the model even weighs in, and it is the first thing to make a vendor answer for.
The three markers of a high risk that holds
The 12 diagnostics judged justified share the same structure. These three markers are not a discovery: they are the criteria used to settle each case. What they teach sits in the negative, in what the other 21 were missing.
- A dated, observed fact. Not an inference, not a projection: an event that took place, with its date.
- A named metric and its value. A silence of 259 days, a health score of 42, a downsell in the pipeline. The number is in the text, checkable at the source.
- No contradiction on the point that matters. A diverging peripheral signal is tolerable; a contradiction on what the whole case rests on is not.
Those 12 diagnostics came out of the same batch, the same engine, without a Customer Success Manager or an Account Manager having to open the account. That is what you buy: not an engine that never gets it wrong, an engine where you know the conditions under which it is right.
One case in the sample comes down to three values: a silence of 259 days, a health score of 42, two independent families of signals converging. There is nothing to interpret, only something to do. That is the diagnostic a Customer Success Manager or an Account Manager should find at the top of the list, and it is exactly the one that too many alerts make invisible.
I tightened two things in the corrected calibration. On paper, here is what it would change.
- Confidence floor. A confidence below 0.5 would no longer carry a high, except on a dated critical trigger, which would override that floor.
- Gray band from 0.5 to 0.6. High would only be allowed there if the signals converge on the point that matters.
- Number of sources. An account tracked by few sources would not be judged less risky: the number of sources would drive the confidence label on display, not the severity.
These values are written as indicative anchors, not as hard cutoffs.
Why this fix will not ship without a re-measurement
On the sample judged, the corrected calibration would move 18 of the 21 over-escalated cases down to medium, exactly half of the 36 diagnostics judged. Where the other three land is not settled in the write-up, and at least one case belongs in opportunity. That is a projection, not a measurement: the re-measurement has not happened.
That is precisely why I do not ship it. The revenue at risk on display would drop mechanically, without a single euro of risk having disappeared. A vendor who deploys this fix without measuring it sells customers an improvement that only happened in the interface. Between a number I know is too high and a reassuring number that is wrong, I keep the first one.
The opposite flaw exists, and it is far worse. An engine that over-escalates wears people out; an engine that is too cautious lets a real churn risk slip past on a large account, and it stays invisible until renewal day.
So I put it through three adversarial reviews, and two returned FAIL. That is exactly what they are for: the first version pulled real risks down to medium and contradicted itself on its own source-count threshold, and it will not ship in that state. A fix that is not attacked before it ships is a fix shipped blind.
The naive fix
An empty last contact or one contradictory peripheral signal would be enough to drop the account back to medium, whatever its health score and its value.
The data gap becomes an excuse.
The two guardrails written to stop it
A very low health score, around 35 or below, combined with a renewal in the next 8 to 10 weeks or a high ARR, would hold the high. And a short silence would not pull an account down to medium when the renewal falls in that same window and the health score is low.
The gap would only de-escalate when the case rests on it.
Seven fixes had to go in before the text held. Two cases are enough to show why. A renewal of existing revenue given a zero probability, on a high ARR account, fell back to high or medium because a date field still pointed to the future: the fix as written would make the content win over the field. And a high ARR account with a health score of 30 ended up de-escalated by a peripheral data gap.
That second case is cited by the reviews but does not exist in the 36 verdicts. The guardrail that comes out of it therefore rests on adversarial reasoning, not on an observation, and it still has to be re-tested against real data.
The corrected calibration is written, reviewed, amended, and kept out of production. The shipping criterion is fixed in advance: a before-and-after re-measurement where the share of high and critical drops without a single high ARR account or imminent renewal leaving the list. Until that is measured, it does not ship. It is a refusal, not a delay.
What this measurement does not prove
- A model judged a model. The better arbiter is still the Customer Success Manager and the Account Manager who run the accounts, or the real outcome of the renewals. Neither settled these 36 cases, and the reviews of the fix come out of the same tooling.
- One portfolio, one date. The 36 diagnostics come from a real portfolio, in July 2026. Nothing says the distribution is the same elsewhere.
- The cases are not published. They cover accounts that are not mine. You cannot redo this judgment, only reuse the protocol.
- The fixes are projected, not measured. The 18 moves down to medium are a projection on the sample, and one of the guardrails comes from a case absent from the 36 verdicts.
What is reproducible here is the protocol, not the result. I publish it anyway, because I know no other way to make the question arguable. An engine that is never judged against written conditions has no reason to produce a fair distribution, and its vendor no reason to notice.
The five questions to put to a vendor, me included
A Customer Success Manager and an Account Manager should not have to audit a prompt. They can, however, ask these five questions before they buy, and listen for whether the answer comes back in numbers or in adjectives.
- The distribution, not the example. Ask for the real spread of severity levels across your accounts. A tool where most accounts come out as high risk is not measuring risk.
- The conditions in writing. Under exactly what conditions does an account become high risk? If that answer does not exist in writing, the level cannot be argued with, so it cannot be checked.
- The evidence inside the alert. Every alert should carry a dated fact and a named metric. Without both, it is not a diagnostic, it is a well written hunch.
- How gaps are handled. What does the tool do when the data is missing? A missing measurement is not a decline, and an empty field must never become evidence.
- The re-measurement after a fix. When a vendor recalibrates, demand the before and the after on your own accounts, and proof that large accounts and near renewals did not vanish on the way.
Here are my answers for Phano, as of the date of this article. You can put the same questions to any other vendor.
| The question | My answer for Phano |
|---|---|
| The real distribution | Measured on the escalated pile of a real portfolio: 21 over-escalations out of 36, 12 justified, 3 borderline. On the escalated pile, not on the whole portfolio, and I say so rather than round it off. Published here, protocol and limits included. |
| The conditions in writing | Three conditions: a dated and observed fact, a named metric with its value, no contradiction on the point that matters. They are in this article, so they can be argued with. |
| The evidence inside the alert | The criterion is written and published: a dated fact, a named metric, no contradiction on the point that matters. It is what surfaced the 21 over-escalations, where at least one of the three conditions was missing. Of the 19 cases tied to a pattern, 12 came from evidence that was absent or contradicted. |
| How gaps are handled | 3 of the 21 over-escalations came from an empty field read as a decline. The fix as written would forbid a missing measurement from carrying a case, and it waits on the re-measurement before it ships. |
| The re-measurement after a fix | Criterion fixed in advance: the share of high and critical has to drop without a single high ARR account or near renewal leaving the list. Until that is measured, nothing goes to production. |
I do not claim the engine is perfect. I claim I know what it produces, that I wrote it down, and that I can show it to you with the numbers in hand. Between a vendor who measures and a vendor who asserts, that is the only difference you can check before signing.
The engine I am describing is the Phano composite diagnostic, delivered every day into your own tools, with the meetings module that prepares the brief before and the recap after. You can put it through this grid on your own accounts on a free trial, then hold the numbers against me.
Frequently asked questions
How do you check that a Customer Success tool does not over-escalate its alerts?
By looking at the distribution, not the examples. Take a sample of diagnostics the tool already rated high risk, write down the conditions a real risk has to meet, then have every case judged against those conditions and count. Of the 36 diagnostics the engine had itself rated high, in a real portfolio, 21 should have come out at a lower level, that is 58.3% of the pile already escalated. When a model judges a model the measurement stays arguable: the judgment of the Customer Success Manager and the Account Manager who run the accounts remains the better arbiter.
What is over-escalation in an AI diagnostic?
It is when an analysis engine rates an account as high risk when nothing warrants it: missing data read as a decline, contradictory signals ignored, low confidence treated as an established fact. The account does not drop out of the analysis, the level attached to it is too high. The consequence is direct: alerts lose their sorting value and the real risks drown among the false ones.
Read the definition: CRM data qualityWhat makes a high risk justified?
Three markers present together: a dated and observed fact, a named metric with its value, and no contradictory signal on the point the case rests on. A silence of 259 days on an account with a health score of 42, confirmed by two independent families of signals, meets all three. An adoption figure recorded as not measured meets none.
Read the definition: Customer Health ScoreIs a drop in revenue at risk after recalibration good news?
Not in itself. When an over-aggressive calibration is corrected, the amount on display drops mechanically without a single euro of risk having disappeared. The only way to check that the risk really went down, and not just the display, is a before-and-after re-measurement, verifying that the high ARR accounts and the imminent renewals stay on the list. That is why the corrected calibration described here is written, reviewed, and kept out of production.
Why publish the distribution of your own severity levels?
Because a vendor who cannot give you that distribution has no reason to produce a fair one, and no way of noticing that it is not. A buyer cannot catch over-escalation in a demo: it only exists in a batch. Publishing it with its protocol is the only proof I can offer someone who does not know me.
Hello,
Your priorities today (2)
● €85,000 ARR · Health 34/100
Silent for 28 days and quote not opened. No meeting in 6 weeks.
→ Escalate to the sponsor
● €42,000 ARR · Health 58/100
Adoption declining on 2 key modules.
→ Schedule a usage review
Opportunities (1)
● €120,000 ARR · Health 88/100
→ Propose an Enterprise upgrade
All diagnostics →
Look at the distribution of your own alerts
Connect your CRM, then put the levels the engine gives your accounts through the grid in this article: critical, high, medium, opportunity, spelled out. On your accounts, not on a demo. Free 30-day trial, no credit card.
Start for freeYour CSMs see the risks, your Account Managers the opportunities. The first diagnostic arrives the same day.
Try for freeGo further
Composite AI
What composite AI means, the 6 crossed techniques, and why crossing detects what a single score misses.
GuideReduce B2B churn
Spot accounts that are slipping before renewal, cross weak signals and act while there is still time.
GuideAI Customer Success
What AI actually does in CS and AM: analyze accounts, predict risk, suggest the action. And its real limits.