CPA Tuning
The default Call Progress Analysis (CPA) thresholds are a calibrated starting point, not a finished configuration. They are derived from broad, mixed call data — but no single campaign looks like the aggregate. Greeting length varies sharply by region, by line type (landline versus mobile versus VoIP), and by local greeting culture: a curt "hello?" behaves nothing like the longer, scripted answer common on business lines or in some markets. When the thresholds that separate a live person from a recorded message are misaligned with the calls a campaign actually places, the cost lands directly on two operational metrics. Classify too aggressively and you drop live humans — an abandoned call, which is a customer-experience failure and, under regimes such as the US TCPA and Ofcom's abandoned-call rules, a compliance exposure. Classify too conservatively and you route voicemail and answering machines to live agents, eroding agent utilization and distorting predictive-dialer pacing. Tuning CPA is the discipline of moving those thresholds toward the population a campaign is dialing, and measuring the trade-off as you go.
Companion guides
Capacity Private Cloud provides analysis tooling to help you understand precisely how a campaign is performing and to refine CPA behavior against real traffic Analysis Portal. For grammar and semantic interpretation testing, use the Deployment Portal Grammar Tester.
Effective tuning is measurement, not guesswork. The workflow is to gather a representative sample of call audio from the campaign you intend to tune — representative of its regions, line types, and time-of-day mix — load it into the platform's tuning tooling, transcribe and label it so you have ground truth, and then measure the actual distribution of greeting lengths. Only then do you adjust the classification thresholds and re-run against the same labeled set to confirm the change reduced misclassification rather than simply shifting it from one error type to the other.
Where the default thresholds come from
CPA classifies an answered call by measuring how long the answering party speaks once media begins, then comparing that duration against two thresholds. The boundaries that separate HUMAN RESIDENCE and HUMAN BUSINESS from the UNKNOWN SPEECH typical of longer recorded greetings were derived by our speech engineers from large volumes of real-world call data over an extended period of analysis. They represent the central tendency of a broad population — which is exactly why they are a defensible default and also why a specific campaign may warrant adjustment.
This is where the default values for CPA_HUMAN_RESIDENCE_TIME (1800 ms) and CPA_HUMAN_BUSINESS_TIME (3000 ms) came from.
As you can see from the above illustration, we calculated the statistical average for HUMAN RESIDENCE is below 1800 ms, HUMAN BUSINESS to be between 1800 ms and at most 3000 ms, and anything longer than that is most likely a recorded message of some sort, so is classified as UNKNOWN SPEECH. The decision as to how to handle UNKNOWN SPEECH classification is very much implementation dependent, so it is up to your application to determine how to respond to these classifications.
An additional classification is also possible, where no human speech was detected at all in the audio stream. This is classified and returned in the results as UNKNOWN SILENCE. Again, how to react to this classification is application-specific, however most treat this as if a machine had answered (similar to UNKNOWN SPEECH)
These settings, documented in the Grammars in CPA and AMD article, can be changed by the application whenever the call population diverges from the default assumptions. A campaign reaching mostly mobile numbers, a market whose greeting conventions run longer or shorter, or a multilingual audience can all shift the real greeting-length distribution away from where the defaults sit. The decision should follow the data: measure the distribution for your own traffic, then move CPA_HUMAN_RESIDENCE_TIME and CPA_HUMAN_BUSINESS_TIME to fit it. Each adjustment is a deliberate trade between false positives (treating a live person as a machine, which drops a real conversation) and false negatives (treating a machine as a live person, which sends voicemail to an agent) — tightening one boundary almost always loosens the other, so validate against labeled audio rather than tuning by feel.
It is also worth reconciling these statistics against your own application logs. A significant divergence between what your call records show and what the platform's tuning and analysis tooling reports often points to an issue elsewhere in the system or call flow — answer-supervision timing, media setup, or signaling — rather than to the CPA thresholds themselves.
Why a call was classified the way it was
Threshold changes are only as trustworthy as your ability to see their effect on individual calls. Beyond reviewing logs and the settings in effect when a CPA interaction was processed, the platform's analysis tooling can describe the reasoning behind a given classification — turning an opaque verdict into a measured account you can check against the audio and against ground truth.
This classification detail is available in the Interactions Properties view, opened by selecting an interaction and viewing its Properties.
On the Answers tab of the dialog, alongside the detailed result information that drives the Interactions List, a Call Progress Analysis Summary section breaks down how the classification was reached against the CPA settings that were in effect at the time.
