afv-library/plugins/builder/salesforce-development/skills/agentforce-generate/references/voice-latency-heuristics.md

9.5 KiB
Raw Blame History

Voice Latency Heuristics

Voice agents live and die on latency. Under ~2 seconds between "caller stops speaking" and "agent starts speaking" feels natural; every additional second erodes the illusion of talking to a person. Text agents can hide a slow tool behind a spinner — voice agents can't. This reference encodes the latency anti-patterns worth catching at authoring time (/agentforce-generate) and diagnosing from traces (/agentforce-observe).

Each pattern has three parts: detection (how to spot it in the .agent bundle or a trace), impact (typical cost), and fix (instructional patch, action/tool change, or flag-for-human).

Fix classification. Only purely instructional fixes (adding an ack phrase, tightening turn length, adding a spoken-form rule) are safe to apply automatically. Anything that changes an action's implementation, its synchronicity, or an org/channel-level setting is flag-only — surface it with the specific object/action named and let a human decide.


1. Synchronous writes on the live-call critical path

The single biggest latency killer in the field. Any write (DML, external callout, trigger cascade) that runs inside the turn between caller-speaks and agent-speaks adds directly to perceived response time.

Detection

  • An action's target performs a write and runs synchronously in the reasoning turn.
  • The action name/description implies mutation: update, insert, log, commit, save, record, create, submit.
  • A trigger fires on the target object (custom or managed-package).

Impact

  • Simple insert with a normal trigger: ~300800ms.
  • Managed-package trigger cascade on a heavy object: multiple seconds.
  • Chained callout inside the trigger: add ~13s per hop.

Fix

  • Preferred: move the write to post-call via a Platform Event or async queue.
  • Interim: wrap the write in a Queueable and return immediately from the action.
  • Instructional (auto): add an ack phrase before the action ("One moment while I get that on file.").
  • Flag for human: if the write MUST happen mid-turn (authentication, payment authorization), don't rewrite — flag with the specific object and trigger names.

2. Large result sets returned raw to the reasoning LLM

Retrieval that returns the full body of every matching record is a hidden latency and accuracy tax — the model has to filter needles from a hayfield every turn, on a bloated context.

Detection

  • Action description mentions search, retrieve, knowledge, articles, documents.
  • Action returns more than ~5 items or ~2KB of text per call by default.
  • No limit / top_k / max_chars parameter surfaced to the planner.
  • Return type is the raw object (article, record), not a summarized shape.

Impact

  • Seconds per call, plus downstream reasoning slowdown on the bloated context.
  • Degraded answer accuracy.

Fix

  • Action change: summarize first (title + 12 sentence excerpt), page the rest behind a follow-up action.
  • Action change: add a limit parameter with a low default (35); name the allowed range in its description (Agent Script has no enum input attribute — see actions-reference.md "Voice-Safe Action Authoring" rule "Enumerate small value sets in the description").
  • Instructional: if the action stays bulky, add an ack phrase ("This can take a few seconds — hang with me.").
  • Flag for human: if the customer intentionally wants wide-net RAG over their whole KB, surface the tradeoff instead of rewriting.

3. Chained external callouts

Every external HTTP hop is a chance for tail latency. Two hops brings the p95 of both to the caller.

Detection

  • An action's implementation makes 2+ external HTTP calls per invocation.
  • A trigger on the target object hits a webhook or external system.

Impact

  • Per hop: ~200800ms typical, seconds at p95.
  • Two chained hops: p95 easily crosses several seconds.

Fix

  • Flag for human with the specific hops named — the topology matters.
  • Where feasible, propose consolidating hops server-side or pre-fetching in parallel.

4. Over-decomposed multi-subagent routing

A monolithic agent is often faster and more accurate than a deep tree of subagents fronted by a classifier. Each planner turn costs ~12s; each subagent handoff can add another. This is the voice cost of the "dead hub" and "over-routing" anti-patterns already tracked in known-issues.md.

Detection

  • The bundle has several subagents whose only job is to route to one other subagent.
  • A classifier/guardrail subagent intercepts every utterance before the primary agent.

Impact

  • Extra planner round-trips per turn; unexpected orchestrator calls when subagents chain.

Fix

  • Flag for human: collapsing multi-subagent → flatter routing is a design decision.
  • Where the structure must stay, ensure each hop has an ack phrase.
  • Consider whether a guardrail subagent can become instruction-level rules in system:.

5. Long agent turns trip the silence/nudge timer

On long agent turns, the platform's "are you still there?" nudge (Speak-Up) can fire before the caller has had a chance to respond, producing an awkward overlap or dead air. This is the same root cause as the response-verbosity check in /agentforce-observe.

Detection

  • Response templates or instructions that produce turns over ~30 words / ~10s of speech.
  • Welcome/greeting message longer than one or two sentences.

Impact

  • Dead air, overlap, or a spurious "still there?" prompt on the first or a long turn.

Fix

  • Instructional (auto): enforce ≤2-sentence turns everywhere (the "Keep responses concise" rule in voice-modality-reference.md); keep the greeting to one sentence.
  • Flag for human: if long turns are legally required (disclaimers), the silent/nudge timing is an org-level setting, not an agent-bundle change.

6. TTS number / currency / ID garble

TTS engines read $172,576.81 or +14155551212 as garble unless the model pre-formats them into spoken form.

Detection

  • Instructions or action outputs surface prices, phone numbers, IDs, dates, or account numbers.
  • No rule telling the model to render numbers in spoken form.

Impact

  • Callers hear unintelligible strings; erodes trust and forces repeats.

Fix

  • Instructional (auto): add the spoken-form number rule (the "Render numbers, prices, and IDs in spoken form" rule in voice-modality-reference.md).
  • Flag for human: if the deployment language has weak TTS number-normalization coverage, note the gap so the instruction rule carries the load.

7. Premature end-of-turn under noise (STT endpointing)

Callers in noisy environments (drive-throughs, call centers, home appliances) get cut off before they finish because the STT endpointer signals end-of-turn too aggressively.

Detection

  • Deployment is telephony (PSTN/SIP) rather than a web client.
  • Customer domain suggests a noisy environment (auto, retail, home services).
  • Endpointing silence threshold left at the default.

Impact

  • Truncated caller utterances, repeated turns, frustration.

Fix

  • Flag for human (org-level): raise the endpointing silence threshold for noisy environments (typical safe range ~300900ms; lean higher when noisy). This is a channel/STT setting, not an agent-bundle change.
  • Instructional (auto): add an ASR repair prompt (Rule for misheard input in voice-modality-reference.md) so a truncated utterance recovers gracefully.

8. Telephony transport hops (SIP / multi-region)

Every trunk hop and region crossing on the SIP path adds latency. Largely outside the agent bundle's control, but worth flagging.

Detection

  • Deployment metadata mentions multiple trunks, a non-local region, or a third-party telephony proxy.

Fix

  • Flag for human only. Not an agent-bundle-level fix.

Automatable checks

Grep the .agent bundle (and, for /agentforce-observe, cross-reference the trace) for these:

Pattern Meaning Severity Fix class
Mutating action (name/description implies write) runs synchronously in the turn Sync write on critical path high flag (+ ack phrase auto)
Retrieval action returns > ~5 items / > ~2KB with no limit parameter Bulky retrieval high flag
Action implementation makes 2+ external HTTP calls Chained callout medium flag
Several subagents that only route to one other subagent Over-routing / dead hub medium flag
Response templates > ~30 words, or greeting > 12 sentences Nudge/long-turn risk medium instructional (auto)
Prices / phone / IDs referenced with no spoken-form rule TTS garble medium instructional (auto)
Slow action (SOQL, external HTTP, retrieval) with no ack phrase in instructions Missing filler high instructional (auto)

What NOT to change automatically

  • Synchronous writes required to happen mid-turn (payment auth, identity verification) — flag with the reason.
  • Retrieval where the customer explicitly wants wide-net RAG — flag with the tradeoff.
  • Endpointing / silence / nudge thresholds — these are channel/org-level settings, not agent-bundle.
  • SIP topology — a platform/telephony decision.
  • voice-modality-reference.md — instruction rules (ack phrases, spoken-form numbers, repair prompts, turn length).
  • actions-reference.md — voice-safe action authoring (descriptions, parameter names, error shapes).
  • known-issues.md — dead-hub / over-routing anti-patterns.
  • /agentforce-observe — trace-driven latency diagnosis (Phase 1 voice checks).