vix.ing · top · new · best · stats · spec

Evidence-in-the-Loop: Trace-Driven Optimization for Customer-Service LLM Agents

2026/07/20 by Chunming Wu, Dafei Qiu, Congde Yuan +7
#cs.IR

paper · pdf

Abstract

Production customer-service bots must improve answer quality across iterative releases, yet large language models must not bypass evidence boundaries, policy rules, or human-handoff safeguards. We present an Evidence-Grounded Customer-Service Agent Workflow deployed in a real-world customer-service setting. BM25 recall, issue-title-vector recall, issue-description-vector recall, weighted RRF fusion, and cross-encoder reranking construct grounded FAQ evidence for controlled LLM decisions. Policy-guided orchestration then combines this RAG evidence with scenario-specific rule evidence, conversation memory, and clarification state inside a fixed LangGraph DAG~\citelanggraph2024. The paper contributes three reusable deployment patterns: hybrid RAG evidence construction, where multi-channel retrieval and reranking produce auditable FAQ candidates; evidence-grounded issue/action decision, where an Evidence-Grounded Decision Module selects an issue/action from typed FAQ evidence and scenario-specific rule evidence; and trace-driven RAG and reranker improvement, where traces diagnose whether failures come from recall, ranking, final candidate selection, clarification, rule-derived evidence, or action policy, and where reranker fine-tuning is evaluated not only for in-domain gain but also for forgetting risk.

Related