Support Architecture for Millions of Conversations: Separating Retrieval, Actions, and Generation
Using a housing booking service as an example, the authors examine the shift from a single model to a constrained orchestrator of typed operations and a generator grounded in verified context; they assess the architecture’s effects separately from accompanying changes. In replayed conversations, booking selection became more accurate, errors in structured actions disappeared, and escalations declined. The optimization also reduced latency and model serving costs, while the volume of conversations handed off to human agents in production remained largely unchanged.








