The Problem: Serial Ticket Handling in High-Volume E-commerce Support
A 2,000+ employee e-commerce company in the USA handles roughly 50,000 support tickets per month. A significant share of those are order and shipment status inquiries: “Where is my package?” “Why is my order delayed?” “I haven’t received my confirmation email.” Each one lands in a shared Gmail inbox, gets picked up by an agent, who logs into the order management system, checks the shipment tracker, drafts a reply, and sends it. Average first-response time sits at 4-6 hours during peak season, and the cost per ticket is driven almost entirely by agent labor.
The problem is not that agents are slow. It is that the workflow is serial: a human must read the ticket, decide what data to pull, pull it from two or three systems, compose a response, and send it. The AI opportunity is not to replace the agent but to collapse the serial steps into a parallel pipeline where the machine does the retrieval and drafting, and the human does the approval. Forfis approaches this as a workflow orchestration problem, not a chatbot problem. The goal is to cut first-response time from hours to minutes while keeping a human in the loop for anything that touches money or a customer commitment.
The Mechanism: LangGraph Orchestration with a RAG Retrieval Layer
The architecture rests on three layers. The orchestration layer uses LangGraph to define a stateful graph where each node is a discrete step: classify the ticket, retrieve order data, draft a response, check the approval gate, and send. Edges between nodes encode the control flow, including branches for escalation to a human agent when confidence is below threshold. LangChain sits underneath, providing the abstractions for LLM calls, prompt management, and document retrieval.
The retrieval layer is a RAG pipeline. The company’s order management system, shipment tracking data, and policy documents are chunked at the record level and embedded into a vector store. When a ticket arrives, the system retrieves the relevant order record and passes it as context to the LLM. The integration layer connects to Google Workspace via the Gmail API and Google Chat API using OAuth 2.0 with least-privilege scopes. The AI does not replace the mailbox; it drafts responses that a human agent reviews and sends through the existing interface.
The model choice is deliberately model-agnostic. Classification and retrieval run on an open-weight model on the client’s hardware where data residency matters. Final response drafting uses a frontier API (OpenAI or Anthropic) for quality. LangGraph abstracts this, so swapping models does not require re-architecting the graph.
Trade-offs: Latency, Data Residency, and Automation Depth
The first trade-off is latency versus accuracy. A frontier API produces better-drafted responses but adds 1-3 seconds of network latency per call. For a first-response-time target of under 10 minutes, this is acceptable. For a real-time voice channel, it would not be. The second trade-off is data residency versus model quality. Running the RAG pipeline on an open-weight model on-premises keeps customer order data inside the building, satisfying ISO 27001 data classification controls, but the model’s drafting quality is lower than a frontier API. The hybrid approach — on-premises retrieval, cloud drafting — splits the difference.
The third trade-off is automation depth versus risk. Auto-approving every AI-drafted response would cut first-response time to under 2 minutes, but it violates the human-in-the-loop requirement for anything touching a refund or a contract. Forfis sets the approval gate at the record level: routine order-status queries auto-approve above a confidence threshold, but any response that mentions a refund, a delay compensation, or a policy exception routes to a human. This keeps the 90% of tickets that are simple status checks fast while protecting the 10% that carry financial or legal risk.
The fourth trade-off is integration scope versus timeline. A four-week sprint cannot rebuild the CRM or the order management system. The integration is read-only on the data sources and write-only on the Gmail outbox. This constraint is a feature: it keeps the pilot reversible and the blast radius small.
Recommendation: Start with a Fixed-Scope Pilot on Order-Status Tickets
For a 2,000+ employee e-commerce company in the USA targeting ISO 27001 compliance, the recommendation is to start with a fixed-scope pilot on order and shipment status tickets only. Do not attempt to automate refund processing, returns, or policy exceptions in the first sprint. The pilot should measure three baselines before the AI goes live: average first-response time, average handling time, and error rate (wrong order number cited, incorrect shipment status, policy misstatement). After four weeks, compare the post-pilot numbers against the baseline.
The integration sprint should follow this sequence: Week one is the process audit and baseline measurement. Weeks two and three build the LangGraph graph, wire the RAG pipeline to the order and shipment data, and connect the Google Workspace API. Week four is the pilot with the human-in-the-loop gate active. The pilot ships with a documented before/after report on cycle time and error rate.
Two specific recommendations. First, chunk the RAG index at the record level, not the paragraph level. Order data is structured; the LLM needs the full order record to answer accurately. Second, log every AI-drafted response, every retrieval, and every approval decision. ISO 27001 requires documented evidence of information security controls, and the audit log is that evidence. The log should capture the ticket ID, the retrieved records, the model used, the confidence score, and the approver’s identity. This log is also the foundation for the managed operation phase after the pilot.
Leave a Reply