The Problem: Manual Triage Is Your Largest Support Cost
You run a 1,200-person fintech in Zurich. Your support team handles 4,000 tickets a month across chargebacks, onboarding, API errors, and account disputes. Every ticket is read, classified, and routed by a human before a specialist touches it. That first pass takes 90 seconds on average, and it is the single largest cost driver in your support operation. You have heard about AI agents, but your data residency requirements mean you cannot send ticket content to a US-hosted API. You need a triage agent that runs on your own hardware, plugs into your existing helpdesk, and gives you a measured cost-per-ticket reduction in two weeks. This is a fixed-scope pilot: one queue, one routing logic, one baseline report, and a go/no-go decision.
Prerequisites: What You Need Before Day One
Before the pilot starts, you need four things in place. First, access to your helpdesk API (Zendesk, Freshdesk, Jira Service Management, or equivalent) with read and write permissions on the target queue. Second, a Notion or Confluence workspace containing your support knowledge base, with API access for retrieval. Third, a GPU server or a private cloud instance with at least 80 GB of VRAM (an A100 80 GB or two A100 40 GB cards) to serve the open-weight model. Fourth, a 200-ticket sample from the last 90 days, exported with timestamps, categories, and resolution notes, to serve as your baseline dataset. If any of these are missing, the two-week timeline slips. Confirm all four with your IT and support leads before day one.
Step 1: Audit the Triage Workflow and Define the Baseline
Spend the first two days mapping the triage workflow. Export 500 historical tickets from your helpdesk. Tag each one with the category a human assigned, the time from creation to routing, and whether the routing was correct. Build a confusion matrix from this data. This tells you which categories the human team already struggles with, and it becomes the ground truth for evaluating the agent. The deliverable is a one-page process map: ticket arrives, human reads, human classifies, human routes, specialist responds. You are automating the first three steps. The specialist response stays human. This boundary is fixed for the pilot.
Step 2: Deploy the Open-Weight Model On-Premise
Deploy the open-weight model on your GPU server. Use vLLM to serve Llama 3 70B or Mistral 8x7B with a 128k context window. The model receives the ticket text, the category taxonomy from your process map, and a retrieval-augmented context pulled from your Notion or Confluence knowledge base. The prompt instructs the model to output a JSON object: {“category”: “chargeback_dispute”, “priority”: “high”, “route_to”: “chargeback_team”, “confidence”: 0.94}. The confidence score is critical: any ticket below 0.80 is flagged for human review instead of auto-routing. This is your human-in-the-loop gate, and it is non-negotiable for a fintech environment.
Step 3: Wire the Agent to Your Helpdesk via API
Build the orchestration layer that connects the model to your helpdesk. Use a lightweight workflow engine (n8n, Temporal, or a custom Python service) to poll the helpdesk API for new tickets in the target queue. For each ticket, the engine calls the model, parses the JSON output, and writes the classification and routing decision back to the helpdesk via the API. The engine also logs every decision, the confidence score, and the timestamp to a local database. This log is your audit trail and your source for the before/after comparison. The integration is read-write on the helpdesk only; no other system is touched in the pilot.
Step 4: Run Shadow Mode and Measure Accuracy
Run the agent in shadow mode for three days. It processes every new ticket in the target queue, but its routing decision is not applied. A support lead reviews each decision against what a human would have done. You track three metrics: classification accuracy (does the agent pick the right category?), routing accuracy (does it send the ticket to the right team?), and cycle time (how fast does the agent classify versus the human average of 90 seconds). After three days, you have 150-300 shadow decisions. If accuracy is below 90%, you tune the prompt, adjust the retrieval context, or narrow the category taxonomy. You do not move to live routing until accuracy is above 90% on the shadow set.
Step 5: Go Live on One Queue with Human-in-the-Loop
Switch the agent to live routing on the target queue. The human-in-the-loop gate remains: any ticket with a confidence score below 0.80 is routed to a human reviewer instead of auto-routed. For the remaining tickets, the agent’s classification and routing are applied directly in the helpdesk. You monitor the queue for five business days. The support lead reviews a random 20% sample of auto-routed tickets each day to catch drift. If the misclassification rate exceeds 5% on any day, you pause live routing and return to shadow mode. The five-day live window gives you enough data to compute a reliable before/after comparison on cycle time and error rate.
Leave a Reply