{"id":69,"date":"2026-10-06T18:59:34","date_gmt":"2026-10-06T18:59:34","guid":{"rendered":"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/"},"modified":"2026-10-06T18:59:34","modified_gmt":"2026-10-06T18:59:34","slug":"ai-ticket-triage-fintech-switzerland-pilot","status":"publish","type":"post","link":"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/","title":{"rendered":"AI Ticket Triage for a Swiss Fintech: A Two-Week On-Premise Pilot"},"content":{"rendered":"<h2>The Problem: Manual Triage Is Your Largest Support Cost<\/h2>\n<p>You run a 1,200-person fintech in Zurich. Your support team handles 4,000 tickets a month across chargebacks, onboarding, API errors, and account disputes. Every ticket is read, classified, and routed by a human before a specialist touches it. That first pass takes 90 seconds on average, and it is the single largest cost driver in your support operation. You have heard about AI agents, but your data residency requirements mean you cannot send ticket content to a US-hosted API. You need a triage agent that runs on your own hardware, plugs into your existing helpdesk, and gives you a measured cost-per-ticket reduction in two weeks. This is a fixed-scope pilot: one queue, one routing logic, one baseline report, and a go\/no-go decision.<\/p>\n<h2>Prerequisites: What You Need Before Day One<\/h2>\n<p>Before the pilot starts, you need four things in place. First, access to your helpdesk API (Zendesk, Freshdesk, Jira Service Management, or equivalent) with read and write permissions on the target queue. Second, a Notion or Confluence workspace containing your support knowledge base, with API access for retrieval. Third, a GPU server or a private cloud instance with at least 80 GB of VRAM (an A100 80 GB or two A100 40 GB cards) to serve the open-weight model. Fourth, a 200-ticket sample from the last 90 days, exported with timestamps, categories, and resolution notes, to serve as your baseline dataset. If any of these are missing, the two-week timeline slips. Confirm all four with your IT and support leads before day one.<\/p>\n<h2>Step 1: Audit the Triage Workflow and Define the Baseline<\/h2>\n<p>Spend the first two days mapping the triage workflow. Export 500 historical tickets from your helpdesk. Tag each one with the category a human assigned, the time from creation to routing, and whether the routing was correct. Build a confusion matrix from this data. This tells you which categories the human team already struggles with, and it becomes the ground truth for evaluating the agent. The deliverable is a one-page process map: ticket arrives, human reads, human classifies, human routes, specialist responds. You are automating the first three steps. The specialist response stays human. This boundary is fixed for the pilot.<\/p>\n<h2>Step 2: Deploy the Open-Weight Model On-Premise<\/h2>\n<p>Deploy the open-weight model on your GPU server. Use vLLM to serve Llama 3 70B or Mistral 8x7B with a 128k context window. The model receives the ticket text, the category taxonomy from your process map, and a retrieval-augmented context pulled from your Notion or Confluence knowledge base. The prompt instructs the model to output a JSON object: {\u201ccategory\u201d: \u201cchargeback_dispute\u201d, \u201cpriority\u201d: \u201chigh\u201d, \u201croute_to\u201d: \u201cchargeback_team\u201d, \u201cconfidence\u201d: 0.94}. The confidence score is critical: any ticket below 0.80 is flagged for human review instead of auto-routing. This is your human-in-the-loop gate, and it is non-negotiable for a fintech environment.<\/p>\n<h2>Step 3: Wire the Agent to Your Helpdesk via API<\/h2>\n<p>Build the orchestration layer that connects the model to your helpdesk. Use a lightweight workflow engine (n8n, Temporal, or a custom Python service) to poll the helpdesk API for new tickets in the target queue. For each ticket, the engine calls the model, parses the JSON output, and writes the classification and routing decision back to the helpdesk via the API. The engine also logs every decision, the confidence score, and the timestamp to a local database. This log is your audit trail and your source for the before\/after comparison. The integration is read-write on the helpdesk only; no other system is touched in the pilot.<\/p>\n<h2>Step 4: Run Shadow Mode and Measure Accuracy<\/h2>\n<p>Run the agent in shadow mode for three days. It processes every new ticket in the target queue, but its routing decision is not applied. A support lead reviews each decision against what a human would have done. You track three metrics: classification accuracy (does the agent pick the right category?), routing accuracy (does it send the ticket to the right team?), and cycle time (how fast does the agent classify versus the human average of 90 seconds). After three days, you have 150-300 shadow decisions. If accuracy is below 90%, you tune the prompt, adjust the retrieval context, or narrow the category taxonomy. You do not move to live routing until accuracy is above 90% on the shadow set.<\/p>\n<h2>Step 5: Go Live on One Queue with Human-in-the-Loop<\/h2>\n<p>Switch the agent to live routing on the target queue. The human-in-the-loop gate remains: any ticket with a confidence score below 0.80 is routed to a human reviewer instead of auto-routed. For the remaining tickets, the agent\u2019s classification and routing are applied directly in the helpdesk. You monitor the queue for five business days. The support lead reviews a random 20% sample of auto-routed tickets each day to catch drift. If the misclassification rate exceeds 5% on any day, you pause live routing and return to shadow mode. The five-day live window gives you enough data to compute a reliable before\/after comparison on cycle time and error rate.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A two-week fixed-scope pilot that deploys an on-premise open-weight AI agent for ticket triage in a Swiss fintech, with measured before\/after baselines on cycle time and error rate.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"AI Ticket Triage for a Swiss Fintech: A Two-Week On-Premise Pilot","rank_math_description":"A two-week fixed-scope pilot that deploys an on-premise open-weight AI agent for ticket triage in a Swiss fintech, with measured before\/after baselines on cycle time and error rate.","rank_math_focus_keyword":"automate monthly reporting ticket triage and routing","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","_yoast_wpseo_focuskw":"","pll_lang":"en","geo_jsonld":"{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/#article\",\"@type\":\"Article\",\"author\":{\"@id\":\"https:\/\/blog.forfis.com#org\"},\"dateModified\":\"2026-10-05T23:45:37.277772055+00:00\",\"datePublished\":\"2026-10-05T23:45:37.277772055+00:00\",\"description\":\"A two-week fixed-scope pilot that deploys an on-premise open-weight AI agent for ticket triage in a Swiss fintech, with measured before\/after baselines on cycle time and error rate.\",\"headline\":\"AI Ticket Triage for a Swiss Fintech: A Two-Week On-Premise Pilot\",\"inLanguage\":\"en\",\"keywords\":[\"Running Isolated Pilots\",\"Open-Weight Models On-Premise\",\"Workflow Orchestration\",\"Customer Support\",\"501-2000\",\"None\",\"Fixed-Scope Pilot\",\"Fintech and Payments\",\"Notion or Confluence\",\"English\",\"Automate Monthly Reporting\",\"Switzerland\",\"2 weeks\",\"Ticket Triage and Routing\"],\"mainEntityOfPage\":\"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/\",\"publisher\":{\"@id\":\"https:\/\/blog.forfis.com#org\"}},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/#faq\",\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"In a fixed-scope pilot, the deliverable is a working triage agent on one queue, a measured before\/after baseline, and a go\/no-go recommendation. The scope is locked in week one: which queue, which ticket categories, which integration points, and which success metrics. If the pilot surfaces a need to automate a second queue or add a voice channel, that is a new engagement, not a scope change. This keeps the two-week timeline honest and the cost predictable.\"},\"name\":\"What does a fixed-scope pilot actually deliver in two weeks?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent reads the ticket, classifies it into a predefined category (e.g., 'chargeback dispute', 'onboarding question', 'API error'), assigns a priority, and routes it to the correct queue or agent in your helpdesk. It does not draft a customer-facing reply in the pilot phase. The human-in-the-loop approval step means a support lead reviews the classification before the ticket moves. This keeps the risk low while you validate accuracy against your historical data.\"},\"name\":\"What does the AI agent actually do in the ticket triage workflow?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes. The pilot uses your existing helpdesk (Zendesk, Freshdesk, Jira Service Management, etc.) via its API. The agent reads new tickets, writes back the classification and routing decision, and logs its confidence score. You do not replace the helpdesk, the CRM, or the ERP. The Notion or Confluence integration is for the knowledge base the agent retrieves from, not for replacing your documentation system.\"},\"name\":\"Do we need to replace our current helpdesk or CRM?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot runs on your own hardware or a private cloud instance you control. The model is an open-weight LLM (e.g., Llama 3 70B or Mistral 8x7B) served via vLLM or TGI. No ticket content, customer PII, or transaction data leaves your infrastructure. The only external calls are to your helpdesk API and your Notion\/Confluence API, both over TLS. If your data residency policy requires it, the entire stack runs inside your Swiss data center.\"},\"name\":\"Where does the model run, and does any data leave our infrastructure?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot includes a baseline measurement: you export 200-500 historical tickets with their actual resolution time and error rate. The agent processes the same set, and you compare cycle time (time from ticket creation to correct routing) and misclassification rate. The deliverable is a one-page report with these two numbers, side by side. This is the evidence you use to decide on rollout and to model the cost-per-ticket reduction.\"},\"name\":\"How do we measure the before\/after baseline in two weeks?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot is a fixed-scope engagement with a defined deliverable and a go\/no-go decision at the end. It is not a subscription. If you proceed to rollout, that is a separate engagement with its own scope, timeline, and pricing. The pilot cost covers the process audit, the agent build, the integration, the baseline measurement, and the recommendation. You pay for the outcome, not for ongoing model hosting during the pilot.\"},\"name\":\"Is the pilot a one-time cost or a subscription?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The agent is model-agnostic. The pilot uses an open-weight model on your hardware because it is the fastest path to a working system in two weeks and keeps data local. If your accuracy requirements exceed what the open-weight model achieves, the rollout can swap in a commercial API (OpenAI, Anthropic) for the classification step while keeping the orchestration layer unchanged. The architecture is designed so the model is a swappable component, not a hard dependency.\"},\"name\":\"Can we switch to a commercial model like GPT-4 after the pilot?\"},{\"@type\":\"Question\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The pilot covers one queue and one routing logic. Monthly reporting automation is a separate workflow that can be built in a follow-up engagement. The process audit done during the pilot will identify which reporting tasks are candidates for automation, and the pilot's baseline data gives you a starting point for scoping that work. You do not need to bundle it into the two-week pilot.\"},\"name\":\"Can the pilot also automate our monthly reporting?\"}]},{\"@id\":\"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/#breadcrumbs\",\"@type\":\"BreadcrumbList\",\"itemListElement\":[{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\",\"name\":\"Home\",\"position\":1},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/\",\"name\":\"Blog\",\"position\":2},{\"@type\":\"ListItem\",\"item\":\"https:\/\/blog.forfis.com\/blog\/ai-ticket-triage-fintech-switzerland-pilot\/\",\"name\":\"AI Ticket Triage for a Swiss Fintech: A Two-Week On-Premise Pilot\",\"position\":3}]},{\"@id\":\"https:\/\/blog.forfis.com#org\",\"@type\":\"Organization\",\"name\":\"Forfis\",\"url\":\"https:\/\/blog.forfis.com\"}]}","geo_content_hash":"4b6b2770649ade437c574bb7a6627cfb353abcbf76f9cee5bd8aa5f2c7842e2c","footnotes":""},"categories":[37],"tags":[69,43,51],"class_list":["post-69","post","type-post","status-publish","format-standard","hentry","category-fintech-and-payments","tag-automate-monthly-reporting","tag-switzerland","tag-ticket-triage-and-routing"],"_links":{"self":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/69","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/comments?post=69"}],"version-history":[{"count":0,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/posts\/69\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/media?parent=69"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/categories?post=69"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.forfis.com\/blog\/wp-json\/wp\/v2\/tags?post=69"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}