Background: A 35-Person Zurich Logistics Firm
This case study is a composite based on patterns observed across multiple engagements. It does not represent a single named client, and no identifying details are disclosed. The scenario reflects recurring operational profiles in the logistics and supply chain sector in Tier-1 European markets.
The company in question is a mid-sized logistics provider based in Zurich, operating 35 employees across operations, customer support, and finance. It manages freight forwarding, last-mile delivery coordination, and customs documentation for B2B clients in DACH and Western Europe. The support team handles approximately 1,200 to 1,800 tickets per month across email, a web portal, and a shared Google Workspace inbox. The stack includes a legacy helpdesk (Zendesk), Google Workspace for email and calendar, and a custom ERP for shipment tracking. The company has no dedicated data science team and had not previously deployed any AI tooling beyond basic keyword filters in the helpdesk.
The Pressure: 1,800 Monthly Tickets and a Q3 Deadline
The support team was the bottleneck. Three senior agents handled the full ticket queue, and each ticket required a human to read, classify, route, and draft a response. The median first-response time was 4.2 hours during business hours and 11 hours for tickets arriving after 17:00 CET. Misrouting to the wrong team occurred in roughly 18 percent of cases, forcing a second handoff and adding 1.5 to 3 hours to resolution. The company was preparing for a 20 percent volume increase tied to a new contract with a retail client, and the operations director had a hard deadline: the support function had to scale without adding headcount before the Q3 peak. GDPR compliance was non-negotiable; the company processes personal data for B2B clients and their end recipients, and the Swiss Federal Act on Data Protection (FADP, revised 2023) applies alongside GDPR for EU-facing operations. The need was specific: free the three senior agents from routine Level-1 triage and drafting so they could focus on escalations, SLA breaches, and client relationship management.
The Approach: Fixed-Scope Pilot on OpenAI with Human-in-the-Loop
The engagement ran as a fixed-scope pilot over 12 weeks, delivered by Forfis as a product studio. The scope was limited to ticket triage and routing: the AI classifies each incoming ticket by category (shipment status, customs query, billing dispute, address correction, other), assigns a priority level, routes it to the correct team, and drafts a first-response reply. The human-in-the-loop rule was explicit: any ticket involving billing, a service-level agreement breach, or personal data in a health or financial context required mandatory human approval before the draft was sent. The AI layer used the OpenAI API (GPT-4o-mini for classification, GPT-4o for drafting) because the ticket volume justified API cost and the multilingual requirement (English and German) was handled natively. The orchestration layer plugged into the existing Zendesk instance via its REST API and into Google Workspace for email-based tickets and calendar scheduling of follow-ups. No new infrastructure was deployed on the client’s side. The pilot included a change-management workshop in week 1 to align the support team on the AI’s role as a drafting and routing assistant, not a replacement.
Outcome: 45 Percent Faster First Response, 9 Percent Misrouting
After 10 weeks of live operation (weeks 3-12), the measured results were as follows. Median first-response time dropped from 4.2 hours to 2.3 hours during business hours and from 11 hours to 5.5 hours for after-hours tickets. The misrouting rate fell from 18 percent to 9 percent. The AI’s triage override rate — the percentage of tickets where a human changed the routing or edited the draft before sending — stabilized at 11 percent after week 6, down from 22 percent in week 3. The three senior agents reported spending roughly 60 percent of their time on escalations and client management rather than Level-1 triage. The company did not add headcount before the Q3 peak. The pilot’s fixed scope meant no feature creep; the client’s request to extend the AI to billing dispute resolution was logged as a separate engagement for Q4. The GDPR compliance review confirmed that the AI’s processing of ticket data met FADP and GDPR requirements, with the record of processing activities updated to reflect the AI’s role.
Lessons for Similar Teams
- Baseline before you build. The 2-4 weeks of historical ticket data with routing labels was the single most valuable input. Without it, the model’s initial accuracy was 71 percent; with it, the starting accuracy was 84 percent. The tuning cycle was shorter and the override rate dropped faster. Teams that skip the baseline measurement cannot prove ROI to their stakeholders.
- Fixed scope is a protection, not a limitation. The client’s instinct to add billing dispute handling during the pilot would have extended the timeline by 4-6 weeks and diluted the pilot’s measurable outcome. The fixed-scope agreement kept the team focused on triage and routing, and the Q4 extension was a natural next step with a clean handover.
- Human-in-the-loop is not a checkbox. The mandatory approval rules for billing and SLA-related tickets were configured in the orchestration layer, not left to agent discretion. This reduced the override rate on high-stakes tickets to under 3 percent and gave the client’s compliance team a clear audit trail.
- Change management is part of the technical delivery. The week-1 workshop with the support team addressed the “will this replace me” concern directly. The agents who engaged with the workshop had a 40 percent lower override rate in the first two weeks than those who did not, suggesting that trust in the tool’s role affects adoption speed.
- Model-agnostic architecture pays off later. The client asked in week 8 whether the system could run on an open-weight model if ticket volume grew and API costs became a concern. Because the orchestration layer was decoupled from the model API, the answer was yes, with a 2-week re-integration. That flexibility was not in the pilot scope, but the architecture made it a non-event.
Leave a Reply