Tag: Free Senior Staff from Routine Work

  • Swiss E-Commerce Firm Cuts Invoice Processing Time 71% with On-Premise AI

    Background: A 300-Person Swiss E-Commerce Firm at Capacity

    This case study is a composite based on patterns observed across Forfis engagements. We do not name real clients. The company described here is a mid-size e-commerce and retail operator based in Zurich, with roughly 300 employees across operations, customer service, and finance. The stack is a mix of a legacy ERP (SAP Business One), a modern CRM (HubSpot), and Slack as the primary internal communication channel. The company had already automated one process — a basic rules-based invoice matching workflow — and was looking to extend AI automation to the next layer of back-office work without adding headcount. The constraint was clear: the finance team was at capacity, and the CTO had a hard deadline to reduce manual data entry before the next fiscal year close.

    Challenge: 12 Hours a Week Lost to Manual Data Entry

    The finance team was spending an estimated 12 hours per week on manual document extraction: pulling supplier invoice fields (vendor name, amount, tax code, line items) from PDFs and entering them into the ERP. The error rate on manual entry was around 8%, and each correction cycle added 45 minutes of rework. The operational pressure was threefold: the fiscal year close was eight weeks away, the team had no budget for additional hires, and the company was in the middle of a PCI DSS re-certification audit, which meant any new system touching payment-related data had to pass a formal risk assessment under Requirement 12.8. The CTO needed a solution that would free senior staff from routine work without introducing a new compliance liability.

    Approach: On-Premise Llama 3 with a Slack Approval Loop

    Forfis ran a two-week process audit that mapped every manual touchpoint in the invoice processing workflow. The audit identified that 70% of the extraction work involved supplier invoices in a consistent PDF format, making them a strong candidate for a fixed-scope pilot. The pilot used an open-weight model (Llama 3 70B) fine-tuned on 500 historical invoice examples, running on the client’s own A100 GPU node inside their VPC. The integration layer connected to Slack: the AI posted extracted fields to a dedicated channel, a human approved or flagged each entry, and approved fields were pushed to the ERP via its REST API. The entire pilot ran in eight weeks, with a measured baseline captured in week one and a shadow run in weeks seven and eight.

    Outcome: 71% Faster Cycle Time, 2.4% Error Rate

    The pilot reduced the average cycle time per invoice from 14 minutes to 4 minutes, a 71% improvement. The field-level error rate dropped from 8% to 2.4%, below the 3% threshold agreed in the pilot scope. The human approval step required intervention on roughly 15% of documents in the first two weeks, tapering to 6% by the end of the shadow run. The finance team reported that the senior staff who had been doing manual entry were now spending that time on supplier negotiations and exception handling. The PCI DSS risk assessment was completed in week six, and the audit trail (every extraction event logged with a document hash) satisfied Requirement 10.2.2 without additional controls.

    Lessons for Teams Scaling AI Without New Hires

    • Baseline before you build. Capturing a 200-document baseline in week one is non-negotiable. Without it, you cannot prove the pilot worked, and the go/no-go decision becomes a gut call. Forfis treats the baseline as a contract: the same sample size, the same measurement method, before and after.
    • Pick the highest-volume, lowest-complexity workflow first. The pilot should target the workflow where the ratio of document volume to format variability is highest. A consistent PDF format with 70% of the volume is a better pilot candidate than a mixed-format pipeline with 30% of the volume.
    • The approval loop is the product, not the model. The Slack channel where a human clicks approve is where the real value lives. The model is a swappable component; the approval workflow is what the team actually uses every day.
    • PCI DSS compliance is a design constraint, not an afterthought. The on-premise architecture and the audit trail were built in from day one, not bolted on after the pilot. Requirement 12.8 risk assessment and Requirement 10.2.2 logging were part of the pilot scope, not a separate workstream.
  • Six Ways Forfis Automates Ticket Triage for Swiss Fintechs in a 3-Month Sprint

    1. Triage eats senior hours that should go to disputes

    A 2,000-employee payments firm in Zurich runs Zendesk as its primary helpdesk. Senior agents spend 40% of their day re-routing misclassified tickets and drafting first responses that follow the same template every time. The process audit identifies ticket triage and routing as the highest-impact workflow: 12,000 tickets per month, a median cycle time of 4.2 hours from receipt to first response, and a 14% error rate on routing. The AI layer classifies by intent, urgency, and department, then routes to the correct queue. After the pilot, median cycle time drops to 38 minutes and routing errors fall to 2.1%. The senior agents who previously handled triage now focus on complex disputes and fraud escalations, work that actually requires their judgment. The 3-month sprint covers audit, pilot, and rollout, with every decision logged for ISO 27001 audit trails.

    2. Workflow orchestration, not a chatbot wrapper

    The orchestration layer sits between Zendesk’s API and the model inference endpoint. Incoming tickets trigger a webhook that passes the ticket body, metadata, and customer history to the classifier. The model returns a structured JSON object with intent, urgency score, and recommended queue. The orchestrator validates the output against a schema, checks confidence thresholds, and routes the ticket accordingly. If confidence falls below 0.85, the ticket flags for human review. Every step logs a timestamp, model version, and input hash. This architecture means the client can swap the classifier model without touching the Zendesk integration or the routing logic. The orchestration layer is the stable contract; the model is a pluggable component.

    3. On-premise open-weight models keep regulated data local

    Swiss data protection law and the client’s ISO 27001 certification require that customer payment data never leaves the building. Forfis deploys an open-weight model on the client’s own GPU cluster, handling all ticket payloads that contain account numbers, transaction IDs, or personal identifiers. The model runs on-premise, so no regulated data crosses a network boundary. For non-sensitive workflows, such as routing a general FAQ ticket, the orchestrator can route to a cloud API where latency and cost are less critical. The model-agnostic design means the client chooses the model per workflow, not per project. This split keeps the ISO 27001 statement of applicability clean: the on-premise path satisfies Annex A.8.22 (use of cryptography) and A.8.15 (access control) without requiring a separate risk assessment for cloud data transfer.

    4. A 3-month sprint with a measured before/after baseline

    The pilot runs on one workflow for six weeks. The baseline is measured in the first two weeks: 12,000 tickets, 4.2-hour median cycle time, 14% routing error rate. The AI layer goes live in week three, handling triage and routing with human approval on any ticket flagged below the confidence threshold. By week six, the metrics show a 38-minute median cycle time and a 2.1% error rate. The before/after comparison is documented in a one-page report that the client’s CFO uses to justify the rollout budget. The pilot also surfaces edge cases: 3% of tickets contain multilingual content that the model misclassifies, prompting a fine-tuning pass before full rollout. This measured approach means the client sees ROI before committing to broader automation across invoice processing or document extraction.

    5. Human-in-the-loop by default, not as an afterthought

    The AI layer classifies and routes, but a human approves any action that touches money, health data, or a contract. In a payments context, this means the model drafts a refund response or flags a fraud-related ticket, but a senior agent signs off before the action executes. The approval threshold is not arbitrary; it is set based on the pilot’s error-rate data. If the model’s routing accuracy on fraud-related tickets is 97%, the human-in-the-loop threshold applies to that subset. For general FAQ tickets where accuracy is 99.5%, the system can auto-route without approval. This tiered approach frees senior staff from routine work while keeping them in the loop for high-stakes decisions. The approval log feeds directly into the ISO 27001 audit trail, showing who approved what and when.

    6. Plugs into Zendesk or Intercom without replacing them

    The integration sprint adds a new processing layer on top of the existing Zendesk or Intercom instance. No data migration is required; the AI layer reads tickets through the helpdesk’s native API and writes routing decisions back through the same API. The client’s existing workflows, SLAs, and reporting dashboards continue to function unchanged. The orchestration layer exposes a REST API that the helpdesk calls, so the integration is a few lines of configuration in Zendesk’s webhook settings. This means the client does not need to retrain agents on a new interface or rebuild their ticket taxonomy. The AI layer is invisible to the end customer; it simply makes the existing system faster and more accurate. The 3-month timeline includes two weeks of integration hardening after the pilot, where edge cases from the pilot are addressed and the system is stress-tested under production load.