Blog

  • Cutting First-Response Time in B2B SaaS Support with a RAG Assistant in Austria

    The Support Team Is Drowning in Status Queries

    The support team at a 120-person B2B SaaS company in Vienna handles 400 to 600 customer queries per week. The majority are order and shipment status updates: “Where is my order?” “When will the shipment arrive?” “Why is my invoice late?” Each query requires the agent to log into the CRM, pull the order record, check the ERP for shipment status, and draft a response. The average first-response time is 6 hours for email and 22 minutes for chat. The team of eight support agents is stretched thin, and the company has no budget to hire more. The pain is not a lack of tools; it is a lack of time. The agents are not unskilled; they are under-resourced. The company needs to scale operations without adding headcount, and the constraint is GDPR: customer data cannot be sent to a US-based API provider without a data processing agreement and a transfer impact assessment.

    Why Off-the-Shelf Chatbots and More Headcount Fail

    The first instinct is to buy a chatbot. Most B2B SaaS companies have tried this. The chatbot handles simple queries but fails on anything that requires cross-referencing the CRM and the ERP. It gives generic answers, and the customer escalates to a human agent, who has to redo the work. The second instinct is to hire more support agents. This works until the volume grows again, and the cost per query rises. The third instinct is to build an internal tool. This takes six to nine months, and the team that builds it is the same team that is supposed to handle the queries. None of these approaches address the root cause: the agents are spending 70% of their time on repetitive, data-retrieval tasks that a machine can do in seconds. The failure mode is not technology; it is a mismatch between the tool and the workflow. The tool must retrieve data from the CRM and ERP, draft a response, and hand it to a human for approval. That is a retrieval-augmented generation task, not a chatbot task.

    A RAG Assistant on the Company’s Own Infrastructure

    The solution is a retrieval-augmented knowledge assistant that plugs into the systems the company already runs. The assistant is deployed on the client’s own hardware using an open-weight model, so customer data never leaves the building. It integrates with the CRM, the ERP, and the helpdesk through their APIs. When a customer query arrives in Slack or Microsoft Teams, the assistant retrieves the relevant order and shipment data, drafts a response, and posts it to the support channel with a flag for human review. The agent approves, edits, or rejects the draft. The approved response is sent to the customer. The entire flow takes under 5 minutes. The architecture is model-agnostic: the open-weight model handles the retrieval and drafting, and if a query requires complex reasoning, the system can escalate to a cloud API provider under a data processing agreement. The pilot is fixed-scope: 8 weeks, one workflow, measured before/after baseline on first-response time and error rate.

    How to Start: Five Concrete Steps in Eight Weeks

    The first step is the process audit. The audit maps the current support workflow: how queries arrive, how they are triaged, which systems the agent accesses, how long each step takes, and where errors occur. The audit identifies the workflows worth automating, prioritized by volume, cycle time, and error rate. For a B2B SaaS company, the highest-impact workflow is order and shipment status updates. The audit takes 1 to 2 weeks and is delivered as a report with a prioritized roadmap. The second step is the fixed-scope pilot. The pilot covers one workflow, integrates with two to three existing systems, deploys the RAG assistant on the client’s infrastructure, and ships with a measured before/after baseline. The third step is the human-in-the-loop approval layer. The model drafts, the human approves. The fourth step is the integration with Slack or Microsoft Teams. The assistant appears as a bot in the support channels. The fifth step is the decision document. At week 8, the client receives the measured metrics, a rollout plan, and a cost model for managed operation.

    Pitfalls That Derail the Pilot

    The most common pitfall is skipping the process audit. The company jumps straight to building the assistant and discovers that the CRM data is incomplete, the ERP fields are mislabeled, and the helpdesk articles are outdated. The assistant retrieves the wrong data, and the human approver has to fix it every time. The second pitfall is underestimating the human-in-the-loop layer. The company assumes that the model will be accurate enough to skip the approval step, and the first batch of automated responses contains errors that damage customer trust. The third pitfall is choosing a cloud API provider without a data processing agreement. The company discovers during the GDPR review that customer data is being sent to a US server, and the project is paused for three weeks while the legal team negotiates the agreement. The fourth pitfall is treating the pilot as a one-off project. The company does not plan for the rollout, and the assistant is never scaled beyond the pilot workflow. The lesson is that the pilot is not the product; it is the proof of concept that unlocks the rollout.

  • 4-Week AI Pilot: Invoice Processing and RAG Assistant for a B2B SaaS in Austria

    The Problem: Manual Back-Office Work and Slow First-Response in a 201-500 Employee B2B SaaS

    Your operations and supply chain team in Vienna processes 1,200 invoices monthly, each taking 14 minutes of manual data entry, and your support desk answers 300 tickets a week with a median first-response time of 4.2 hours. The back-office work is repetitive, error-prone, and consuming 3.5 FTEs that could be redeployed. The EU AI Act, in force since August 2024, requires you to document your AI risk assessment before deploying any automated system that touches financial data. You need a fixed-scope pilot that delivers a measured before/after baseline in 4 weeks, not a 6-month transformation program. The pilot must work within your existing stack — Notion for documentation, your CRM for customer records, your ERP for invoice data — and must keep regulated data on Austrian infrastructure.

    Prerequisites Before Step 1

    • Process audit completed: You have mapped the invoice processing workflow from receipt to payment, timed each step, and counted error types. The audit output is a one-page document with baseline metrics: average cycle time (hours), error rate (%), and FTE hours consumed.
    • n8n instance deployed: A self-hosted n8n instance runs on your Austrian cloud or on-premises server. You have API credentials for your CRM, ERP, and helpdesk. The n8n version is 1.40 or later for stable webhook and AI node support.
    • RAG source material ready: Notion or Confluence contains at least 50 pages of operational documentation — vendor onboarding, invoice coding rules, escalation paths, SLA definitions. The content is current (updated within the last 30 days).
    • Human-in-the-loop approvers identified: You have named 2–3 people who will approve AI-drafted invoice entries and ticket responses. They understand the approval criteria and have access to the n8n approval UI.
    • EU AI Act risk assessment drafted: A one-page document classifying your RAG assistant as a limited-risk system, noting the transparency obligations, and confirming no special-category data is processed without consent.
    • Fixed-scope statement of work signed: The pilot scope, success metrics, and 4-week timeline are locked. No scope changes without a change order.

    Step 1: Run the Process Audit and Lock the Baseline

    Run a 2-hour process audit with your operations lead. Map every step from invoice receipt (email, portal, or EDI) to payment posting in the ERP. Time each step with a stopwatch or screen-recording tool. Count error types over the last 30 days: wrong vendor code, duplicate entry, missing tax ID, incorrect tax rate. Record the baseline: average cycle time in hours, error rate as a percentage, and total FTE hours consumed. Output: a one-page audit document with a workflow diagram and a table of error types with frequencies. This document is your before/after measurement anchor. Do not proceed to Step 2 until the baseline is signed off by the operations lead.

    Step 2: Build the n8n Invoice Extraction Workflow

    Build the n8n workflow for invoice extraction. Create a webhook node that receives the invoice PDF via email or ERP API. Add an AI node using OpenAI’s GPT-4o or Anthropic’s Claude 3.5 Sonnet for extraction — these models handle multi-column invoice layouts with 94–97% field accuracy on standard B2B invoices. Configure the extraction schema: vendor name, vendor tax ID, invoice number, line items, tax rate, total amount, due date. Add a validation node that checks for missing fields and flags anomalies (e.g., tax ID format mismatch, total exceeds PO amount by more than 5%). Route flagged invoices to a human approval node in n8n; route clean invoices to the ERP write-back node. Test with 50 historical invoices before going live.

    Step 3: Build the RAG Knowledge Assistant Over Notion or Confluence

    Set up the RAG index over your Notion or Confluence documentation. In n8n, create a workflow that pulls pages on an hourly schedule using the Notion API node or Confluence Cloud API. Chunk the content at 512 tokens with 64-token overlap. Embed using BGE-M3 or Cohere embed-v3 — both handle English and German, which matters for your Austrian team. Store embeddings in pgvector on your PostgreSQL instance. Build the RAG query workflow: receive a ticket or question, retrieve the top-5 chunks, pass them as context to the LLM, and return a grounded answer with source citations (page title and URL). Test with 20 real questions from your support team. If retrieval hit-rate is below 85%, re-chunk or re-embed. The RAG assistant must never answer without a source citation.

    Step 4: Integrate with CRM, ERP, and Helpdesk

    Integrate the n8n workflows with your existing systems. For the invoice workflow: connect the ERP write-back node to your ERP’s API (SAP, NetSuite, or similar) using the vendor’s REST or SOAP endpoint. For the RAG assistant: connect the helpdesk (Zendesk, Freshdesk, or Jira Service Management) via webhook so that incoming tickets trigger the RAG query workflow. The RAG workflow drafts a response, attaches the retrieved context, and routes it to the human approver. The approver edits or approves in the n8n UI, and the approved response sends via the helpdesk API. All integrations use your existing API credentials — no new accounts, no new systems. Test each integration with 10 real transactions in a staging environment before moving to production.

    Step 5: Run the 4-Week Pilot with Human-in-the-Loop Approval

    Run the pilot in shadow mode for 2 weeks. The n8n workflows process real invoices and tickets, but the human approver reviews every output before it reaches the ERP or the customer. Track three metrics daily: cycle time (from invoice receipt to ERP posting, or from ticket creation to first response), error rate (AI-drafted entries rejected or edited by the approver), and human-override rate (percentage of AI outputs that required manual correction). At the end of 2 weeks, compare against the Step 1 baseline. The pilot report must show: cycle time reduction in hours, error rate change in percentage points, and FTE hours saved. If cycle time drops by 40% or more and error rate stays below 5%, the pilot is a success. If not, diagnose the failure mode before proceeding to rollout.

  • Swiss E-commerce Cuts Invoice Cycle Time 92% in a Two-Week ISO 27001-Safe Pilot

    Background: A Swiss Retail Group Under Audit Pressure

    This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The details are plausible and reflect the range of outcomes seen in the field, but they do not describe a single real company.

    The client is a Swiss e-commerce and retail group with roughly 2,400 employees, operating in German, French, and Italian markets. The finance and accounting team handles 18,000 to 22,000 supplier invoices per month across three ERP instances. The stack is a mix of SAP S/4HANA for the core ledger, a legacy document management system for invoice images, and Confluence for internal runbooks and audit documentation. The company holds ISO 27001 certification and is in the middle of a renewal audit. The finance director’s mandate was clear: reduce the average cycle time from invoice receipt to ERP posting without introducing a compliance gap.

    Challenge: 20,000 Invoices a Month and a 90-Day Audit Clock

    The finance team was processing invoices manually: a clerk downloaded the PDF, typed the vendor name, amount, tax code, and cost center into the ERP, and flagged discrepancies for review. The average cycle time was 4 to 6 hours per invoice, with a 3 to 5 percent error rate on a sample of 500 invoices. The error rate was not just a cost issue; it was a compliance issue. ISO 27001 requires documented controls over financial data, and a 4 percent error rate on 20,000 invoices per month meant roughly 800 mis-posted entries that had to be caught in a secondary review. The secondary review was itself a manual process, adding another 2 to 3 hours per flagged invoice. The finance director had a deadline: the ISO 27001 renewal audit was 90 days out, and the auditor had already flagged the manual process as a control weakness.

    Approach: A Two-Week Pilot on the Top Five Vendors

    The engagement started with a three-day process audit. The team mapped the invoice lifecycle from receipt to posting, identified the 12 vendor categories that accounted for 78 percent of volume, and pulled a historical sample of 1,200 invoices for calibration. The pilot scope was fixed: one ERP instance, one vendor category (the top 5 suppliers by volume), and a two-week window. The architecture used the OpenAI API for extraction, with a human-in-the-loop approval queue. The model extracted vendor name, invoice number, amount, tax code, and cost center. A reviewer saw the proposed entry alongside the original PDF and could approve, correct, or reject. The approval log was written to Confluence and to the ERP audit trail. The pipeline connected to the ERP via its REST API and to the document store via SFTP. No new infrastructure was required. The client’s existing IT team handled the API credentials and network access.

    Outcome: 92 Percent Cycle-Time Reduction in 12 Days

    The pilot ran for 12 business days. The model processed 1,840 invoices from the top five vendors. The average cycle time dropped from 4.2 hours to 22 minutes, a 92 percent reduction. The error rate on the pilot sample was 0.8 percent, down from the 3.4 percent baseline. Of the 1,840 invoices, 1,612 were approved with zero edits. The remaining 228 required human correction, mostly on tax codes for cross-border invoices. The approval queue averaged 14 minutes per invoice for the corrected entries. The ISO 27001 audit trail showed 100 percent of inferences logged with timestamp, user ID, and confidence score. The finance director presented the pilot results to the audit committee. The auditor accepted the AI-assisted workflow as a control improvement, conditional on the managed operations SLA being in place before the renewal audit.

    Lessons for Teams Running Similar Pilots

    • The historical sample matters more than the model. The 1,200-invoice calibration sample was the single biggest factor in the 0.8 percent error rate. A team that skips this step and goes live with a generic prompt will see error rates of 8 to 12 percent and lose the human trust needed for the approval workflow.
    • Fix the scope before you start. The two-week window only worked because the pilot was limited to one ERP instance and five vendors. A team that tries to cover all 12 vendor categories in two weeks will spend the time on integration edge cases and miss the baseline measurement.
    • The approval queue is the product, not the model. The model’s extraction quality was good, but the reviewer interface was what made the workflow usable. A team that ships a model without a clean approval UI will see reviewers bypass the system and go back to manual entry.
    • ISO 27001 is a design constraint, not a post-hoc checkbox. The audit trail, the data processing agreement, and the access controls were built into the architecture from day one. Retrofitting them after go-live is 3 to 4 times more expensive and often fails the audit.
    • Managed operations is where the value compounds. The pilot proved the concept. The managed operations SLA, with monthly reports on confidence distribution and error rate, is what keeps the error rate at 0.8 percent instead of drifting to 3 percent as vendor formats change.
  • Swiss E-Commerce Firm Cuts Invoice Processing Time 71% with On-Premise AI

    Background: A 300-Person Swiss E-Commerce Firm at Capacity

    This case study is a composite based on patterns observed across Forfis engagements. We do not name real clients. The company described here is a mid-size e-commerce and retail operator based in Zurich, with roughly 300 employees across operations, customer service, and finance. The stack is a mix of a legacy ERP (SAP Business One), a modern CRM (HubSpot), and Slack as the primary internal communication channel. The company had already automated one process — a basic rules-based invoice matching workflow — and was looking to extend AI automation to the next layer of back-office work without adding headcount. The constraint was clear: the finance team was at capacity, and the CTO had a hard deadline to reduce manual data entry before the next fiscal year close.

    Challenge: 12 Hours a Week Lost to Manual Data Entry

    The finance team was spending an estimated 12 hours per week on manual document extraction: pulling supplier invoice fields (vendor name, amount, tax code, line items) from PDFs and entering them into the ERP. The error rate on manual entry was around 8%, and each correction cycle added 45 minutes of rework. The operational pressure was threefold: the fiscal year close was eight weeks away, the team had no budget for additional hires, and the company was in the middle of a PCI DSS re-certification audit, which meant any new system touching payment-related data had to pass a formal risk assessment under Requirement 12.8. The CTO needed a solution that would free senior staff from routine work without introducing a new compliance liability.

    Approach: On-Premise Llama 3 with a Slack Approval Loop

    Forfis ran a two-week process audit that mapped every manual touchpoint in the invoice processing workflow. The audit identified that 70% of the extraction work involved supplier invoices in a consistent PDF format, making them a strong candidate for a fixed-scope pilot. The pilot used an open-weight model (Llama 3 70B) fine-tuned on 500 historical invoice examples, running on the client’s own A100 GPU node inside their VPC. The integration layer connected to Slack: the AI posted extracted fields to a dedicated channel, a human approved or flagged each entry, and approved fields were pushed to the ERP via its REST API. The entire pilot ran in eight weeks, with a measured baseline captured in week one and a shadow run in weeks seven and eight.

    Outcome: 71% Faster Cycle Time, 2.4% Error Rate

    The pilot reduced the average cycle time per invoice from 14 minutes to 4 minutes, a 71% improvement. The field-level error rate dropped from 8% to 2.4%, below the 3% threshold agreed in the pilot scope. The human approval step required intervention on roughly 15% of documents in the first two weeks, tapering to 6% by the end of the shadow run. The finance team reported that the senior staff who had been doing manual entry were now spending that time on supplier negotiations and exception handling. The PCI DSS risk assessment was completed in week six, and the audit trail (every extraction event logged with a document hash) satisfied Requirement 10.2.2 without additional controls.

    Lessons for Teams Scaling AI Without New Hires

    • Baseline before you build. Capturing a 200-document baseline in week one is non-negotiable. Without it, you cannot prove the pilot worked, and the go/no-go decision becomes a gut call. Forfis treats the baseline as a contract: the same sample size, the same measurement method, before and after.
    • Pick the highest-volume, lowest-complexity workflow first. The pilot should target the workflow where the ratio of document volume to format variability is highest. A consistent PDF format with 70% of the volume is a better pilot candidate than a mixed-format pipeline with 30% of the volume.
    • The approval loop is the product, not the model. The Slack channel where a human clicks approve is where the real value lives. The model is a swappable component; the approval workflow is what the team actually uses every day.
    • PCI DSS compliance is a design constraint, not an afterthought. The on-premise architecture and the audit trail were built in from day one, not bolted on after the pilot. Requirement 12.8 risk assessment and Requirement 10.2.2 logging were part of the pilot scope, not a separate workstream.
  • 8-Week AI Automation Audit: Cutting First-Response Time in a UAE Medtech Firm

    1. Map the ticket flow before touching the model

    The audit phase is where most 11-50 person firms stall. Forfis starts by mapping every ticket that hits the support queue over a 10-day window, tagging each by topic, resolution path, and time-to-first-response. For a UAE medtech company, the data typically shows 60-70% of tickets are “where is the protocol for X” or “what is the warranty window for Y” questions that live in Confluence or Notion but are buried under 200+ pages. The audit output is a ranked list of the top five question categories by volume and time cost, with a measured baseline: average first-response time of 4.2 hours, error rate of 12% on a 200-ticket sample. This baseline is the number the pilot must beat, and it is documented in a one-page report the team signs off on before any code is written.

    2. Build the RAG layer on LangGraph, not a monolith

    The RAG pipeline indexes Confluence and Notion pages into a vector store, chunking at 512 tokens with 64-token overlap. LangGraph orchestrates the retrieval, generation, and scoring nodes. The predictive scoring module evaluates each draft on three axes: retrieval relevance (cosine similarity of the top-3 chunks), answer coherence (a secondary LLM call that checks the draft against the retrieved context), and historical approval rate (a running average from the pilot’s first 50 tickets). Responses scoring below 0.85 route to a human; those above auto-post to the helpdesk. For a 15-person team, this means the AI handles roughly 75% of tickets, and the human agent reviews the remaining 25% in under 5 minutes each. The scoring threshold is tunable in the LangGraph config without redeploying.

    3. Run the pilot with a measured before/after baseline

    The pilot runs for two weeks on a live subset of tickets. The team uses the agent in production, and every interaction is logged: the ticket ID, the retrieved chunks, the draft answer, the predictive score, and whether the human approved, edited, or rejected it. By the end of the soak period, the team has a 200-ticket dataset with before/after metrics. For a UAE medtech firm, the typical result is first-response time dropping from 4.2 hours to 18 minutes, with error rate holding at 11% or below. The 8-week timeline includes a one-week buffer for model tuning if the initial scoring threshold is too aggressive or too conservative. The final deliverable is a one-page baseline report with the numbers, the model used, the cost per 1,000 tokens, and a recommendation on whether to scale to all ticket categories or adjust the scope.

    4. Keep the model layer swappable from day one

    The architecture calls the LLM through an abstraction layer in LangChain, so the model is a config parameter, not a hard dependency. For a UAE healthcare firm with no compliance mandate, starting with OpenAI’s GPT-4o API is the fastest path: no hardware procurement, no MLOps overhead. The audit phase documents the cost per 1,000 tokens (typically $0.03-0.06 for GPT-4o) and the latency (18-25 ms for a 512-token response). If the team later decides to move to an open-weight model like Llama 3.1 70B on their own hardware, the LangGraph nodes do not change. The swap is a one-line config update. This matters for a 15-person team because it removes the risk of being locked into a single vendor’s pricing or API changes mid-engagement.

    5. Plug into the helpdesk, not around it

    The agent does not replace the helpdesk. It plugs into the existing ticketing system via API. When a ticket arrives, the agent retrieves relevant chunks, drafts a response, and posts it as a suggested reply in the ticket. The human agent sees the draft, approves or edits it, and sends it. The agent logs the retrieval context and the predictive score in the ticket metadata, so the team can audit why a particular answer was suggested. For a 15-person team, this means no new UI to learn, no workflow redesign, and no training beyond a 30-minute onboarding session. The agent operates inside the tools the team already uses, which is critical for adoption in a small firm where every hour of context-switching is expensive.

    6. Plan for the knowledge base to change

    The most common failure mode is treating the pilot as a one-time deliverable. For a 15-person UAE medtech firm, the knowledge base changes weekly: new protocols, updated warranty terms, revised SOPs. The RAG pipeline must re-index Confluence and Notion on a schedule (daily or on webhook trigger) to keep the chunks current. The predictive scoring model also drifts: the approval rate that was 75% in week 6 may drop to 60% in week 10 if the team starts asking different questions. The 8-week engagement includes a handover document that specifies the re-indexing cadence, the scoring threshold review schedule (monthly), and the escalation path if error rate exceeds 15% on a rolling 50-ticket window. Without this, the agent degrades silently within 60 days.

    7. Define the success metric before the pilot starts

    The 8-week engagement is not a product launch; it is a measured experiment with a clear success criterion. For a UAE medtech firm, the success criterion is: first-response time under 30 minutes on 80% of tickets, error rate under 12%, and the team reporting that the agent saves at least 3 hours per week per agent. The audit phase sets the baseline, the pilot measures against it, and the final report states whether the criterion was met. If it was, the team decides whether to scale to all ticket categories, add a voice channel, or extend the RAG layer to other internal tools. If it was not, the report identifies which axis failed (retrieval, generation, or scoring) and what the next iteration should target. The engagement ends with a decision, not a demo.

  • 12-Step Checklist: RAG Pilot for Lead Qualification in German Insurance

    Pre-Pilot: Scope and Compliance Setup

    A 2-week fixed-scope pilot in a German insurance firm must produce a working RAG assistant on one workflow, a GDPR-compliant data-flow document, and a measured before/after baseline. The checklist below is operational: each item is a task a team can mark done or not done. It assumes the team uses LangChain and LangGraph, integrates with Slack or Microsoft Teams, and targets lead qualification to cut first-response time. The pilot is not a production deployment; it is a scoped experiment with a clear exit criterion. Work through the items in order. Skipping the audit or the baseline measurement invalidates the pilot’s value as a decision input for rollout.

    Build the RAG Pipeline on LangChain and LangGraph

    The RAG pipeline is the core of the pilot. Build it on LangChain for document chunking, embedding, and vector search, and on LangGraph for the stateful workflow that routes queries, handles multi-turn context, and triggers the human-approval gate. Keep the graph simple: one retrieval node, one generation node, one approval gate. Use a managed vector store in an EU region for the pilot. If the client’s data cannot leave the building, switch to an on-premises vector store and an open-weight model on the client’s GPU hardware. The RAG code is identical; only the embedding and inference endpoints change. Test the pipeline against 20 real lead queries before integrating with Slack or Teams.

    Integrate with Slack or Microsoft Teams

    The pilot must integrate with the channel the team already uses: Slack or Microsoft Teams. Build a bot that receives the lead query, calls the RAG pipeline, and returns the draft qualification score and suggested next step. The bot must include a human-approval gate: if the AI’s confidence drops below a threshold, or if the lead involves health-related data, the bot flags the query for a human agent. Log every human override. The integration must not replace the existing CRM or helpdesk; it plugs into them via their APIs. For a 2-week pilot, use a single OpenAI or Anthropic API endpoint for the LLM layer. Keep the model-agnostic layer thin: a single abstraction over the API call so switching providers later requires only a config change.

    Measure the Before/After Baseline

    Before the pilot starts, measure the baseline: cycle time from lead entry to qualified status, and error rate (misclassified leads) for the 2 weeks prior. Document the sample size, the definition of ‘error,’ and the measurement method. During the pilot, measure the same metrics for the 2 weeks of the pilot. The before/after comparison is the pilot’s primary deliverable. Without it, the client has no objective basis for the rollout decision. The baseline report must include: the number of leads processed, the average cycle time before and after, the error rate before and after, and the number of human overrides. This report is the exit criterion for the pilot.

    Define the Pilot Exit Criterion

    The pilot is a fixed-scope engagement: the vendor delivers a defined set of artifacts within the 2-week deadline. It is not a subscription or managed service. After the pilot, the client decides whether to proceed to rollout. The pilot includes a measured before/after comparison on cycle time and error rate, giving the client objective data to justify or reject the full deployment. The exit criterion is clear: if the pilot reduces cycle time by at least 30% and error rate by at least 20%, the client proceeds to rollout. If not, the pilot ends, and the client retains the baseline report and the RAG pipeline code. The vendor does not retain any client data after the pilot ends.

    Maintain the Checklist Over Time

    The checklist is a living document. After the pilot, review each item: mark what worked, what did not, and what needs adjustment. If the pilot proceeds to rollout, update the checklist to reflect the new scope: additional workflows, multi-language support, production monitoring. If the pilot ends, archive the checklist with the baseline report. Revisit the checklist before any new pilot: the GDPR landscape, the LLM provider landscape, and the integration landscape change. The checklist is not a one-time artifact; it is a tool for continuous improvement in AI-native operations. Keep it in the team’s project management tool, not in a static PDF.

  • Six Ways Forfis Automates Ticket Triage for Swiss Fintechs in a 3-Month Sprint

    1. Triage eats senior hours that should go to disputes

    A 2,000-employee payments firm in Zurich runs Zendesk as its primary helpdesk. Senior agents spend 40% of their day re-routing misclassified tickets and drafting first responses that follow the same template every time. The process audit identifies ticket triage and routing as the highest-impact workflow: 12,000 tickets per month, a median cycle time of 4.2 hours from receipt to first response, and a 14% error rate on routing. The AI layer classifies by intent, urgency, and department, then routes to the correct queue. After the pilot, median cycle time drops to 38 minutes and routing errors fall to 2.1%. The senior agents who previously handled triage now focus on complex disputes and fraud escalations, work that actually requires their judgment. The 3-month sprint covers audit, pilot, and rollout, with every decision logged for ISO 27001 audit trails.

    2. Workflow orchestration, not a chatbot wrapper

    The orchestration layer sits between Zendesk’s API and the model inference endpoint. Incoming tickets trigger a webhook that passes the ticket body, metadata, and customer history to the classifier. The model returns a structured JSON object with intent, urgency score, and recommended queue. The orchestrator validates the output against a schema, checks confidence thresholds, and routes the ticket accordingly. If confidence falls below 0.85, the ticket flags for human review. Every step logs a timestamp, model version, and input hash. This architecture means the client can swap the classifier model without touching the Zendesk integration or the routing logic. The orchestration layer is the stable contract; the model is a pluggable component.

    3. On-premise open-weight models keep regulated data local

    Swiss data protection law and the client’s ISO 27001 certification require that customer payment data never leaves the building. Forfis deploys an open-weight model on the client’s own GPU cluster, handling all ticket payloads that contain account numbers, transaction IDs, or personal identifiers. The model runs on-premise, so no regulated data crosses a network boundary. For non-sensitive workflows, such as routing a general FAQ ticket, the orchestrator can route to a cloud API where latency and cost are less critical. The model-agnostic design means the client chooses the model per workflow, not per project. This split keeps the ISO 27001 statement of applicability clean: the on-premise path satisfies Annex A.8.22 (use of cryptography) and A.8.15 (access control) without requiring a separate risk assessment for cloud data transfer.

    4. A 3-month sprint with a measured before/after baseline

    The pilot runs on one workflow for six weeks. The baseline is measured in the first two weeks: 12,000 tickets, 4.2-hour median cycle time, 14% routing error rate. The AI layer goes live in week three, handling triage and routing with human approval on any ticket flagged below the confidence threshold. By week six, the metrics show a 38-minute median cycle time and a 2.1% error rate. The before/after comparison is documented in a one-page report that the client’s CFO uses to justify the rollout budget. The pilot also surfaces edge cases: 3% of tickets contain multilingual content that the model misclassifies, prompting a fine-tuning pass before full rollout. This measured approach means the client sees ROI before committing to broader automation across invoice processing or document extraction.

    5. Human-in-the-loop by default, not as an afterthought

    The AI layer classifies and routes, but a human approves any action that touches money, health data, or a contract. In a payments context, this means the model drafts a refund response or flags a fraud-related ticket, but a senior agent signs off before the action executes. The approval threshold is not arbitrary; it is set based on the pilot’s error-rate data. If the model’s routing accuracy on fraud-related tickets is 97%, the human-in-the-loop threshold applies to that subset. For general FAQ tickets where accuracy is 99.5%, the system can auto-route without approval. This tiered approach frees senior staff from routine work while keeping them in the loop for high-stakes decisions. The approval log feeds directly into the ISO 27001 audit trail, showing who approved what and when.

    6. Plugs into Zendesk or Intercom without replacing them

    The integration sprint adds a new processing layer on top of the existing Zendesk or Intercom instance. No data migration is required; the AI layer reads tickets through the helpdesk’s native API and writes routing decisions back through the same API. The client’s existing workflows, SLAs, and reporting dashboards continue to function unchanged. The orchestration layer exposes a REST API that the helpdesk calls, so the integration is a few lines of configuration in Zendesk’s webhook settings. This means the client does not need to retrain agents on a new interface or rebuild their ticket taxonomy. The AI layer is invisible to the end customer; it simply makes the existing system faster and more accurate. The 3-month timeline includes two weeks of integration hardening after the pilot, where edge cases from the pilot are addressed and the system is stress-tested under production load.

  • How a 15-Person UK Medtech Firm Cut Back-Office Errors 40% in Two Weeks

    1. The pilot scope is one workflow, not a platform

    A 15-person medtech company in Manchester was losing 11 hours per week to manual invoice data entry and document extraction. The finance lead typed supplier invoices into the ERP, cross-checked line items against purchase orders, and flagged discrepancies in a shared spreadsheet. Error rate: 6.2% on a sample of 200 invoices. Cycle time: 4.3 hours per batch.

    The fix was not a new hire. It was a fixed-scope pilot with a two-week deadline: automate the extraction and validation step for one supplier, integrate it into the existing ERP via API, and measure the before/after delta. The pilot used the Anthropic Claude API for document parsing because the invoice formats were inconsistent and required nuanced field mapping. A human approved every extracted record before it hit the ERP. The result: error rate dropped to 1.8%, cycle time fell to 1.1 hours per batch, and the finance lead spent the freed time on supplier negotiations instead of data entry.

    2. The integration lives in Slack, not a new dashboard

    The pilot ran inside the team’s existing Slack workspace. A bot posted extracted invoice fields into a dedicated channel, tagged the finance lead for approval, and logged the decision. No new UI, no new login, no training session. The integration used the Slack API and the ERP’s REST endpoint — both already in production.

    This matters because a 15-person team does not have the bandwidth to adopt a new tool. The workflow orchestration layer sat between the Claude API and the ERP: it handled retries, format validation, and the approval gate. When the finance lead approved a record in Slack, the orchestration layer pushed it to the ERP. When they rejected it, the bot asked for the correction and re-processed. Every interaction was logged for the GDPR audit trail. The team never left Slack. The AI never replaced the ERP. It filled the gap between the two.

    3. GDPR compliance is a design constraint, not an afterthought

    The pilot processed supplier invoices, which contain no patient data. But the company’s broader documentation — SOPs, regulatory checklists, clinical trial protocols — does. The architecture was designed from day one to be model-agnostic: the orchestration layer could route a request to the Anthropic Claude API for general document work, or to an open-weight model running on the company’s own server for anything touching special-category data under GDPR Article 9.

    The DPIA was completed before the pilot started. It documented: what data the AI processes, where it is stored, who can access it, and how a human can override any automated decision. The Data Processing Agreement with Anthropic was signed. The open-weight model (a 7B-parameter Llama variant) ran on a single GPU workstation in the office. No patient data left the building. The pilot’s scope was narrow enough that the compliance overhead was a one-day task, not a multi-week project.

    4. The baseline is measured, not assumed

    The pilot’s success metric was not “the AI works.” It was: error rate drops from 6.2% to under 3%, and cycle time drops from 4.3 hours to under 2 hours per batch. The baseline was measured in week one, before any automation was live. The team processed 50 invoices manually and logged every error and every minute. In week two, the AI processed the same 50 invoices, and the finance lead approved or corrected each one. The delta was the deliverable.

    This is what separates a pilot from a demo. A demo shows the AI extracting fields from a sample PDF. A pilot measures whether the extraction is accurate enough to trust in production, and whether the human approval step is fast enough to be worth the overhead. The 4.3-hour to 1.1-hour drop was not theoretical. It was logged in the ERP’s audit trail, timestamped, and attributable to the automation.

    5. The rollout is a sequence of fixed-scope engagements

    The pilot’s scope was one supplier, one document type, one integration point. The rollout plan was explicit: week three adds the second supplier, week four adds the third, week five adds the document extraction for purchase orders. Each expansion was a separate fixed-scope engagement with its own baseline and success metric.

    This is how a 15-person company scales operations without new hires. The finance lead’s role did not change — she still approved every record. But the time she spent typing dropped from 4.3 hours to 1.1 hours per batch. The freed capacity went to supplier management, which had been neglected for two years. The company did not hire a data entry clerk. It did not buy a new ERP. It added an AI layer to the workflow it already ran, measured the delta, and expanded only when the numbers justified it.

    6. The knowledge search is a byproduct, not the goal

    The pilot’s real value was not the 40% error reduction. It was the internal knowledge search capability that emerged from the same architecture. The orchestration layer that routed invoice data to the ERP was repurposed to route queries to the company’s document store. A support agent in Slack could now ask, “What is the recall procedure for device X?” and get a cited answer from the SOP, with the relevant section highlighted. The agent still reviewed the answer before sending it to a customer. The AI did not replace the agent. It cut the search time from 12 minutes to 90 seconds.

    The synthesis: a 15-person UK medtech company did not need a new hire, a new ERP, or a new helpdesk. It needed a two-week fixed-scope pilot that measured a real delta, ran inside the tools the team already used, and kept a human in the loop for every high-stakes action. The AI was a layer, not a replacement. The compliance was a constraint, not a blocker. The rollout was a sequence, not a big bang. That is the pattern that works when the team is small, the data is regulated, and the timeline is two weeks.