Tag: Switzerland

  • Forfis AI Automation Audit and Pilot for Swiss Healthcare and Medtech Operations

    The Problem: Manual Back-Office Work in Swiss Healthcare and Medtech

    You run a 300-person healthcare or medtech company in Switzerland. Your operations team processes 400-600 invoices per month, each taking 12-18 minutes to key into the ERP. Your customer support team handles 150-250 tickets per week, with a median first-response time of 4.2 hours. You want to cut first-response time to under 30 minutes and reduce invoice processing cycle time by 60%, but you cannot send patient-adjacent data to a public cloud API. You need an AI-native operations layer that runs on your own hardware, integrates with your existing ERP and Google Workspace, and ships in 4 weeks. This is the exact scenario Forfis is built for: a fixed-scope pilot on one workflow, measured against a before/after baseline, with human-in-the-loop approval for anything touching money or health data.

    Prerequisites: What You Need Before the Audit Starts

    Before the audit begins, you need four things in place. First, access to your ERP system with read permissions on the invoice module and write permissions on the posting queue. Second, a sample of 50-100 recent invoices in PDF or image format, including at least 10 with line-item errors or missing fields. Third, access to your helpdesk or ticketing system with read permissions on the last 90 days of tickets, including timestamps for first response and resolution. Fourth, a named business owner who can approve scope changes and sign off on the pilot success criteria. You do not need to clean your data before the audit; the audit itself identifies data readiness gaps. You do need to confirm that your IT team can provision a virtual machine or container on your on-premise network for the open-weight model deployment.

    Step 1: Run the Process Audit and Select the Pilot Workflow

    Days 1-5. Forfis reviews your invoice processing workflow end-to-end: how invoices arrive (email, portal, paper), how they are keyed, how errors are handled, and where they sit in the ERP. The deliverable is a process map with cycle time and error rate baselines. You select one workflow for the pilot based on the audit’s prioritization matrix. The pilot scope is fixed: one workflow, one model configuration, one integration point. If you want to automate both invoice processing and customer triage, you run two separate pilots, not one combined engagement.

    Step 2: Deploy the Open-Weight Model on Your On-Premise Hardware

    Days 6-10. Forfis provisions an open-weight model, typically Llama 3 70B or Mistral 8x7B, on your on-premise hardware. The model is fine-tuned on your invoice samples or ticket history, depending on the pilot scope. For invoice processing, the model is trained to extract vendor name, invoice number, line items, tax amounts, and due date from PDF or image input. For customer triage, the model is trained to classify ticket urgency and draft a first response. The fine-tuning dataset is built from your historical data, not synthetic data. You review the model’s output on a holdout set of 20-30 items before it goes live.

    Step 3: Integrate the Agent with Your ERP and Google Workspace

    Days 11-15. Forfis connects the AI agent to your ERP and Google Workspace through their native APIs. For invoice processing, the agent reads the invoice PDF from your email or shared drive, extracts the fields, and posts a draft entry to the ERP posting queue. A human approver reviews the draft in the ERP and clicks approve or reject. For customer triage, the agent reads new tickets from your helpdesk, classifies them, and drafts a first response in Google Workspace. The human agent reviews the draft and sends it. The integration is read-write, so the agent logs its actions in your existing tools without requiring your team to switch platforms.

    Step 4: Run the Pilot in Parallel with Your Existing Process

    Days 16-20. The pilot runs in parallel with your existing process. For invoice processing, the agent processes a subset of invoices, say 20% of the daily volume, while your team continues to process the rest manually. For customer triage, the agent drafts first responses for a subset of tickets, say 30% of the weekly volume, while your team handles the rest. You measure cycle time and error rate for both the agent and the manual process. The success criteria are defined in the audit: for example, a 60% reduction in invoice processing cycle time and a 95% accuracy rate on field extraction. If the agent misses the criteria, Forfis adjusts the model configuration or the integration logic and re-tests.

    Step 5: Validate the Pilot and Roll Out to Full Volume

    Days 21-25. You review the pilot results against the success criteria. If the agent meets the criteria, you proceed to rollout. The rollout expands the agent’s scope from the pilot subset to 100% of the workflow volume. For invoice processing, this means the agent processes all incoming invoices, with human approval still required for anything touching money. For customer triage, this means the agent drafts first responses for all new tickets, with human review before sending. The rollout takes 3-5 business days, during which Forfis monitors the agent’s performance and adjusts thresholds as needed. You do not change your team’s daily workflow; the agent works in the background, and your team approves or rejects its output in the tools they already use.

  • AI Workflow Automation vs. Compliance-Safe Rollout for Ticket Triage in B2B SaaS

    What Is Being Compared

    The two options under comparison are AI workflow automation and a compliance-safe AI rollout, both applied to ticket triage and routing in a B2B SaaS company with 2,000+ employees in Switzerland. AI workflow automation refers to the technical layer: an orchestration engine that classifies incoming support tickets, routes them to the correct queue, and drafts a first response using the OpenAI API. It integrates with the existing helpdesk and pulls context from Notion or Confluence via API. The compliance-safe rollout is the delivery and governance layer: a dedicated AI team runs a fixed-scope pilot over 8 weeks, with human-in-the-loop approval on every ticket that touches a customer, and a measured before/after baseline on cycle time and error rate. The two are not alternatives; they are the technical build and the delivery wrapper. The comparison below judges them against the criteria that matter for a 2,000+ employee organization scaling AI across departments.

    Criteria for Judgment

    The following criteria determine which approach fits the scenario. Each is judged against the specific dimensions: B2B SaaS, Switzerland, 2,000+ employees, 8-week timeline, ticket triage and routing, OpenAI API, Notion or Confluence integration, dedicated AI team delivery, and the goal of reducing error rate in the back office.

    • Cycle time reduction: measured from ticket creation to first routed response.
    • Error rate: percentage of misrouted or misclassified tickets.
    • Integration depth: how the AI connects to the helpdesk, Notion/Confluence, and CRM without replacing them.
    • Human-in-the-loop overhead: time a support agent spends approving AI-drafted actions.
    • Timeline feasibility: whether the 8-week window is realistic for pilot and baseline measurement.
    • Scalability across departments: whether the architecture extends to invoice processing, document extraction, and other workflows.
    • Vendor lock-in: whether the model-agnostic design allows swapping OpenAI for an open-weight model if data residency rules change.
    • Cost per ticket: API token cost plus human review time, compared to the current manual triage cost.

    Comparison Table

    Criterion AI Workflow Automation Compliance-Safe Rollout
    Cycle time reduction 40-60% reduction in triage-to-response time Same reduction, but gated by human approval step (adds 5-10 sec per ticket)
    Error rate 30-50% reduction in misrouting Same reduction, with human catch on low-confidence tickets (<0.85)
    Integration depth API connections to helpdesk, Notion/Confluence, CRM Same integrations, plus audit log and approval workflow
    Human-in-the-loop overhead Minimal if confidence threshold is high 5-10 sec per ticket for agent review; scales with ticket volume
    Timeline feasibility 8 weeks for pilot build and baseline 8 weeks includes audit, pilot, tuning, and handover
    Scalability across departments Model-agnostic; new workflows are new integrations Dedicated team runs process audit per department; 2-3 pilots in parallel
    Vendor lock-in OpenAI API; swappable to open-weight model Same; architecture is model-agnostic by design
    Cost per ticket ~EUR 0.02-0.05 in API tokens per ticket Same API cost plus ~EUR 0.10-0.20 in human review time

    Scenario-by-Scenario Verdict

    For a B2B SaaS company in Switzerland with no specific compliance mandate, the AI workflow automation layer is the primary value driver. The OpenAI API handles English-language ticket classification with high accuracy, and the Notion or Confluence integration provides the RAG context for first-response drafting. The 8-week timeline is feasible because the scope is limited to one workflow: ticket triage and routing. The dedicated AI team builds the orchestration, connects the APIs, and runs the pilot. The compliance-safe rollout adds the governance wrapper: human-in-the-loop approval, baseline measurement, and audit logging. For a company with 2,000+ employees, this wrapper is not optional; it is what makes the pilot acceptable to the support leadership and the finance team. The two layers are inseparable in practice: the automation without the rollout wrapper is a demo, not a production system.

    When the company scales across departments, the compliance-safe rollout becomes the scaling mechanism. The dedicated AI team runs a process audit for each new department—invoice processing, document extraction, data entry—and identifies the highest-ROI workflow. The 8-week timeline applies per workflow, not to the entire company. The model-agnostic architecture means each new workflow can use the same orchestration engine, with the OpenAI API for quality-critical tasks and open-weight models on the client’s hardware if a department handles regulated data. The dedicated AI team model ensures continuity: the same team that built the ticket triage pilot runs the next pilot, reducing onboarding friction and maintaining the baseline measurement methodology.

    Recommendation

    The recommendation is to run both layers as a single engagement, not as separate projects. The AI workflow automation is the technical build: an orchestration engine using the OpenAI API that classifies and routes tickets, pulls context from Notion or Confluence, and drafts first responses. The compliance-safe rollout is the delivery and governance wrapper: a dedicated AI team runs the 8-week pilot with human-in-the-loop approval, measures the before/after baseline on cycle time and error rate, and hands over to managed operation. For a 2,000+ employee B2B SaaS company in Switzerland with no compliance constraints, this combined approach is the only one that fits the 8-week timeline and the goal of reducing error rate in the back office. The automation layer delivers the speed and accuracy; the rollout wrapper delivers the trust and the measurement. Neither works without the other. The dedicated AI team owns the technical execution; the client’s support team owns the business outcomes and the human-in-the-loop approval. This split is the standard delivery model for Forfis engagements and is the one that scales across departments without re-architecting the stack.

  • Deploying a RAG Assistant for Order Status Updates in Swiss E-Commerce

    The Problem: Senior Support Staff Buried in Routine Order Status Tickets

    Your support team at a 2,000+ employee e-commerce company in Switzerland handles thousands of order and shipment status inquiries weekly. Senior agents spend 40-60% of their time on routine lookups: “Where is my package?” “Why is my order delayed?” This work does not require judgment, but it consumes the people who should be handling complex escalations, refund disputes, and customer retention conversations. The EU AI Act, which applies to Swiss companies serving EU customers, requires transparency when AI systems interact with users. You need a retrieval-augmented knowledge assistant that drafts accurate responses from your order-management system and shipping carrier data, integrates with Zendesk or Intercom, and keeps a human in the loop for anything touching refunds or contract terms. The goal: cut first-response time from hours to minutes, reduce error rate on shipping information, and free senior staff for high-value work within a 3-month integration sprint.

    Prerequisites: What You Need Before the Sprint Starts

    Before starting the integration sprint, confirm these are in place:

    • Zendesk or Intercom API access: OAuth 2.0 tokens with read/write permissions for tickets, macros, and webhooks. Test with a sandbox account first.
    • Order-management system (OMS) API: Read access to order status, tracking numbers, and shipping carrier data. If you use Shopify, SAP Commerce, or a custom OMS, document the endpoint schema.
    • Shipping carrier APIs: Integration with at least your top two carriers (e.g., Swiss Post, DHL) for real-time tracking events.
    • PostgreSQL 15+ with pgvector extension: CREATE EXTENSION vector; Run on a dedicated instance with at least 16 GB RAM for 500k+ vectors.
    • LLM endpoint: OpenAI API key (gpt-4o or claude-3-5-sonnet) for drafting, or an on-prem Llama 3 70B instance if customer PII cannot leave your infrastructure.
    • EU AI Act compliance documentation: A data-protection impact assessment (GDPR Article 35) and a model card for each LLM endpoint.
    • Baseline metrics: Export 30 days of ticket data from Zendesk/Intercom. Calculate average first-response time, resolution rate, and error rate on shipping-related tickets.

    Step 1: Audit the Workflow and Establish a Baseline

    Run a process audit on your last 90 days of support tickets. Filter for order and shipment status inquiries: “Where is my order?” “Tracking number not working” “Delivery delayed.” Count the volume, measure average handling time, and identify the top five questions. For a 2,000+ employee e-commerce company, this typically represents 35-50% of total ticket volume. Export the data to a CSV with columns: ticket_id, subject, category, first_response_time, resolution_time, agent_id, error_flag. Calculate the baseline: if your average first-response time is 4 hours and error rate on shipping information is 8%, those are your targets to beat. Document this baseline in a one-page report. This becomes the measurement framework for the pilot and rollout phases.

    Step 2: Build the RAG Pipeline with pgvector

    Build the retrieval layer using pgvector. Chunk your knowledge base: shipping policies, carrier SLAs, return procedures, and order status definitions. Use a 512-token chunk size with 50-token overlap. Generate embeddings with OpenAI text-embedding-3-small (1536 dimensions) or bge-base-en-v1.5 (768 dimensions) if you prefer open-weight models. Load into PostgreSQL:

    CREATE TABLE documents (
      id SERIAL PRIMARY KEY,
      content TEXT,
      metadata JSONB,
      embedding vector(1536)
    );
    CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
    

    Set ef_search = 64 for sub-10 ms recall. Test with 20 sample queries: “Where is my order with tracking number XYZ?” Verify that the top-5 retrieved chunks contain the relevant shipping policy and carrier SLA. If recall is below 90%, adjust chunk size or add metadata filters (e.g., WHERE metadata->>'carrier' = 'DHL').

    Step 3: Integrate with Zendesk or Intercom via Webhooks

    Connect the RAG pipeline to Zendesk or Intercom. For Zendesk: create a webhook on ticket creation that triggers your RAG service. The service retrieves relevant chunks, calls the LLM endpoint with a system prompt: “You are a support assistant for [Company]. Use only the retrieved context to draft a response. If the context does not contain the answer, say so. Do not invent tracking numbers or delivery dates.” Post the drafted response to the ticket via the API with a RAG-drafted tag. For Intercom: use the Events API to trigger on ticket.created and the Agent Inbox API to post the draft. Store the correlation ID (ticket_id + timestamp) in a log table for audit trails. This satisfies EU AI Act Article 50 transparency requirements: users are informed they are interacting with an AI, and every response is traceable to its source documents.

    Step 4: Add Human-in-the-Loop Approval for Sensitive Actions

    Implement the human-in-the-loop approval workflow. Any RAG-drafted response that touches refunds, address changes, or contract terms must be approved by a human before sending. In Zendesk, create a custom field ai_approval_status with values: pending, approved, rejected. When the RAG service posts a draft, set ai_approval_status = pending and assign the ticket to a supervisor queue. The supervisor reviews the draft, the retrieved context, and the LLM’s confidence score. If approved, the ticket moves to approved and the response sends. If rejected, the supervisor edits or reassigns. Log every approval decision with the supervisor’s user ID and timestamp. This workflow is mandatory under EU AI Act Article 50 for any AI system that makes decisions affecting consumers. For a 3-month sprint, build a simple approval UI in React or use Zendesk’s built-in ticket views filtered by ai_approval_status = pending.

    Step 5: Pilot with 10-20% of Tickets and Measure

    Run the pilot with 10-20% of order-status tickets for two weeks. Route a subset of tickets (e.g., all tickets tagged order_status from a specific region or carrier) to the RAG assistant. Measure: first-response time (target: under 15 minutes vs. baseline 4 hours), resolution rate (target: 80%+ first-contact resolution), and error rate on shipping information (target: under 2% vs. baseline 8%). Sample 5% of AI-drafted responses weekly. Compare each against the OMS and carrier API data. If the assistant states a delivery date, verify it matches the carrier’s tracking event. If error rate exceeds 2%, pause the pilot, re-index the knowledge base, and adjust the LLM prompt to require citation of specific tracking events. Document every error in a log with the ticket ID, the incorrect claim, and the correct data from the OMS. This log feeds into the EU AI Act model card and the GDPR Article 35 impact assessment.

  • 4-Week Pilot: LangGraph Ticket Triage Agent for Swiss Professional Services

    The Problem: Manual Ticket Triage in a Swiss Professional Services Firm

    You run a 501-2000 employee professional services firm in Switzerland. Your operations team spends 12-18 hours per week manually triaging client tickets, routing them to the wrong queue, and re-keying data into the CRM. The EU AI Act does not directly apply to Swiss firms, but your EU-based clients will contractually demand Article 50 transparency for any AI system that touches their data. You have already automated one back-office process (invoice processing), and now you want to extend AI to customer-facing channels. The specific use case is ticket triage and routing: classify incoming tickets, extract key entities, route to the correct queue, and draft a first response. The constraint is a 4-week fixed-scope pilot with a measurable before/after baseline on cycle time and error rate. The architecture must plug into your existing helpdesk and CRM via custom REST API and webhooks, not replace them.

    Prerequisites Before Step 1

    • Helpdesk API access: Your helpdesk (e.g., Zendesk, Freshdesk, or a custom system) must expose a REST API with endpoints for: listing tickets, fetching ticket details, updating ticket status, and creating webhooks for new ticket events. You need OAuth 2.0 or API key authentication.
    • CRM integration: Your CRM (e.g., Salesforce, HubSpot, or a custom system) must expose a REST API for reading and writing client records. The agent will need to fetch client context (contract type, SLA tier, historical tickets) to inform routing decisions.
    • Model access: You need API keys for at least one LLM provider (OpenAI, Anthropic, or a self-hosted open-weight model). For the pilot, one model is sufficient; the architecture should support swapping models later.
    • Human approval UI: A simple web interface where a human can review the agent’s proposed classification, extracted entities, and draft response, then approve, edit, or escalate. This can be a lightweight React app or a form in your existing internal tool.
    • Baseline data: At least 200 historical tickets with timestamps, queue assignments, and resolution notes. This is your before/after measurement set.
    • Legal review: A 1-hour consultation with your legal team to confirm EU AI Act applicability and any Swiss-specific data protection requirements under the FADP (Federal Act on Data Protection).

    Step 1: Process Audit and Baseline Measurement

    Spend 3-4 days mapping the current triage workflow. Document: (1) the average cycle time from ticket creation to first human response, (2) the error rate (tickets misrouted or requiring rework), (3) the top 5 ticket categories by volume, and (4) the decision rules humans use to route tickets. For a 501-2000 employee firm, you should sample at least 200 tickets over 2 weeks. Record the baseline metrics in a spreadsheet: ticket_id, created_at, first_response_at, assigned_queue, final_queue, rework_flag. This baseline is the primary deliverable that justifies the pilot. Without it, you cannot measure improvement. The audit also identifies which ticket categories are worth automating: focus on the top 2-3 categories that account for 60-70% of volume and have clear, rule-based routing logic.

    Step 2: Build the LangGraph Agent with Intent Classification

    Set up the LangGraph agent with 3-5 intent classes corresponding to your top ticket categories. Each node in the graph represents a discrete action: classify_intent, extract_entities, fetch_client_context, route_to_queue, draft_response. The classify_intent node calls the LLM with a system prompt that defines each intent class and few-shot examples from your historical tickets. The extract_entities node pulls out key fields: client name, ticket ID, issue type, urgency. The fetch_client_context node calls your CRM REST API to get the client’s contract type and SLA tier. The route_to_queue node uses conditional edges: if urgency == 'high' or contract_type == 'enterprise', route to the human queue; otherwise, route to the automated queue. The draft_response node generates a first response using the client context and ticket details. The entire graph should be under 500 lines of Python code.

    Step 3: Integrate with Helpdesk via REST API and Webhooks

    Integrate the agent with your helpdesk via custom REST API and webhooks. The helpdesk sends a webhook to your agent’s endpoint when a new ticket is created. The agent’s endpoint receives the ticket ID, fetches the full ticket details via the helpdesk REST API, runs the LangGraph agent, and returns the proposed classification, extracted entities, and draft response. The agent then calls the helpdesk REST API to update the ticket status to ‘awaiting_human_approval’ and creates a task in your human approval UI. The human reviews the task, clicks ‘Approve’, ‘Edit’, or ‘Escalate’. If approved, the agent calls the helpdesk REST API to assign the ticket to the correct queue and post the draft response. If escalated, the agent assigns the ticket to a senior agent and logs the escalation reason. All API calls should be logged with timestamps for audit.

    Step 4: Implement Human-in-the-Loop Approval Workflow

    The human approval UI is a simple web app with three actions: ‘Approve’, ‘Edit’, ‘Escalate’. The UI displays: (1) the proposed intent classification with confidence score, (2) the extracted entities (client name, ticket ID, issue type, urgency), (3) the client context fetched from the CRM (contract type, SLA tier, historical tickets), (4) the draft response. The human can edit any field before approving. Every action is logged: ticket_id, action, timestamp, user_id, edited_fields. This log is your audit trail for EU AI Act compliance. The UI should be accessible from the helpdesk: add a ‘View AI Suggestion’ button on the ticket detail page that opens the approval UI in a new tab. The approval workflow adds 15-30 seconds per ticket, but it ensures accountability and builds trust during the pilot. For the 4-week pilot, target a 90% approval rate (humans approve without editing) as a success metric.

    Step 5: Measure Before/After Baseline and Ship the Report

    Run the pilot for 2 weeks with the agent in shadow mode: the agent processes every ticket, but the human approval workflow is the only path to action. After 2 weeks, measure the same 200 tickets (or an equivalent sample) with the agent in place. Compare: (1) cycle time from ticket creation to first human response, (2) error rate (misrouted tickets or rework), (3) human effort saved (hours per day). The before/after report should show: cycle time reduction (target: 40-60%), error rate change (target: <5% misclassification), and human effort saved (target: 3.5 hours per day). This report is the primary deliverable that justifies rollout to additional ticket categories. If the pilot meets the targets, the next step is a 6-week rollout to the remaining ticket categories, with the same human-in-the-loop workflow and baseline measurement. If the pilot misses the targets, iterate on the intent classification prompt or the routing rules before proceeding.

  • Cutting Contract Review Cycle Time in Swiss E-Commerce: A 2-Week AI Pilot

    The Contract Review Bottleneck in Swiss E-Commerce

    A 201-500 employee e-commerce company in Switzerland runs its legal and compliance function on a small team. Contract review for vendor agreements, data processing agreements, and customer-facing terms consumes 4 to 6 hours per document. The legal team tracks cycle time manually in a spreadsheet, and error rate on standard clauses sits at 12 to 18 percent because reviewers work through queues without a consistent precedent library. First-response time on internal compliance queries from the sales and operations teams averages 2 to 3 business days because the legal team is buried in contract work. The cost per support ticket that touches a contract question runs 35 to 50 Swiss francs in legal time, and the team has no baseline to measure improvement. The company has run two isolated AI pilots in the last 18 months, neither of which reached production because the scope was undefined and the integration with existing systems was never planned.

    Why Isolated Pilots Stall in Legal and Compliance

    Most companies in this position reach for one of three approaches, and each fails in a predictable way. The first is a generic LLM wrapper: a legal team member pastes a contract into ChatGPT and asks for a summary. This produces plausible-sounding output that misses jurisdiction-specific clauses, Swiss data protection requirements under the revised nFADP, and the company’s own precedent language. The second is a RAG pipeline built on a single document store without a structured extraction layer. The retrieval step finds relevant clauses, but the extraction step that pulls out party names, payment terms, and liability caps is brittle and requires manual correction on 30 to 40 percent of documents. The third is a full vendor platform that replaces the existing CRM and document management system. The integration cost alone exceeds the annual legal budget for a 201-500 employee firm, and the migration timeline stretches past 12 months. None of these approaches ship a measured before/after baseline, so the company cannot prove the pilot reduced cycle time or error rate.

    A Fixed-Scope Pilot That Ships in Two Weeks

    The fix starts with a 2-week AI automation audit that maps the contract review workflow end to end. Forfis interviews the legal team, identifies the top 3 to 5 document types by volume and error rate, and scores each on automation feasibility and data sensitivity. The audit delivers a fixed-scope pilot proposal on the single workflow with the best risk-to-reward ratio, typically standard vendor contracts. The pilot architecture uses a model-agnostic stack: OpenAI or Anthropic APIs for classification and drafting where quality matters, open-weight models on the client’s own hardware where regulated data cannot leave the building. A pgvector embeddings search layer indexes the company’s contract templates, precedent clauses, and compliance checklists from Notion or Confluence, so the AI agent retrieves relevant language before drafting. The system plugs into the existing CRM and helpdesk through their APIs rather than replacing them. Every pilot ships with a measured before/after baseline on cycle time and error rate, and the human-in-the-loop approval step ensures no contract touches a counterparty without legal sign-off.

    How to Start: Five Concrete Steps

    Week 1 of the audit: Forfis maps the current contract review process, identifies the top 3 to 5 document types by volume, and records baseline cycle time and error rate on a sample of 50 to 100 historical contracts. The team interviews the legal and compliance staff to understand which clauses are non-negotiable and which can be auto-classified. Week 2: the team builds a proof-of-concept extraction pipeline on the sample, measures the before/after delta, and delivers a fixed-scope pilot proposal with cost, timeline, and EU AI Act compliance controls. The pilot itself runs 4 to 6 weeks and ships with a measured baseline. From there, rollout extends to additional document types and the managed operation phase handles model updates, drift monitoring, and compliance reporting. The first step is to schedule the audit. The second is to gather 50 to 100 historical contracts in a shared Notion or Confluence workspace. The third is to identify the single workflow with the highest volume and error rate. The fourth is to define the success metric: cycle time reduction and error rate drop. The fifth is to assign a legal owner who will approve every AI-drafted output during the pilot.

  • Swiss E-commerce Retailer Cuts Reporting Cycle Time 70% with AI Automation

    Background and Challenge

    This case study is a composite based on patterns observed in the field. It does not represent a single named customer but reflects common challenges and solutions in the e-commerce and retail sector in Switzerland.

    Background
    A mid-sized Swiss e-commerce retailer with 1,200 employees operates across DACH markets. The company uses a custom-built CRM and ERP system, with data stored in on-premise servers. The sales team of 45 handles lead qualification and monthly reporting manually, using spreadsheets and email. The company has no AI in production yet and is looking to reduce manual back-office work while improving lead qualification accuracy.

    Challenge
    The sales team spends 12 hours per week on monthly reporting, manually aggregating data from the CRM, ERP, and web analytics. The process is error-prone, with a 10% error rate in data entry. Lead qualification is inconsistent, with 30% of leads being misclassified, leading to lost opportunities. The company faces GDPR compliance requirements and a deadline to implement improvements before the Q4 peak season.

    Approach
    Forfis conducted an AI automation audit, identifying monthly reporting and lead qualification as high-impact use cases. A fixed-scope pilot was designed to automate these workflows using LangChain and LangGraph for workflow orchestration. The system integrates with the existing CRM and ERP via custom REST APIs and webhooks. A human-in-the-loop model ensures that AI-generated reports and lead scores are reviewed by a human before finalization. The pilot was deployed in two weeks, with a measured before/after baseline on cycle time and error rate.

    Outcome
    The pilot reduced monthly reporting cycle time from 5 days to 1 day, a 70% improvement. The error rate decreased from 10% to 2%, an 80% reduction. Lead qualification accuracy improved from 70% to 95%, with a 25% increase in qualified leads passed to sales. The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building.

    Lessons

    • Start with a fixed-scope pilot to demonstrate ROI quickly.
    • Use a human-in-the-loop model to ensure accuracy and compliance.
    • Integrate with existing systems via APIs rather than replacing them.
    • Measure before/after baselines to quantify impact.
    • Choose a model-agnostic architecture to future-proof the solution.

    Approach: AI Automation Audit and Pilot Design

    The AI automation audit identified two high-impact use cases: monthly reporting and lead qualification. The audit mapped existing workflows, identified bottlenecks, and evaluated the feasibility of automating specific tasks. The results were a prioritized list of use cases, with estimated ROI and implementation complexity.

    Monthly Reporting
    The current process involves manually aggregating data from the CRM, ERP, and web analytics. The sales team spends 12 hours per week on this task, with a 10% error rate in data entry. The AI system automates data collection, validation, and report generation. It uses LangChain to chain prompts and tools, and LangGraph to define stateful, multi-step workflows. The system integrates with the existing CRM and ERP via custom REST APIs and webhooks, ensuring data integrity and real-time updates.

    Lead Qualification
    The current process is inconsistent, with 30% of leads being misclassified. The AI system uses a classification model to score leads based on predefined criteria, such as company size, industry, and engagement level. The model is trained on historical data and fine-tuned using feedback from the sales team. A human-in-the-loop model ensures that AI-scored leads are reviewed by a human before they are passed to sales, ensuring accuracy and context.

    GDPR Compliance
    The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building. Data minimization is implemented, and data subjects can exercise their rights. The legal basis for processing is documented, and third-party AI APIs are GDPR-compliant. The system uses open-weight models on the client’s own hardware, ensuring that regulated data does not leave the building.

    Outcome: Measured Impact on Cycle Time and Error Rate

    The pilot was deployed in two weeks, with a measured before/after baseline on cycle time and error rate. The system was integrated with the existing CRM and ERP via custom REST APIs and webhooks, ensuring seamless data flow. The human-in-the-loop model was implemented, with a review dashboard for the sales team to approve AI-generated reports and lead scores.

    Cycle Time
    The monthly reporting cycle time was reduced from 5 days to 1 day, a 70% improvement. The AI system automates data collection, validation, and report generation, eliminating manual data entry and aggregation. The sales team spends 2 hours per week on review and approval, compared to 12 hours previously.

    Error Rate
    The error rate in monthly reporting decreased from 10% to 2%, an 80% reduction. The AI system validates data in real-time, flagging anomalies and inconsistencies. The human-in-the-loop model ensures that errors are caught and corrected before the report is finalized.

    Lead Qualification Accuracy
    Lead qualification accuracy improved from 70% to 95%, with a 25% increase in qualified leads passed to sales. The AI system scores leads based on predefined criteria, and the human-in-the-loop model ensures that misclassified leads are corrected. The sales team reports a 15% increase in conversion rates, attributed to more accurate lead qualification.

    GDPR Compliance
    The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building. The legal basis for processing is documented, and data subjects can exercise their rights. The system uses open-weight models on the client’s own hardware, ensuring that regulated data does not leave the building.

    Lessons for Similar Teams

    The pilot demonstrated significant improvements in cycle time, error rate, and lead qualification accuracy. The system is GDPR-compliant and integrated with existing systems via APIs. The human-in-the-loop model ensures accuracy and compliance, while the model-agnostic architecture provides flexibility and future-proofing.

    Scalability
    The system can be scaled to automate other workflows, such as invoice processing and document extraction. The model-agnostic architecture allows for switching between different AI models, based on cost, performance, and compliance requirements. The system can be extended to other departments, such as marketing and customer service, with minimal changes.

    Cost Efficiency
    The pilot reduced manual back-office work by 80%, saving 10 hours per week. The cost of the AI system is offset by the reduction in manual effort and the increase in qualified leads. The system is cost-effective, with a payback period of less than 3 months.

    Risk Mitigation
    The human-in-the-loop model mitigates the risk of errors and ensures compliance with regulations. The model-agnostic architecture mitigates vendor lock-in and allows for future-proofing. The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building.

    Next Steps
    The company plans to roll out the system to other departments, such as marketing and customer service. The system will be extended to automate other workflows, such as invoice processing and document extraction. The company will continue to measure the impact of the system on key metrics, such as cycle time, error rate, and lead qualification accuracy.

  • Swiss E-commerce Cuts Invoice Processing to 3 Hours with On-Premise AI

    Background: A Swiss E-commerce Operator at 300 Headcount

    This case study is a composite based on patterns observed across Forfis engagements. We do not name real customers. The company described here is a mid-sized Swiss e-commerce operator with roughly 300 employees, running a multi-channel retail operation across DACH and Western Europe. The stack is a mix of a legacy ERP for inventory and finance, a modern CRM for customer relationships, and a helpdesk platform for internal and supplier communications. The operations team handles 1,200 to 1,800 supplier invoices per month, plus a monthly consolidated report that feeds into the finance close. The company is in the AI-native operations stage: leadership has approved AI investment, but the team has not yet built internal capability to deploy and maintain AI workflows. The engagement ran over 8 weeks, delivered by a dedicated Forfis AI team embedded with the client’s operations group.

    Challenge: 12 Hours a Week of Manual Invoice Entry and a Fixed Monthly Close

    The operations team spent an estimated 12 to 15 hours per week on manual invoice processing: extracting line items from PDFs, matching them against purchase orders in the ERP, flagging discrepancies, and entering validated data. The monthly consolidated report required pulling data from three systems, reconciling it, and formatting it for the finance close. The error rate on the baseline was 4.2 percent on invoice line items, with a 3-day average cycle time from receipt to posting. The pressure was twofold: the monthly close deadline was fixed, and the team had lost two senior operators to attrition in the prior quarter. Leadership wanted to reduce manual back-office work without replacing the existing ERP or CRM, and without sending supplier or financial data to a third-party cloud. The compliance posture was internal: no regulatory mandate, but the finance director required that all financial data remain on-premise.

    Approach: On-Premise Open-Weight Models with a Human-in-the-Loop Approval Layer

    Forfis ran a two-week process audit to map the invoice workflow end-to-end and capture baseline metrics. The pilot scope was fixed: automate invoice extraction, PO matching, and discrepancy flagging, plus generate the monthly consolidated report from the same data pipeline. The architecture used open-weight models deployed on the client’s own hardware, so all invoice and financial data stayed on-premise. The AI layer connected to the ERP and helpdesk through custom REST APIs and webhooks: the ERP pushed new invoices via webhook, the AI service processed them, and validated records were written back through the ERP’s REST API. Discrepancies were pushed to the helpdesk as tickets for human review. The human-in-the-loop layer was built into the workflow: the model drafted and classified, a person approved anything touching a financial transaction. The dedicated Forfis team handled technical planning, product design, and full-cycle development over the 8-week timeline.

    Outcome: Cycle Time Down 75 Percent, Error Rate Under 1 Percent

    After the 8-week engagement, the measured results were: cycle time on invoice processing dropped from 12 to under 3 hours per week, a reduction of roughly 75 percent. The error rate on invoice line items fell from 4.2 percent to under 1 percent. The monthly consolidated report, which previously took 2 to 3 days of manual reconciliation, was generated automatically from the same data pipeline and required only a 30-minute human review. The human-in-the-loop approval queue handled roughly 8 to 12 percent of invoices that required manual review, down from 100 percent. The operations team redirected the freed capacity to supplier relationship management and exception handling. The finance director confirmed that all data remained on-premise throughout the pilot and rollout, and the monthly close process was unchanged in structure but faster in execution. The system is now in managed operation with Forfis monitoring model performance and handling drift.

    Lessons for Similar Teams

    • Baseline before you build. The 4.2 percent error rate and 12-hour cycle time were captured during the audit, not estimated. Without that baseline, the outcome metrics would be unverifiable. Any team automating a back-office workflow should measure the current state before touching the process.
    • One workflow, not five. The pilot scope was fixed to invoice processing and monthly reporting. Attempting to automate the entire back-office in 8 weeks would have diluted the team’s focus and made the baseline unmeasurable. Sequence the rollout: prove one workflow, then expand.
    • On-premise is not a constraint, it is a design choice. The open-weight model on the client’s hardware was not a compromise. It was the right fit for the data residency requirement, and the model-agnostic architecture meant the team could swap models without re-architecting the integration layer.
    • Human-in-the-loop is the default, not a fallback. The approval layer was built into the workflow from day one, not added after a failure. The 8 to 12 percent manual review rate is a feature, not a bug: it keeps the team in control of financial transactions while the AI handles the volume.
    • Integration through existing APIs, not replacement. The custom REST API and webhook layer connected to the ERP and helpdesk without requiring data migration. This kept the project within the 8-week timeline and avoided the risk of a parallel system.
  • AI Contract Review for Swiss Insurance: A 2-Week LangGraph Pilot

    The Problem: Contract Review at Scale in Swiss Insurance

    A 2,000+ employee Swiss insurer processes thousands of contracts annually. Each contract requires manual review by legal and compliance teams, taking 4-8 hours per document. The bottleneck is not the legal review itself, but the pre-review work: extracting key terms, classifying risk, and flagging missing clauses. This is where AI can help. The goal is not to replace lawyers, but to reduce the manual back-office work that precedes legal review. The pilot focuses on one process: contract review. The output is a system that extracts terms, scores risk, and flags issues, with a human approving anything that touches money, health data, or legal obligations. The timeline is 2 weeks, which is tight but feasible for a scoped pilot. The architecture is model-agnostic, using OpenAI and Anthropic APIs where quality matters, and open-weight models on the client’s own hardware where GDPR-sensitive data cannot leave the building. The integration is with Google Workspace, where contracts live in Gmail, Drive, and Docs. The system must support multilingual coverage: German, French, Italian, and English, reflecting Switzerland’s linguistic landscape. The delivery model is an integration sprint, not a full product build. The output is a working prototype with a measured before/after baseline on cycle time and error rate.

    The Mechanism: LangGraph Workflow and Predictive Scoring

    The system uses LangChain and LangGraph to orchestrate the contract review workflow. LangGraph provides a stateful, cyclic graph structure that maps well to the review process. Each node represents a step: extraction, classification, scoring, human review. Edges define transitions based on conditions. For example, if the risk score is above 70, the contract goes to legal review. If below 30, it may auto-approve. The extraction node uses a fine-tuned model to pull key terms: parties, dates, amounts, clauses. The classification node categorizes the contract type: life, health, property, liability. The scoring node assigns a risk score (0-100) based on detected clauses, missing terms, and historical data. The human review node presents the extracted terms, risk score, and flagged issues to a legal reviewer. The reviewer approves, rejects, or requests changes. The system logs every decision for audit trails. The integration with Google Workspace uses the Google Workspace API, with OAuth 2.0 for authentication. The AI reads contracts from Gmail, Drive, and Docs, drafts responses, and logs actions. Data stays within the client’s Google tenant, and the AI only accesses what the user has permission to see. The model-agnostic architecture routes documents to the appropriate model based on data sensitivity. GDPR-sensitive data goes to open-weight models on the client’s hardware. Non-sensitive data can use OpenAI or Anthropic APIs.

    Trade-offs: Model Choice, Automation Level, and Multilingual Support

    The architect faces several trade-offs. First, model choice: OpenAI and Anthropic APIs offer higher quality but raise GDPR concerns. Open-weight models on the client’s hardware are GDPR-compliant but may have lower accuracy. The solution is routing: a simple classifier determines which model handles each document based on data sensitivity. Second, automation level: full automation is faster but riskier. Human-in-the-loop is slower but safer. The pilot uses human-in-the-loop by default, with the option to auto-approve low-risk contracts after a period of measured accuracy. Third, multilingual support: supporting German, French, Italian, and English increases complexity. The model’s accuracy may vary by language, so test thoroughly. Use language-specific models or fine-tune on multilingual data. Fourth, integration depth: a shallow integration (read-only) is faster but less useful. A deep integration (read-write) is more useful but requires more time and testing. The pilot uses a shallow integration, with the option to deepen in subsequent phases. Fifth, scope: a broad scope (all contract types) is more ambitious but harder to deliver in 2 weeks. A narrow scope (one contract type) is more feasible but less impactful. The pilot focuses on one contract type, with the option to expand in subsequent phases.

    Recommendation: A Scoped Pilot with Measured Baselines

    For a 2,000+ employee Swiss insurer, the recommendation is to start with a scoped pilot on one contract type. Use LangGraph to orchestrate the workflow, with nodes for extraction, classification, scoring, and human review. Use predictive scoring to assign a risk score to each contract, with high scores triggering human review. Integrate with Google Workspace to read contracts from Gmail, Drive, and Docs. Use a model-agnostic architecture, routing GDPR-sensitive data to open-weight models on the client’s hardware and non-sensitive data to OpenAI or Anthropic APIs. Support multilingual coverage: German, French, Italian, and English. Use human-in-the-loop by default, with the option to auto-approve low-risk contracts after a period of measured accuracy. Measure before/after baselines on cycle time and error rate. The goal is to reduce the manual back-office work that precedes legal review, not to replace lawyers. The output is a working prototype with a measured baseline, not a production-ready system. Full rollout and managed operation follow in subsequent phases. The key is to automate one process well before attempting multiple. The 2-week timeline is tight but feasible for a scoped pilot. Week 1 covers the process audit, data mapping, and environment setup. Week 2 focuses on building the LangGraph workflow, integrating with Google Workspace, and running the first 50-100 test documents.

  • 2-Week AI Support Sprint for Swiss Professional Services Firms

    The Problem: Round-the-Clock Support Without Tripling Headcount

    Swiss professional services firms with 201-500 employees face a specific problem: customer support teams cannot provide round-the-clock coverage across German, French, Italian, and English without tripling headcount. The EU AI Act, which entered into force in August 2024, classifies customer support bots as limited-risk systems, requiring disclosure, human oversight, and documented evaluation metrics. A 2-week integration sprint addresses this by deploying an AI agent that handles first-response and ticket triage in Slack or Microsoft Teams, with a human-in-the-loop approval for anything touching money, health data, or contracts. The sprint delivers a measured baseline on cycle time and error rate, giving you concrete data before committing to full rollout. The architecture uses Anthropic Claude API for quality-critical tasks and open-weight models on local hardware for regulated data that cannot leave the building.

    Week 1: Process Audit and Prompt Engineering

    The sprint begins with a process audit that identifies which support workflows are worth automating. For a professional services firm, this typically includes ticket triage, first-response drafting, and internal knowledge search over CRM records and project documentation. The audit takes 2-3 days and produces a prioritized list of workflows ranked by volume, complexity, and compliance risk. The next phase is prompt engineering and API integration. The AI agent connects to your existing Slack or Microsoft Teams through their APIs, reads from your CRM and helpdesk, and drafts responses or classifies tickets. For multilingual coverage, the agent must be configured to detect the customer’s language and respond in German, French, Italian, or English as appropriate. Anthropic Claude supports 100+ languages, but you must ensure your internal documentation is available in each language for accurate retrieval-augmented responses.

    Week 2: Pilot Deployment and Baseline Measurement

    Week 2 focuses on the pilot and baseline measurement. The AI agent runs on one workflow, typically ticket triage or first-response drafting, while a human agent handles the same queue in parallel. The pilot measures cycle time (time from ticket creation to first response) and error rate (percentage of responses requiring human correction). For a professional services firm, a typical baseline shows a 40-60% reduction in cycle time and a 15-25% error rate for the AI agent, compared to 100% human handling. The human-in-the-loop approval ensures that anything touching money, health data, or contracts requires human sign-off before the response is sent. The pilot ships with a written report documenting the baseline metrics, the EU AI Act compliance checklist, and a recommendation for full rollout or scope adjustment. This fixed-scope structure protects you from open-ended costs and gives you concrete data to evaluate performance before committing to additional workflows.

    Compliance: EU AI Act and Swiss Data Protection

    The EU AI Act requires that customer support bots disclose they are AI, maintain human oversight for sensitive queries, and document the model’s training data and evaluation metrics. For Swiss firms, the Swiss Federal Act on Data Protection (revFADP) also applies to any personal data processed in the support flow. The architecture is deliberately model-agnostic: Anthropic Claude API handles quality-critical tasks where the model’s reasoning matters, while open-weight models on the client’s own hardware handle regulated data that cannot leave the building. This dual approach satisfies both quality requirements and data residency constraints. The AI agent plugs into existing CRMs, ERPs, helpdesks, and messaging through their APIs rather than replacing them, so your team continues working in the interfaces they already use. This integration approach minimizes disruption and training overhead, which is critical for a 2-week sprint.

    Pricing and Scope: What the 2-Week Sprint Delivers

    The 2-week sprint delivers a working pilot with measured baseline metrics, not a production-ready system. Full rollout across all support channels typically adds 4-6 weeks and includes additional workflows, multilingual coverage, and managed operation. The sprint cost for a 201-500 employee professional services firm ranges from EUR 15,000 to EUR 30,000 depending on integration complexity. This covers the process audit, prompt engineering, API integration, and pilot with baseline metrics. Ongoing managed operation and rollout to additional workflows are separate engagements with recurring costs based on API usage or hardware maintenance. The fixed-scope structure means you can evaluate performance before committing to full rollout. If the pilot does not meet your targets, you can adjust the scope or terminate the engagement without open-ended costs. This structure protects you from the common pitfall of AI projects that expand in scope and cost without delivering measurable results.

  • Open-Weight RAG vs. Cloud LLM APIs: Swiss Insurance Knowledge Search

    What Is Being Compared

    The two options under comparison are: (A) a retrieval-augmented knowledge assistant built on open-weight models (Llama 3 70B or Mistral 8x7B) deployed on the client’s own hardware, integrated into Microsoft Teams or Slack; and (B) the same RAG architecture but powered by OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet via their public APIs. Both options serve the same use case: internal knowledge search over policy documents, claims procedures, and regulatory updates for a 501–2,000-person insurance or insurtech firm in Switzerland. The pilot scope is identical in both cases: one workflow, four weeks, a measured before/after baseline on cycle time and error rate, and a human-in-the-loop approval layer for compliance-sensitive queries. The difference is where the model runs and what that implies for latency, cost, data residency, and accuracy.

    Criteria for Judgment

    We judge the two options against six criteria that matter for a Swiss insurance firm operating under GDPR and FINMA supervision:

    • Data residency and GDPR compliance: whether personal data or special-category data (Article 9) can leave the client’s infrastructure.
    • Latency: end-to-end response time from query to answer, measured in milliseconds.
    • Accuracy on domain-specific retrieval: measured as top-k recall on a 200-query test set drawn from the client’s actual policy documents.
    • Cost at pilot scale: total cost of ownership for the 4-week pilot, including infrastructure, API calls, and integration work.
    • Vendor lock-in: how easily the client can swap models or providers after the pilot.
    • Operational overhead: who manages model updates, prompt tuning, and pipeline maintenance during the managed operations phase.

    Comparison Table

    Criterion Option A: Open-Weight On-Premise Option B: Cloud LLM API
    Data residency All data stays on client hardware; no external transmission Data transmitted to OpenAI or Anthropic servers (US/EU regions)
    GDPR Article 32 compliance Satisfied by default; no third-party processor Requires DPA and SCCs; Article 9 data requires additional safeguards
    Latency (p95) 180–350 ms (local inference, 8x A100 or equivalent) 400–900 ms (network round-trip + inference)
    Top-k recall (200-query test) 82–88% 91–95%
    Pilot cost (4 weeks) CHF 18,000–25,000 (hardware amortized + integration) CHF 8,000–12,000 (API calls + integration)
    Vendor lock-in Low; model weights are open, pipeline is portable Medium; prompt engineering and fine-tuning tied to provider
    Operational overhead Client manages hardware; Forfaq manages pipeline Forfaq manages pipeline; client manages API keys and billing

    Scenario-by-Scenario Verdict

    When Option A wins: The client’s knowledge base contains GDPR Article 9 special-category data (health-related policy terms, claims involving medical records) or Swiss data-residency requirements mandate that no data leaves the building. In this case, the 15–30% accuracy gap is acceptable because the queries are retrieval-heavy—finding the correct policy clause or regulatory citation—rather than complex multi-step reasoning. The 180–350 ms latency is well within the 2-second threshold for a back-office agent waiting for an answer in Teams. The 4-week pilot fits because the hardware is already provisioned or the client has existing GPU infrastructure.

    When Option B wins: The knowledge base is purely internal (policy terms, claims procedures, FINMA regulatory updates) with no personal data, and the client prioritizes accuracy over data residency. The 91–95% top-k recall matters when the assistant is used for compliance review, where a missed citation has regulatory consequences. The lower pilot cost (CHF 8,000–12,000 vs. CHF 18,000–25,000) makes it attractive for a first engagement. The 400–900 ms latency is acceptable for a back-office workflow where the agent is not on a live customer call.

    Recommendation

    For a 501–2,000-person Swiss insurance firm with one process already automated and a 4-week pilot timeline, Option A (open-weight on-premise) is the recommended choice if the knowledge base includes any GDPR Article 9 data or if Swiss data-residency policy prohibits external transmission. The accuracy gap is manageable for retrieval-heavy queries, and the data-residency advantage is non-negotiable for compliance. If the knowledge base is purely internal and the client’s primary goal is reducing error rate in compliance review, Option B (cloud API) is the better fit for the pilot, with a clear migration path to on-premise if the client later expands the assistant to handle personal data. In both cases, the human-in-the-loop approval layer is mandatory, and the managed operations agreement covers pipeline maintenance, prompt updates, and a 4-hour SLA for critical issues from week 5 onward.