Tag: Switzerland

  • How a Zurich Professional Services Firm Cut Monthly Close From 14 Days to 4

    Background: A Zurich Professional Services Firm at 1,200 Headcount

    This case study is a composite based on patterns observed across multiple engagements. We do not name real clients. The firm described here is a 1,200-person professional services company based in Zurich, operating in legal, tax, and consulting. It runs a mid-market ERP, a CRM, and Microsoft Teams as its primary collaboration layer. The finance department has 14 FTEs, and the firm is ISO 27001 certified. The engagement ran over six months, from process audit through pilot to managed rollout, with a fixed-scope pilot on three workflows: invoice extraction, contract clause flagging, and monthly reporting assembly.

    The Challenge: 14-Day Close, Frozen Headcount, and a Board Deadline

    The finance director’s problem was specific: the monthly close took 14 days, and 60% of that time went to manual data entry from invoices and contracts. The firm was growing at 18% year-over-year, but the finance department had a hiring freeze. Two open requisitions sat unfilled because the budget line was tied to revenue growth that had not yet materialized. The deadline was the next quarterly board report, which required a 30% reduction in close time. The operational pressure was not hypothetical: the finance team was working 50-hour weeks during close periods, and the director had flagged burnout risk in a Q3 planning memo. The need was not to replace the finance team but to remove the repetitive extraction and entry work that consumed their time without adding analytical value.

    The Approach: n8n Orchestration, On-Prem Models, and a Fixed-Scope Pilot

    The engagement started with a two-week process audit that mapped the monthly close workflow end-to-end. The audit identified three workflows worth automating: invoice data extraction from PDFs, contract clause classification for the legal review queue, and a monthly reporting dashboard that pulled from the ERP and CRM. The pilot was scoped to these three workflows with a fixed six-week timeline. The architecture used n8n as the orchestration layer, running on the firm’s own infrastructure to satisfy ISO 27001 requirements. An open-weight model handled document extraction on-prem; an API-based model handled contract clause classification. The human-in-the-loop approval step was built into the n8n workflow as a mandatory gate for anything touching money or contracts. Slack and Microsoft Teams notifications routed approval requests to the relevant analysts.

    Outcome: 14 Days to 4, with Measured Error Rate Reduction

    The pilot measured cycle time and error rate for each workflow before and after automation. Invoice extraction dropped from 45 minutes per invoice to 8 minutes, with error rate falling from 3.2% to 0.4%. Contract clause flagging reduced review time per contract from 90 minutes to 22 minutes. The monthly reporting dashboard cut the time to assemble the board report from 3 days to 4 hours. The finance director approved the full rollout within two weeks of the pilot’s completion. The rollout extended the n8n workflows to cover the remaining invoice types and added a second contract classification category. The managed operation phase included a 30-day support window and a runbook handed to the firm’s IT team, who had prior n8n experience from an internal tooling project. The total engagement ran six months from audit to steady-state operation.

    Lessons for Similar Teams

    • Scope the pilot to three workflows, not the whole department. The fixed scope kept the six-week timeline intact and gave the finance director a clear go/no-go decision point. Trying to automate the entire close process in one pilot would have stretched the timeline and diluted the baseline metrics.
    • Run the orchestration layer on your own infrastructure if you are ISO 27001 certified. n8n on-prem satisfied the data residency requirement without requiring a separate compliance review for each model. The model-agnostic design meant that switching from an API-based model to a different one required only a connector change, not a full rebuild.
    • Build the human-in-the-loop gate into the workflow, not as a separate review step. The n8n workflow routed approval requests to Slack and Teams with a mandatory gate before data entered the ERP. This kept the compliance posture intact while still capturing the time savings.
    • Measure cycle time and error rate before and after, not just time saved. The error rate drop from 3.2% to 0.4% on invoice extraction was as valuable to the finance director as the time savings, because it reduced the risk of misstated financials in the board report.
    • Hand over to the client’s IT team with a runbook, not a managed service contract. The firm’s IT team had prior n8n experience, which reduced handover friction. A 30-day support window was enough to cover the initial stabilization period.
  • 4-Week AI Contract Review Pilot for a 15-Person Swiss E-Commerce Team

    The problem: contract review at 6.2 hours per document in a 15-person Swiss e-commerce team

    A 15-person e-commerce and retail company in Switzerland reviews vendor onboarding agreements, customer return-policy acknowledgments, and marketplace seller terms by hand. Each contract takes a median of 6.2 hours from receipt to signed approval, and 11% of contracts ship with a missed clause or an incorrect term. The legal and compliance function is a single person who also handles GDPR inquiries and tax filings. The company needs multilingual coverage across English, German, and French, and it wants to lower the cost per support ticket without adding headcount. The constraint is a 4-week fixed-scope pilot: no open-ended discovery, no multi-department rollout in the first engagement. The deliverable is a measured before/after baseline on cycle time and error rate for one contract-review workflow, plus a 12-month scaling roadmap across departments.

    Prerequisites before step 1

    • PostgreSQL 15 or later with the pgvector extension installed (CREATE EXTENSION vector;). The extension must be available on the client’s own instance; do not use a managed vector database for this pilot.
    • A contract library of at least 200 historical contracts in English, German, and French, exported as PDF or DOCX. These become the embedding index.
    • Google Workspace with API access enabled: the Drive API for document storage, the Gmail API for notifications, and the Chat API for approval workflows. The service account needs drive.file and gmail.send scopes.
    • An LLM API key for OpenAI (GPT-4o) or Anthropic (Claude 3.5 Sonnet). The key must have access to the text-embedding-3-small endpoint for the embedding step.
    • A single VM with 16 GB RAM and either an A10G GPU (24 GB VRAM) for batch embedding or a CPU-only setup if contract volume is under 500 per month.
    • One named reviewer from the legal and compliance function who will approve or reject every LLM-drafted clause during the pilot. This person must be available for 2 hours per day during weeks 3 and 4.

    Step 1: Build the pgvector contract index

    Export the 200 historical contracts from Google Drive to a local directory. Run a Python script that splits each contract into clauses using a regex on section headers (e.g., ^\d+\.\d+\s+[A-Z]). For each clause, call the text-embedding-3-small endpoint with the clause text and store the 1,536-dimensional vector in a contract_clauses table with columns id, contract_id, clause_text, embedding vector(1536), language, and created_at. The script should log the embedding latency per clause; expect 18 ms per call on a GPT-4o endpoint. After indexing, run a sanity check: embed a known clause and query the top-5 matches. If the original clause does not appear in the top-5, the index is broken and you must re-run the embedding step.

    Step 2: Wire the workflow orchestration layer

    Define the state machine in a YAML file with five states: received, embedded, drafted, awaiting_approval, and approved. The received state triggers the embedding step. The embedded state calls the LLM with the top-5 pgvector matches as context and the incoming contract clause as the query. The LLM returns a JSON object with suggested_revision, confidence_score, and flagged_terms. The drafted state sends a Google Chat message to the reviewer with the clause text, the suggested revision, and a link to the Google Doc. The awaiting_approval state pauses for 48 hours. If the reviewer approves, the state moves to approved and the contract is marked complete. If the reviewer rejects, the state returns to drafted with the reviewer’s comment appended to the LLM prompt. Log every state transition in a workflow_log table with the reviewer’s Google Workspace ID, the clause hash, and the timestamp.

    Step 3: Run the human-in-the-loop review for 10 business days

    Run the pilot on the highest-volume contract type: vendor onboarding agreements. For each incoming contract, the orchestration layer embeds the clauses, queries pgvector, and calls the LLM. The LLM drafts a revision for any clause that does not match the company’s standard template. The reviewer receives a Google Chat notification with the flagged clause and the suggested revision. The reviewer opens the contract in Google Docs, sees the flagged clause highlighted in yellow, and clicks approve or reject. The state machine records the decision. Run the pilot for 10 business days. Track three metrics per contract: cycle time (hours from receipt to approved), error rate (percentage of clauses the reviewer had to edit), and cost per ticket (LLM API cost + reviewer time × hourly rate). The baseline from the audit is 6.2 hours, 11% error rate, and CHF 42 per contract.

    Step 4: Measure cycle time, error rate, and cost per ticket

    At the end of the 10-day pilot, compare the measured metrics against the baseline. The go/no-go criteria are defined in the pilot contract: if cycle time drops by at least 50% (to 3.1 hours or less) and error rate drops by at least 40% (to 6.6% or less), the client proceeds to rollout. If either criterion is not met, the pilot is extended by 5 business days with a revised LLM prompt or a different embedding model. The measurement report includes a per-clause breakdown: which clause types the LLM handled well (e.g., payment terms, liability caps) and which still require human review (e.g., IP assignment, termination clauses). The report also includes the cost per ticket for the pilot period and a projection for 12 months at the current contract volume. The 12-month scaling roadmap identifies the next two workflows to automate: customer return-policy acknowledgments and marketplace seller terms.

    Common pitfalls and how to detect them

    • Embedding drift: if the contract template changes (e.g., a new liability clause is added), the pgvector index becomes stale. Detect this by running a weekly job that embeds the current template and compares it against the index. If the top-5 match score drops below 0.82, re-index the affected clauses.
    • Reviewer bottleneck: if the reviewer does not respond within 48 hours, the workflow stalls. Detect this by monitoring the awaiting_approval state duration. If the median wait exceeds 36 hours, escalate to the team lead via a Gmail API email.
    • Language misclassification: if a German contract is misclassified as English, the LLM may produce a low-quality draft. Detect this by logging the detected language per contract and flagging any contract where the detected language does not match the contract’s metadata field.
    • LLM hallucination: if the LLM invents a clause that does not exist in the contract library, the reviewer will reject it. Detect this by logging the confidence_score and flagging any draft with a score below 0.70 for manual review before it reaches the reviewer.
  • Cutting First-Response Time in Swiss Medtech: A 6-Month AI Integration Sprint

    The Problem: First-Response Time in a 250-Person Medtech Firm

    A 250-person medtech company in Zurich runs its lead pipeline on a CRM that was configured in 2019. Leads arrive from trade-show badges, partner referrals, and web forms. The sales team manually qualifies each lead, enriches missing fields (company size, regulatory context, product interest), and logs the outcome. The average first-response time is 4.2 hours. The error rate on field-level data is 18%—missing, malformed, or inconsistent values that force a second pass. The company wants to cut first-response time without adding headcount. The constraint is not the model; it is the integration. The CRM exposes a custom REST API and webhook endpoints, but the data model is inconsistent, and the qualification logic is tribal knowledge in three sales reps’ heads. The audit must surface that logic before any automation can be built. The pilot must run on live data with a measured baseline, not a synthetic dataset. The rollout must not replace the CRM; it must plug into it through the existing API layer.

    How the LangGraph Pipeline Works

    The pipeline is a LangGraph stateful graph with four nodes: Fetch, Enrich, Qualify, and Write. The Fetch node calls the CRM’s GET /leads/{id} endpoint and loads the raw record into the graph state. The Enrich node runs a conditional branch: if the company size field is missing, it calls an external data provider API; if the regulatory context is missing, it queries the company’s internal documentation via a retrieval-augmented generation (RAG) call. The Qualify node sends the enriched record to an LLM (OpenAI gpt-4o or Anthropic claude-3-5-sonnet) with a structured prompt that outputs a JSON object: {"score": 0-100, "reason": "...", "fields_to_fix": [...]}. The Write node calls PATCH /leads/{id} to update the enriched fields and POST /leads/{id}/qualification to set the score. A webhook on the CRM fires on status change, which triggers the next pipeline run if the lead is re-submitted. The graph state persists between nodes, so a failed enrichment call does not lose the qualification context. The entire pipeline runs in under 3 seconds for a typical record.

    Trade-offs: Model Choice, Human-in-the-Loop, and Integration Scope

    Three architectural choices drive the cost and risk profile. First: model selection. OpenAI and Anthropic APIs are used for the qualification and enrichment steps because their classification and extraction quality is higher than open-weight models at the same latency. The cost is approximately EUR 0.02-0.05 per lead, which is negligible at 250-person scale. If the client later extends the system to handle patient-adjacent data, the same LangGraph pipeline can be pointed at an open-weight model (Llama 3 70B or Mistral 8x7B) running on the client’s own hardware. The API layer is abstracted, so the switch is a configuration change. Second: human-in-the-loop. The AI drafts the qualification score and enriches fields, but the final status requires a human click. This adds 30-60 seconds per record to the approval queue, but it preserves accountability for any record that touches a contract or pricing. Third: integration scope. The sprint touches only the CRM’s REST API and webhook endpoints. It does not modify the CRM’s data model, does not replace the helpdesk, and does not build a new frontend. The scope is fixed: one workflow, one CRM, one set of endpoints.

    Recommendation: The 6-Month Integration Sprint

    The 6-month timeline is fixed-scope and non-negotiable. Month 1-2: Process audit. Map the lead flow, identify data gaps, quantify manual effort, and capture the baseline: average first-response time (4.2 hours), field-level error rate (18%), and manual hours per 100 leads. The output is a prioritized roadmap with one pilot workflow selected. Month 3-4: Integration sprint. Connect the custom REST API and webhook endpoints to the LangGraph pipeline. Build the enrichment and qualification logic. Run unit tests on the API layer. Month 5: Pilot. Run the pipeline on a live lead stream. Human-in-the-loop approval for any record that touches a contract or pricing. Measure against the baseline. Month 6: Rollout and handoff. Extend the pipeline to the full lead stream. Document the handoff to managed operation. Run a 30-day hypercare period. The timeline assumes the CRM API is stable. If the CRM is mid-migration, add 2-3 weeks to the sprint phase. The pilot ships with a one-page summary: before/after cycle time, error rate, and the raw data attached for verification.

  • Cut First-Response Time in a Swiss Healthcare Company: A 3-Month AI Pilot

    1. Pick the highest-volume, lowest-complexity workflow first

    The first workflow to automate is the one with the highest volume and the lowest complexity. For a 100-person healthcare and medtech company in Switzerland, that is almost always order and shipment status updates. The operations team receives 40 to 60 inquiries per day from hospitals, clinics, and distributors asking where an order is. Each inquiry requires a human to log into SAP or Microsoft Dynamics, check the order status, and draft a response. The average first-response time is 4 to 6 hours. The error rate is 8 to 12 percent because humans copy data from the ERP into the response and make transcription mistakes. This workflow is the ideal first pilot because it is high-volume, low-complexity, and the data is structured. The AI reads the ERP directly, so there is no transcription step. The response is a template with the order number, the status, and the expected delivery date. The human approval gate is simple: if the status is ‘shipped’ or ‘delivered’, the AI sends the response automatically. If the status is ‘delayed’ or ‘exception’, a human reviews it. This single workflow, automated, cuts the first-response time from 4 hours to 60 seconds and the error rate to under 2 percent.

    2. Integrate with the ERP through its native API, not a custom connector

    The AI layer does not replace the ERP. It reads order and shipment records through the SAP or Dynamics API, classifies the status, and writes the response back to the helpdesk or messaging channel. The ERP remains the system of record for inventory, billing, and shipping. The AI orchestration layer sits between the ERP and the customer-facing channel, handling the translation and the human approval gate. No data is duplicated; the AI reads and writes through the existing API endpoints. The integration is built on the ERP’s standard API, not a custom connector. For SAP, that is the OData API or the BAPI layer. For Microsoft Dynamics, that is the Web API or the Business Central API. The integration is tested against the client’s staging environment before it goes live. The client’s IT team provisions the API credentials and the network access in the first two weeks. The Forfis team builds the orchestration layer in the next four weeks. The result is a system that plugs into the existing infrastructure without replacing it.

    3. Run open-weight models on-premise to keep PHI inside the building

    The model-agnostic architecture means Forfis can use OpenAI or Anthropic APIs for tasks where quality matters and the data is not regulated, and open-weight models on the client’s hardware for tasks where regulated data cannot leave the building. For a Swiss healthcare company, the order status workflow uses open-weight models on-premise because the ERP contains patient identifiers. The model never sees raw patient identifiers; the orchestration layer strips PHI before the prompt is constructed. The model’s output is a structured JSON object with a status code and a template ID, not free text. A human operator reviews any output that triggers an exception rule before it is sent. This architecture satisfies HIPAA’s minimum necessary standard and Swiss FADP Article 6(2) on data minimization. The client’s IT team provisions a single A100 or H100 GPU server in the first two weeks. The Forfis team fine-tunes the model on the client’s historical order data in the next four weeks. The model runs on the client’s hardware, so no data leaves the building.

    4. Build the human-in-the-loop approval gate into the existing helpdesk

    The AI drafts the response, but a human approves anything that touches money, health data, or a contract. For order status updates, the approval rule is simple: if the status is ‘shipped’ or ‘delivered’, the AI sends the response automatically. If the status is ‘delayed’, ‘returned’, or ‘exception’, a human reviews and approves before the response goes out. The approval queue is integrated into the existing helpdesk, so the operations team does not need a new tool. The human-in-the-loop design is not a compromise; it is the default. The model is a draft, not a decision. The human is the decision-maker. This design reduces the risk of a bad response going out, and it builds trust with the operations team. The approval rate for ‘shipped’ and ‘delivered’ statuses is 95 to 98 percent, so the human only reviews the 2 to 5 percent of responses that are exceptions. The average approval time is 30 to 60 seconds. The total first-response time, from inquiry to response, is under 2 minutes.

    5. Measure the before/after baseline in the first two weeks

    The pilot ships with a measured baseline: the average first-response time and error rate before automation, and the same metrics after. For a 100-person healthcare company, the typical baseline is 4 to 6 hours for a human to check the ERP and draft a response. After automation, the AI drafts the response in under 2 seconds, and a human approves it in 30 to 60 seconds. The error rate drops from 8 to 12 percent to under 2 percent because the AI reads the ERP directly rather than relying on a human to copy data correctly. The baseline is measured in the first two weeks of the pilot, before the AI is live. The after-metrics are measured in the last two weeks, after the AI has been running for at least four weeks. The client gets a one-page report with the before/after numbers, the error rate breakdown, and the approval rate. This report is the basis for the decision to scale to the next workflow. The 3-month timeline is realistic because the scope is fixed and the metrics are measured from day one.

    6. Fix the scope and the price before the pilot starts

    The pilot is fixed-scope and fixed-price. The scope is defined in the contract: the number of API endpoints, the number of response templates, and the approval rules. The cost covers the process audit, the integration with the ERP, the build of the orchestration layer, the model fine-tuning, and the 3-month managed operation. The client’s cost is the GPU hardware for the on-premise model, typically a single A100 or H100 server, and the time of the operations lead and IT contact. For a 100-person company, the total cost of the pilot is typically in the range of EUR 40,000 to EUR 60,000, depending on the complexity of the ERP integration. The fixed-scope model prevents scope creep. If the client wants to expand to shipment tracking or returns, that is a second pilot with its own scope and timeline. The 3-month timeline is realistic because the scope is fixed and the team is dedicated. The client does not need to hire new staff; the existing operations team handles the approval queue, and the IT team provisions the hardware and the API credentials.

    7. Ship the pilot as a measured baseline, not a transformation

    The pilot is one workflow, not a transformation. The operations team still handles the exceptions, the escalations, and the complex inquiries. The AI handles the 80 to 90 percent of inquiries that are routine status checks. The human-in-the-loop approval gate ensures that the AI does not make a mistake that a human would have caught. The model-agnostic architecture means the client is not locked into one vendor; if the open-weight model is not good enough, Forfis can switch to a commercial API for the non-PHI tasks. The integration with the ERP means the client does not need to replace its system of record. The 3-month timeline is realistic because the scope is fixed and the team is dedicated. The result is a measurable reduction in first-response time and error rate, with no new hires and no new tools. The operations team gets its time back for the work that actually requires a human.

  • RAG Assistant for B2B SaaS: 4-Week GDPR-Compliant Rollout in Switzerland

    The Problem: Routine Work Consuming Senior Staff Time

    A 20-person B2B SaaS company in Switzerland faces a common problem: senior staff spend too much time on routine tasks, such as answering order and shipment status queries. This reduces their capacity for high-value work, such as product development and strategic account management. The solution is a Retrieval-Augmented Generation (RAG) assistant that can handle these routine queries autonomously. The assistant retrieves relevant documents from a vector database and uses them to ground the LLM’s response, ensuring accuracy and reducing hallucinations. The goal is to free up senior staff from routine work, allowing them to focus on complex issues. This deep dive explores how to implement such a system in 4 weeks, using pgvector for embeddings search and integrating with Google Workspace.

    Mechanism: How the RAG Assistant Works

    The RAG assistant works by retrieving relevant documents from a vector database and using them to ground the LLM’s response. The process starts with ingesting documents, such as order records, shipment logs, and policy documents. These documents are split into chunks, and each chunk is converted into an embedding using a model like OpenAI’s text-embedding-3-small. The embeddings are stored in pgvector, a PostgreSQL extension that enables vector similarity search. When a user asks a question, the question is also converted into an embedding, and the vector database retrieves the most similar chunks. These chunks are then passed to the LLM, which uses them to generate a response. The LLM is prompted to use only the retrieved chunks, reducing the risk of hallucination. The response is then sent to the user via Google Workspace, such as Gmail or Chat.

    Trade-offs: Model Choice and Data Privacy

    The main trade-off is between using a third-party API (like OpenAI) and an open-weight model on your own hardware. Third-party APIs offer higher quality and lower maintenance but raise GDPR concerns due to data leaving your control. Open-weight models (like Llama 3 or Mistral) can run on your own hardware, ensuring data stays in Switzerland, but require more technical expertise and may have lower quality. For a small company, a hybrid approach is often best: use third-party APIs for non-sensitive tasks and open-weight models for sensitive data. Another trade-off is between accuracy and speed. More complex retrieval strategies, such as hybrid search (combining vector and keyword search), improve accuracy but increase latency. For a 20-person company, a simple vector search is often sufficient.

    Recommendation: A 4-Week Implementation Plan

    Week 1: Conduct a process audit to identify high-volume, low-complexity tasks. Define success metrics: cycle time, error rate, and customer satisfaction. Build a baseline by measuring current performance. Week 2: Ingest data, generate embeddings, and set up pgvector. Test the retrieval process to ensure accuracy. Week 3: Integrate with Google Workspace and test the assistant with internal users. Refine prompts and data sources based on feedback. Week 4: Conduct user acceptance testing and GDPR compliance checks. Hand over the system to the client and provide training. This timeline assumes the client has clean, accessible data and dedicated staff available for interviews and testing. If data quality is poor, additional time may be needed for cleaning and preprocessing.

  • Swiss Fintech AI Automation: A 6-Month Sprint to Cut Back-Office Cycle Time

    The Back-Office Bottleneck in Swiss Fintech Operations

    A 300-person fintech in Zurich processes roughly 12,000 payment instructions and 4,500 support tickets per month. The operations team of 48 people spends an estimated 3,200 hours monthly on data entry, document re-keying, and first-response triage. The cost is not just the salary bill; it is the cycle time. A payment instruction received at 09:00 often does not reach the ERP until 14:30, and a support ticket in German or French waits 4 to 6 hours for a first response. The company has tried adding headcount twice in the last 18 months, but the volume grew faster than the team. The constraint is not talent availability in the Swiss market; it is the structural mismatch between linear headcount growth and sub-linear process improvement.

    The question is not whether to adopt AI. The question is which workflows to automate first, how to integrate them into the existing SAP or Dynamics ERP without a rip-and-replace, and how to measure whether the automation actually reduced cycle time and error rate rather than just shifting work to a different queue. A 6-month integration sprint is the right scope: long enough to run a real pilot with a before/after baseline, short enough to avoid the scope creep that kills most AI projects in the second quarter.

    The LangGraph Pipeline: From Raw Document to ERP Post

    The pipeline has five stages. First, document ingestion pulls PDFs, emails, and scanned images from the existing intake channels. Second, OCR and field extraction uses a multilingual LLM to identify and extract structured fields: payer name, IBAN, amount, currency, reference number, and date. The extraction prompt is version-controlled and includes few-shot examples in German, French, and Italian. Third, validation checks the extracted fields against business rules: IBAN format per ISO 13616, amount range, currency code per ISO 4217. Fourth, routing sends high-confidence extractions directly to the ERP via the OData API and flags low-confidence ones for human review. Fifth, human-in-the-loop approval presents the flagged items in a queue with the AI’s suggested values pre-filled; the reviewer confirms or corrects and the system logs the override.

    For ticket triage, the graph is simpler: classification assigns the ticket to a category (payment dispute, onboarding, technical issue, regulatory inquiry), language detection tags the ticket, and routing sends it to the appropriate queue. The LangGraph state object carries the ticket text, detected language, assigned category, and confidence score. Conditional edges route regulatory inquiries directly to a senior agent, bypassing the AI entirely. The entire graph is defined in Python and version-controlled in Git, so every change to the routing logic is auditable.

    Model-Agnostic Architecture and the On-Premises Question

    The first trade-off is model choice. OpenAI’s GPT-4o and Anthropic’s Claude 3.5 Sonnet handle multilingual extraction well, but the data leaves the client’s infrastructure. For a fintech in Switzerland, even without a specific regulatory mandate, the data residency question is real. The alternative is an open-weight model like Llama 3.1 70B or Mistral Large running on the client’s own GPU hardware. The open-weight model costs roughly EUR 18,000 to 25,000 in initial hardware and EUR 2,000 to 3,500 per month in electricity and maintenance, but it keeps all data on-premises. The quality gap for structured extraction is small; for nuanced ticket classification, the proprietary models still edge ahead by 3 to 5 percent on F1 score.

    The second trade-off is integration depth. A shallow integration reads from the ERP and writes back via the OData API. A deep integration embeds the AI layer inside the ERP’s workflow, which requires custom ABAP or X++ development. The shallow approach is faster to ship and easier to maintain, but it adds 200 to 400 milliseconds of latency per API call. For a batch process running at 02:00, that latency is irrelevant. For a real-time ticket triage, it matters. The recommendation is shallow integration for document extraction and a hybrid approach for ticket triage, where the AI layer runs as a microservice in front of the helpdesk API.

    Human-in-the-Loop as the Quality Gate, Not the Fallback

    The human-in-the-loop step is not a fallback; it is the primary quality gate. The threshold for automatic approval is set per field. For payment instructions, the IBAN and amount fields require a confidence score of 0.95 or higher; the payer name field requires 0.90. Below the threshold, the item goes to the review queue. The reviewer sees the AI’s suggested values, the source document, and the confidence scores. They confirm, correct, or reject. Every override is logged with the reviewer’s ID, timestamp, and the correction made.

    This log is the training data for the next iteration. After four weeks of operation, the override log contains 800 to 1,500 corrections. These are used to refine the extraction prompt, add new few-shot examples, or adjust the confidence thresholds. The system does not retrain the base model; it adjusts the prompt and the validation rules. This is faster, cheaper, and more auditable than fine-tuning. The human-in-the-loop step also serves as the audit trail: every automated decision is traceable to a human approval or a confidence threshold, which matters when a payment instruction is disputed six months later.

    The 6-Month Sprint: Phases, Gates, and Exit Criteria

    The 6-month sprint breaks into four phases. Phase 1 (weeks 1 to 6): Process audit and baseline. The team maps the current manual workflow step by step, samples 300 real transactions over two weeks, and measures cycle time, error rate, and cost per transaction. The output is a prioritized list of workflows ranked by volume, error cost, and data availability. The client selects one workflow for the pilot.

    Phase 2 (weeks 7 to 14): Pilot on one workflow. The LangGraph pipeline is built, tested against the sample data, and run in shadow mode alongside the existing manual process. The before/after baseline is measured on the same 300 transactions. The pilot must show a 40 percent or greater reduction in cycle time and a 25 percent or greater reduction in error rate to proceed.

    Phase 3 (weeks 15 to 22): Second workflow and ERP integration. The second workflow is added, and the OData integration with SAP or Dynamics is built and tested. The multilingual coverage is validated on real German, French, and Italian documents.

    Phase 4 (weeks 23 to 26): Monitored rollout. The system goes live with daily error-rate reviews, a 24-hour rollback plan, and a weekly report to the operations lead. The final deliverable is a measured before/after report with the raw data, so the client can verify the numbers independently.

    Pitfalls That Kill the Sprint and How to Avoid Them

    The most common failure mode is scope creep in the pilot phase. The client wants to automate three workflows instead of one, or add a new integration with a third-party payment provider mid-sprint. The fix is contractual: the pilot scope is fixed at the start of Phase 2, and any change triggers a change order with a revised timeline. The second failure mode is insufficient sample data. If the client cannot provide 300 clean, labeled examples of the target workflow, the baseline is unreliable and the pilot results are meaningless. The fix is to start the data collection in week 1, not week 5.

    The third failure mode is ERP API access delays. SAP and Dynamics API access requires security reviews, firewall changes, and sometimes custom development. If the API is not available by week 10, the pilot cannot run in shadow mode and the timeline slips. The fix is to request API access in the first week of the engagement and assign a dedicated ERP administrator on the client side. The fourth failure mode is multilingual edge cases. German compound nouns, French abbreviations, and Italian date formats break extraction models that were trained primarily on English. The fix is to include language-specific few-shot examples in the prompt from day one and to test on real multilingual documents, not synthetic ones.

  • AI Agent Development in Insurance: A Glossary

    AI Agent Development

    AI agent development refers to the design and deployment of autonomous software systems that perform specific tasks, such as classifying customer inquiries or extracting data from documents. In insurance, these agents are typically built using frameworks like LangChain and LangGraph, integrated with existing systems via APIs, and operated with human-in-the-loop oversight to ensure compliance and accuracy. The goal is to automate routine work, freeing senior staff to focus on high-value activities.

    Running Isolated Pilots

    Running isolated pilots is a strategy for managing AI maturity by deploying automation in a controlled, limited scope before broader rollout. This approach allows the organization to establish baseline metrics for cycle time and error rates, validate GDPR compliance, and refine the model without disrupting core operations. It is a standard practice for large enterprises, ensuring that the AI system is reliable and compliant before scaling.

    LangChain and LangGraph

    LangChain is a framework for building applications that use large language models, providing abstractions for prompts, memory, and tool use. LangGraph extends this by allowing developers to define stateful, multi-step workflows as graphs, which is essential for complex insurance processes like claims adjudication that require conditional logic and human-in-the-loop approvals. Together, they enable the construction of robust, scalable AI agents.

    Data Enrichment and Cleanup

    Data enrichment involves augmenting raw customer or claim records with external data sources, such as credit scores or vehicle history, to improve decision-making. Cleanup refers to standardizing inconsistent formats, removing duplicates, and correcting errors in existing datasets. For a 2,000+ employee insurer, this ensures that AI agents operate on high-quality, GDPR-compliant data, reducing the risk of errors and non-compliance.

    Scaling Operations Without New Hires

    Scaling operations without new hires involves using AI automation to handle increased workloads, such as a surge in insurance claims, without proportional increases in headcount. By automating routine tasks like ticket triage and data entry, the organization can maintain service levels and reduce operational costs while freeing senior staff to focus on strategic initiatives. This approach is particularly valuable for large enterprises managing growth and efficiency.

    Operations and Supply Chain

    Operations and supply chain in insurance refer to the back-office processes that support policy administration, claims processing, and customer service. These functions are often labor-intensive and prone to errors, making them ideal candidates for AI automation. By integrating AI agents with existing CRMs and ERPs, insurers can streamline these processes, reduce cycle times, and improve data accuracy, ultimately enhancing customer satisfaction and operational efficiency.

  • Swiss Fintech Cuts Candidate Screening Cost 78% with On-Prem AI in 4 Weeks

    Background: A 300-Person Swiss Payments Firm Stuck in Pilot Limbo

    This case study is a composite built from patterns Forfis has observed across multiple engagements in Swiss fintech and payments. No named customer is represented. The company described here is a mid-size payments processor in Zurich, roughly 300 employees, operating in the Running Isolated Pilots stage of AI maturity. It runs a standard on-prem ERP, a mid-market ATS, and Google Workspace as its primary collaboration suite. The team had tried two earlier AI pilots in 2023, both scoped to marketing copy generation, and had not moved past the pilot phase. The CTO wanted a third attempt that would actually change a cost line, not just produce a demo. The constraint was non-negotiable: candidate data could not leave the building, and the solution had to work inside the tools the recruiting team already used.

    Challenge: 120 Applications a Month, 14 Minutes Each, and a Q3 Deadline

    The recruiting team of six handled roughly 120 applications per month across four open roles. Each application required a recruiter to read the CV, extract key fields, compare them against the role criteria, and write a short assessment. The average time per application was 14 minutes, and the monthly reporting cycle for the CTO’s ops dashboard took two full days of manual spreadsheet work. The cost per processed application, loaded with recruiter salary and overhead, sat around CHF 18. The team was not understaffed in absolute terms, but the volume was growing 15% quarter-over-quarter as the firm expanded into new payment corridors. The CTO’s deadline was the end of Q3: a working pilot that reduced the cost per ticket and the monthly reporting effort, delivered in four weeks, with GDPR compliance documented before any candidate data was touched.

    Approach: Four-Week Integration Sprint with an On-Prem Open-Weight Model

    Forfis ran a one-week process audit that mapped the screening workflow end to end: application intake from the ATS, CV parsing, field extraction, criteria matching, recruiter review, and the monthly report. The audit confirmed that 70% of the recruiter’s time went to extraction and initial scoring, not to judgment calls. The pilot scope was fixed: build a document and data extraction pipeline that ingests CVs from the ATS, runs them through an open-weight model on the client’s own A100 GPU node, scores each application against weighted criteria, and writes the result back to the ATS and into a Google Docs template for the recruiter’s review. The model was a 7B-parameter Llama 3.1 8B fine-tuned on the client’s historical screening decisions. No candidate data left the building. The integration sprint ran four weeks: audit and baseline in week one, pipeline build in week two, shadow test in week three, and human-in-the-loop approval workflow plus handover in week four.

    Outcome: 79% Less Time per Application, 78% Lower Cost per Ticket

    The pilot processed 340 applications over a six-week shadow period, compared to the 120 the team handled manually in the same window. The model agreed with the recruiter’s accept/reject decision on 89% of cases. On the 11% where it disagreed, a structured review found the model was correct in 4 of 12 cases, the recruiter in 7, and 1 was genuinely ambiguous. The error rate on structured field extraction was 2.3% across 340 documents, down from the 8% baseline of the previous manual process. The recruiter’s manual time per application dropped from 14 minutes to 3 minutes for review, a 79% reduction. The cost per processed application fell from roughly CHF 18 to CHF 4, a 78% reduction, before accounting for the one-time GPU hardware cost. The monthly reporting cycle, which had taken two days of spreadsheet work, was reduced to a 20-minute review of an auto-generated summary in Google Docs. The CTO’s Q3 deadline was met on the fourth Friday.

    Lessons for Teams Running Isolated Pilots in Regulated Fintech

    • Baseline before you build. The 8% manual error rate and the 14-minute cycle time were measured in week one, not assumed. Without that baseline, the 2.3% and 3-minute results would have been unprovable. Every pilot should ship with a measured before/after on cycle time and error rate.
    • On-prem is not a technical constraint, it is a compliance constraint. The client’s DPO required a documented data flow map before any candidate data was processed. The one-page diagram showing that all data stayed on the A100 node and that no external API calls were made was the single most important artifact in the engagement. GDPR Article 35 DPIA updates were handled in week one, not after the model was built.
    • Integrate into the tools the team already uses. The recruiter’s review happened in a Google Docs template linked from a Gmail notification. No new dashboard, no new login. The adoption rate was 100% because the workflow lived inside the tools the team already used every day.
    • Human-in-the-loop is not optional for regulated data. Every candidate decision required a recruiter’s approval. The model drafted, ranked, and flagged; the person decided. This satisfied both the GDPR accountability requirement and the team’s trust threshold.
    • Fixed scope, four weeks, one workflow. The pilot touched one workflow, one model, one integration point. The CTO’s Q3 deadline was met because the scope was fixed in week one and did not expand.
  • Cut HR First-Response Time in a 2,000+ B2B SaaS Company: A Two-Week RAG Pilot

    The Problem: HR First-Response Time in a 2,000+ Employee B2B SaaS Company

    A 2,000+ employee B2B SaaS company in Switzerland runs HR and recruiting operations on a mix of Confluence, Notion, and a helpdesk. Employees ask the same 40 questions every week: how to request PTO, how to file an expense report, how to access the staging environment. The current first-response time is 4–6 hours because the answer lives in a Confluence page that no one can find quickly. The goal is to cut first-response time to under 10 minutes by building a retrieval-augmented knowledge assistant that searches the company’s own documentation and returns a sourced answer. The pilot runs for two weeks, uses Anthropic Claude API for generation, and ships with ISO 27001-compliant access controls and audit logging. The delivery model is managed AI operations: Forfis builds, deploys, and monitors the system, and the client’s team owns the content and the feedback loop.

    Prerequisites Before Step 1

    • Knowledge base access: API credentials for Confluence or Notion, with read access to the relevant workspaces. Confirm the workspace contains the 40 most-asked questions.
    • Anthropic API key: A production key with usage limits set. Store it in a secrets manager (HashiCorp Vault, AWS Secrets Manager, or GCP Secret Manager), not in code.
    • Communication channel: Slack or Microsoft Teams workspace where employees ask questions. Confirm webhook or API access is available.
    • Vector database: A managed instance (Pinecone, Weaviate, or pgvector on Postgres) with sufficient capacity for the knowledge base size. For a 2,000+ employee company, expect 5,000–20,000 documents.
    • ISO 27001 documentation: Access control policies, audit logging requirements, and data retention rules. The pilot must comply with these before go-live.
    • Baseline data: A one-week log of HR questions, current first-response times, and resolution rates. This is the before/after measurement point.

    Step 1: Ingest the Knowledge Base

    Export all relevant Confluence or Notion pages to a structured format. Use the Confluence REST API (/rest/api/content?spaceKey=HR) or the Notion API (/v1/databases/{database_id}/query) to pull pages. Store the output as JSON files in a staging directory. Each document should include: id, title, body (Markdown), last_updated, and owner. For a 2,000+ employee company, expect 5,000–20,000 pages. Filter out pages marked as deprecated or restricted. The ingestion script should run in under 30 minutes for a typical workspace. Log the number of pages ingested and any errors to a CSV file for the audit trail.

    Step 2: Build the Retrieval Pipeline

    Split each document into chunks of 256–512 tokens, with a 50-token overlap. Use a semantic chunking strategy: split on headings first, then on paragraphs. For each chunk, generate an embedding using the text-embedding-3-small model (OpenAI) or bge-large-en (open-weight, if the data cannot leave the building). Store the embeddings in the vector database with metadata: document_id, chunk_index, title, last_updated. For a 10,000-document knowledge base, expect 50,000–100,000 chunks. The indexing process should take under 2 hours on a managed vector database. Verify the index by running 10 test queries and confirming that the top-5 results are relevant.

    Step 3: Configure the Generation Layer

    Configure the Anthropic Claude API call with the following parameters: model: claude-sonnet-4-20250514, max_tokens: 1024, temperature: 0.2. The system prompt should instruct the model to answer only from the retrieved context, cite the source document, and say “I don’t know” if the answer is not in the context. The user prompt should include: the employee’s question, the top-5 retrieved chunks (with titles and URLs), and a request for a concise answer with a source link. Test the pipeline with 20 real questions from the baseline log. Measure: (1) retrieval precision (are the top-5 chunks relevant?), (2) generation accuracy (is the answer correct?), (3) latency (should be under 3 seconds end-to-end). Iterate on the chunking and prompt until accuracy is above 80%.

    Step 4: Integrate with the Communication Channel

    Integrate the assistant with Slack or Microsoft Teams. In Slack, create a custom app with a /ask slash command. The command sends the question to the RAG pipeline, waits for the response, and posts it back to the channel. In Teams, use a bot framework (Microsoft Bot Framework) with a similar flow. The response should include: the answer, a link to the source document, and a feedback button (thumbs up/down). The feedback button sends a structured event to a logging endpoint. For ISO 27001 compliance, log every query with: timestamp, user_id, question, retrieved_chunks, model_response, feedback. Store the logs in a read-only database with a 12-month retention policy. Restrict access to the assistant via SSO: only authenticated employees can use it.

    Step 5: Run the Two-Week Pilot

    Run the pilot for two weeks with a defined scope: one department (HR or recruiting), one knowledge source (Confluence or Notion), one channel (Slack or Teams). Track five metrics daily: (1) first-response time (target: under 10 minutes, baseline: 4–6 hours), (2) resolution rate (target: 70%, baseline: 30–40%), (3) accuracy (target: 80%, measured by user feedback), (4) retrieval precision (target: 85%, measured by manual review of 50 queries), (5) user satisfaction (target: 4/5, measured by post-answer rating). At the end of week two, produce a report with: before/after metrics, a list of the top 10 unanswered questions, and a recommendation for rollout. The report should be reviewed by the client’s HR lead and the Forfis delivery team.

  • OpenAI API vs On-Prem Models for a Swiss Fintech Pilot

    What Is Being Compared

    The two options under comparison are the OpenAI API as a hosted inference service and an open-weight model running on the client’s own hardware. The OpenAI API is a managed service where prompts are sent over HTTPS and completions are returned; the client does not manage the model weights or the inference infrastructure. The on-prem option uses a model such as Llama 3 or Mistral, deployed on the client’s servers or a private cloud, where the model weights are downloaded and the inference runs locally. Both options can serve the same two workflows: a retrieval-augmented knowledge assistant over Confluence or Notion, and a ticket triage and routing system for the helpdesk. The comparison is framed for a Swiss fintech with 501 to 2000 employees, operating under ISO 27001, with a two-week fixed-scope pilot as the delivery vehicle. The goal is to free senior staff from routine work in operations and supply chain, specifically by reducing manual back-office tasks and automating first-response triage.

    Criteria for the Comparison

    The evaluation uses seven criteria that matter to a Swiss fintech under ISO 27001. Latency is measured as the time from prompt submission to first token, which affects the user experience in a RAG assistant. Cost per unit is the total expense per ticket triaged or per document extracted, including API fees, compute, and human review time. Data residency is whether the data leaves the client’s network, which is a hard constraint for payment data under FINMA guidance. Compliance fit is how well the option aligns with ISO 27001 controls, particularly access control, logging, and data processing agreements. Integration effort is the number of API calls and configuration steps needed to connect to Confluence, Notion, and the helpdesk. Model quality is measured on a defined evaluation set of 200 tickets and 100 documents, scored by a human reviewer. Vendor lock-in is the cost and effort of switching to a different model or provider after the pilot. Each criterion is scored in the table below with concrete numbers where available.

    Comparison Table

    Criterion OpenAI API On-Prem Open-Weight Model
    Latency (first token) 180 to 400 ms over HTTPS 50 to 150 ms on local GPU
    Cost per ticket triaged 0.02 to 0.05 USD per ticket 0.005 to 0.02 USD per ticket after amortized hardware
    Data residency Data leaves client network to OpenAI infrastructure Data stays on client hardware
    ISO 27001 fit Requires DPA and data flow documentation Easier to document; no external data transfer
    Integration effort 3 to 5 API calls; standard HTTPS 8 to 12 steps; requires GPU provisioning and model loading
    Model quality (200-ticket eval) 92 percent accuracy on triage 85 to 88 percent accuracy on triage
    Vendor lock-in Low; prompt templates are portable Low; model weights are open, but inference stack is tied to hardware

    The latency difference is small for batch processing but noticeable in a live RAG assistant where the user is waiting for a response. The cost difference is significant at scale: for 10,000 tickets per month, the OpenAI API costs 200 to 500 USD, while the on-prem model costs 50 to 200 USD after the initial hardware investment. The data residency row is the deciding factor for a fintech handling payment data.

    When the OpenAI API Wins

    For ticket triage and routing, the OpenAI API wins on quality and speed of deployment. The 92 percent accuracy on the 200-ticket evaluation set means fewer misroutes, which directly reduces the time senior staff spend correcting errors. The 180 to 400 ms latency is acceptable for a triage system where the user is not waiting for a real-time response; the ticket is routed asynchronously. The integration effort is lower: three to five API calls to the helpdesk and the OpenAI endpoint, with no GPU provisioning. For a two-week pilot, this means the team can focus on the classification logic and the human-in-the-loop approval step rather than on infrastructure setup. The cost of 0.02 to 0.05 USD per ticket is negligible at the pilot scale of a few hundred tickets.

    When the On-Prem Model Wins

    For the retrieval-augmented knowledge assistant over Confluence or Notion, the on-prem model is the stronger choice when the indexed documents contain payment data, customer identifiers, or internal financial records. The data residency constraint is non-negotiable: FINMA guidance for Swiss fintechs requires that personal data and payment data be processed within the client’s control. The on-prem model keeps the embeddings and the prompts on the client’s hardware, so no data leaves the building. The 50 to 150 ms latency is faster than the OpenAI API, which improves the user experience in a live assistant. The 85 to 88 percent accuracy is lower than the OpenAI API, but for a RAG assistant the quality is more dependent on the retrieval step than on the model itself. The integration effort is higher, requiring GPU provisioning and model loading, but this is a one-time setup that pays off over the life of the assistant.

    Recommendation for the Swiss Fintech Pilot

    The recommendation is a hybrid architecture that uses the OpenAI API for ticket triage and the on-prem model for the RAG assistant. This split is driven by the data residency constraint: ticket data in a helpdesk is less sensitive than the financial documents in Confluence, so the OpenAI API is acceptable for triage. The RAG assistant indexes Confluence and Notion, which contain internal financial records and payment data, so the on-prem model is required. The model-agnostic architecture means the application layer is decoupled from the model provider, so the team can swap models without re-implementing the business logic. The two-week pilot should deliver a measured baseline for both workflows: cycle time and error rate for ticket triage, and retrieval accuracy and response quality for the RAG assistant. The pilot should also include a data flow diagram that maps exactly which fields go to the OpenAI API and which stay on the client’s hardware, satisfying the ISO 27001 documentation requirement.