Author: Forfis

  • AI Assistant for Austrian Insurance: Fixed-Scope Pilot with EU AI Act Compliance

    Process Audit and Baseline Measurement

    A 51-200 employee insurance firm in Austria faces a specific constraint: senior staff spend 40 to 60 percent of their week on routine lookups, document extraction, and first-response triage. The process audit that opens a fixed-scope pilot identifies which of these workflows have the highest volume and the clearest before/after metrics. For most mid-size insurers, the audit targets three areas: invoice processing and document extraction in the back office, customer-facing ticket triage on support channels, and internal knowledge search over policy manuals and CRM records. The pilot then focuses on one of these workflows, not all three, to prove value within a 6 to 10 week window. The baseline is measured before any AI touches the workflow: cycle time per ticket, error rate on document extraction, and the number of tickets that require a human agent. This baseline is the reference point for the after measurement, and it is what the pilot report will show to the board or the compliance officer.

    Customer-Facing Assistant on Support Channels

    The customer-facing assistant handles first-response triage on the firm’s support channels. It reads the incoming ticket, classifies it by policy type and urgency, and drafts a first response using the company’s own documentation and CRM records. The architecture uses LangChain for chaining LLM calls and retrieval, and LangGraph for stateful, cyclic workflows that let the assistant loop through retrieval, classification, and escalation steps. The assistant connects to the existing helpdesk and CRM through their native REST APIs and webhooks; it does not replace these systems. For an Austrian firm handling health data, the model layer is deliberately model-agnostic: OpenAI or Anthropic APIs handle tasks where quality matters, while open-weight models run on the client’s own hardware when regulated data cannot leave the building. The human-in-the-loop default means the model drafts or classifies, and a person approves anything that touches money, health data, or a contract. Every pilot ships with a measured before/after baseline on cycle time and error rate, so the cost per ticket reduction is quantified, not estimated.

    Internal Knowledge Search for Legal and Compliance

    The internal knowledge search assistant lets legal and compliance staff query the company’s own documentation, policy manuals, and CRM records in natural language. It returns cited answers from the source documents, reducing the time staff spend searching through PDFs and legacy systems. The retrieval layer uses a vector index over the firm’s document corpus, built with LangChain’s retrieval primitives. The assistant is model-agnostic: for documents that contain personal data or health records, the retrieval and generation steps run on open-weight models on the client’s own hardware. For general policy documentation, a commercial API may be used. The key design constraint is that the assistant does not make decisions; it retrieves and cites. A compliance officer reviews the cited answer before acting on it. This keeps the system within the lower-risk categories of the EU AI Act, which requires transparency for AI systems that assist human decision-making but does not mandate conformity assessment for purely retrieval-based tools.

    Predictive Scoring for Claim and Ticket Triage

    Predictive scoring assigns a probability to each incoming ticket or claim based on historical data. In the pilot, the scoring model is trained on the firm’s past 12 to 24 months of ticket and claim data, using features such as policy type, claim amount, and historical resolution time. The model flags high-risk or high-value cases for immediate human review. For example, a claim with a fraud likelihood score above 0.7 is routed to a senior adjuster before the first response is drafted. The scoring model runs as a separate service, called by the LangGraph workflow at the classification step. It does not replace the human decision; it prioritizes the queue. The before/after baseline for the pilot includes the number of high-risk cases that were missed in the manual process versus the number flagged by the scoring model. This metric is what the compliance officer will review when assessing whether the system meets the firm’s internal risk thresholds.

    EU AI Act Compliance and Data Residency

    The EU AI Act, which entered into force in August 2024 and applies in phases through 2026, classifies AI systems by risk level. A customer-facing assistant that handles health data or makes decisions affecting policyholders may fall under high-risk categories, requiring conformity assessment, logging, and human oversight. A purely internal knowledge search tool is generally lower risk but still subject to transparency obligations. For an Austrian insurance firm, the practical compliance steps are: document the intended use of each AI component, ensure that human-in-the-loop approval is in place for anything touching money, health data, or contracts, and maintain logs of model inputs and outputs for the period required by the Act. The fixed-scope pilot includes a compliance review as part of the handover documentation. The firm’s legal team reviews the pilot report before the system moves to managed operation. The architecture is designed so that the compliance controls are built into the workflow, not bolted on after deployment.

    Pilot Timeline and Delivery Model

    The fixed-scope pilot runs 6 to 10 weeks for a 51-200 employee insurance firm. The first two weeks cover the process audit and baseline measurement. The next four to six weeks build and test the pilot on one workflow, with weekly check-ins between the delivery team and the firm’s operations and compliance staff. The final week handles handover, documentation, and the before/after report. The pilot is delivered by a product studio with eight years of delivery experience, working with founders and operators across fintech, healthcare, e-commerce, B2B SaaS, logistics, insurance, and professional services in Tier-1 markets. The delivery model is fixed-scope: the features, the timeline, and the success metrics are defined before the pilot starts. If the pilot meets the baseline targets, the firm moves to rollout and managed operation. If it does not, the firm has a documented reason and a measured baseline to decide the next step. The cost of the pilot is fixed and agreed in advance, with no open-ended scope.

  • Fixed-Scope Pilot vs. In-House Build: Lead Qualification for a UK Fintech

    What Is Being Compared

    The two options are distinct in scope and risk profile. Option A is a fixed-scope pilot delivered by an external product studio: a 6-8 week engagement on one workflow—lead qualification—using the Anthropic Claude API as the model layer, integrated via custom REST API and webhooks into the existing CRM. The studio handles technical planning, product design, and full-cycle development. The pilot ships with a measured before/after baseline on cycle time and error rate. Option B is a fully in-house build: the company’s own engineering team designs, develops, and operates the agent, using the same model API or an open-weight model on internal hardware. The in-house team owns the architecture, the integration, and the ongoing operation. Both options target the same use case—lead qualification for a 201-500 employee fintech in the UK—but they differ in who bears the delivery risk, how fast the first working system ships, and what the company must maintain after the pilot.

    Criteria for the Comparison

    The comparison is judged against seven criteria that matter to a fintech scaling operations without new hires:

    • Time to first working system — how many weeks from kickoff to a live agent handling real leads.
    • Total cost of ownership over 6 months — including model API costs, integration work, and ongoing operation.
    • PCI DSS scope impact — whether the agent’s data boundary touches cardholder data and what that means for compliance.
    • Error rate reduction — the measured delta in misclassified leads between the manual baseline and the agent.
    • Cycle time reduction — the measured delta in time from lead creation to qualified status.
    • Vendor lock-in — how easily the company can switch model providers or take the system in-house after the pilot.
    • Operational burden — who monitors, tunes, and maintains the agent after the pilot ends.

    Comparison Table

    Criterion Option A: Fixed-Scope Pilot (External Studio) Option B: In-House Build
    Time to first working system 6-8 weeks from kickoff; studio has delivery templates and prior fintech experience 12-16 weeks minimum; team must design architecture, build integration, and tune the model from scratch
    Total cost over 6 months Fixed pilot fee (typically £25,000-£40,000) plus Anthropic API usage (approx. £1,500-£3,000/month at 500-1,000 leads/month); no new hires 2-3 FTEs at £60,000-£80,000/year each plus API costs; total £150,000-£250,000 over 6 months including salaries
    PCI DSS scope impact Studio designs data boundary to exclude cardholder data; client retains compliance ownership Same design principle, but in-house team must validate the boundary against PCI DSS 4.0 requirements; no external review
    Error rate reduction Measured in pilot; studio ships with baseline and delta report; typical delta: 30-50% reduction in misclassification Measured after build; no external baseline; team must design the measurement framework themselves
    Cycle time reduction Measured in pilot; typical delta: 40-60% reduction in time-to-qualified Measured after build; no external baseline; team must design the measurement framework themselves
    Vendor lock-in Low: model-agnostic architecture; client can switch to OpenAI or an open-weight model post-pilot Low: in-house team controls the stack; no external dependency
    Operational burden Studio provides handover documentation and a 30-day post-pilot support window; client takes over operation In-house team owns all operation, monitoring, and tuning from day one

    Scenario-by-Scenario Verdict

    When Option A wins: The company has no dedicated AI engineering team and needs a working lead qualification agent within 6-8 weeks to hit a quarterly sales target. The fixed-scope pilot removes delivery risk: the studio has delivered similar systems for fintech and payments clients in Tier-1 markets, and the pilot’s measured baseline gives the sales team a concrete number to report to leadership. The 6-month timeline is tight for an in-house build, and the pilot’s fixed fee is a smaller commitment than hiring 2-3 engineers. For a 201-500 employee company where every new hire is a significant cost, the pilot’s cost profile is easier to justify.

    When Option B wins: The company already has a strong engineering team with experience in API integrations and LLM applications, and the lead qualification workflow is one of several AI initiatives the team is building. The in-house build gives the team full control over the architecture, which matters if the company plans to extend the agent to other workflows (invoice processing, document extraction) over the next 12-18 months. The in-house team can also choose to run an open-weight model on internal hardware if the data residency requirements tighten, without renegotiating a vendor contract.

    Recommendation

    For a 201-500 employee UK fintech with a 6-month timeline and no dedicated AI engineering team, Option A—the fixed-scope pilot on the Anthropic Claude API—is the better fit. The pilot’s 6-8 week delivery window fits the 6-month timeline with room for a rollout phase after the pilot. The fixed fee is a smaller financial commitment than hiring 2-3 engineers, and the studio’s prior experience with fintech and payments clients in Tier-1 markets reduces the risk of a failed pilot. The measured baseline on cycle time and error rate gives the sales team a concrete business case for scaling. The model-agnostic architecture means the company is not locked into Anthropic; if the data residency requirements change, the team can switch to an open-weight model on internal hardware without rebuilding the integration. The in-house build is the right choice only if the company already has the engineering capacity and the lead qualification agent is part of a broader AI roadmap that justifies the longer build time and higher cost.

  • Swiss E-Commerce Team Cuts Invoice Cycle Time 47% with a Claude Extraction Pilot

    Background: A Swiss E-Commerce Operations Team at the Pilot Stage

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements. We do not name real clients. The company described here is a plausible representative of a profile we have worked with repeatedly: a mid-sized Swiss e-commerce and retail operations firm, roughly 120 employees, running a mixed stack of SAP Business One for ERP, Microsoft Teams for internal communication, and a legacy document management system for incoming supplier invoices. The team was in the “running isolated pilots” stage of AI maturity: they had experimented with a generic OCR tool on a small sample of invoices, seen promising results, but had no structured process to move from experiment to production. The finance and operations leads wanted a repeatable path, not another one-off test.

    Challenge: 1,800 Invoices a Month, No Headroom, and a Compliance Clock

    The operations team processed roughly 1,800 supplier invoices per month across 14 business days. Each invoice required a clerk to open the PDF, transcribe vendor name, line items, tax codes, and payment terms into SAP Business One, then flag discrepancies for review. The average cycle time from receipt to ERP entry was 3.2 days, with a field-level error rate of 11% on a 200-invoice sample. Two pressures made the status quo untenable: first, the EU AI Act’s transparency and human-oversight obligations (Articles 13 and 14) meant that any automated system handling financial data needed a documented approval workflow, and the team had no such process in place. Second, the operations lead was managing a 20% volume increase tied to a new retail distribution agreement that closed in six weeks. Hiring two additional clerks would have cost roughly CHF 14,000 per month in fully loaded salary, and the onboarding cycle for a new finance clerk in the Swiss market was 4 to 6 weeks.

    Approach: A Two-Week Pilot on One Workflow, Built on Claude and Teams

    Forfis scoped a two-week, fixed-scope pilot on a single workflow: supplier invoice extraction and ERP entry. The architecture used the Anthropic Claude API for extraction, chosen for its 200K-token context window, which handled multi-page invoices and attached purchase orders in a single inference call without chunking. The model output was constrained to a JSON schema matching SAP Business One’s field structure. The integration path was deliberately thin: incoming invoices arrived via email to a monitored mailbox, a lightweight ingestion service pulled the PDFs, the Claude API extracted and classified the fields, and the result was pushed to SAP via its REST API. Approval requests and status updates routed through Microsoft Teams, where the finance team reviewed extractions above a CHF 5,000 threshold. The human-in-the-loop rule was explicit: any invoice touching a payment, a contract clause, or a tax code required a named approver’s sign-off before the ERP write. The pilot team included one Forfis engineer, one product designer, and the client’s operations lead, working as a dedicated AI team embedded in the client’s daily standup.

    Outcome: 47% Faster Cycle Time, 5.8% Error Rate, Zero Re-Keys

    The pilot ran for 10 business days on a live subset of 320 invoices. The measured results, compared against the pre-pilot baseline: cycle time from receipt to ERP entry dropped from 3.2 days to 1.7 days, a 47% reduction. The field-level error rate fell from 11% to 5.8% on the same 200-invoice verification sample. The finance team approved 94% of extractions without correction; the remaining 6% were flagged by the model’s own confidence score and routed to a human reviewer before ERP entry. No invoice required a full re-key. The operations lead reported that the two clerks who had been doing manual entry were redeployed to handle the 20% volume increase from the new distribution agreement without additional hiring. The pilot’s measured baseline and post-pilot metrics were delivered as a one-page report, which the client used in a board presentation to justify a rollout to the remaining 12 invoice workflows. The EU AI Act compliance documentation, including the human-oversight log and transparency disclosures, was included as an appendix.

    Lessons for Teams Running Isolated Pilots

    • Scope the pilot to one workflow, not one document type. The client initially wanted to pilot invoices, credit notes, and purchase orders simultaneously. Forfis pushed back: a single workflow with a full integration chain (ingestion, extraction, approval, ERP write-back, Teams notification) produces operationally meaningful metrics. A multi-document pilot with a partial integration chain produces vanity numbers. The client agreed, and the focused scope is why the two-week timeline held.
    • The baseline is a contractual deliverable, not an afterthought. Without the pre-automation measurement of cycle time and error rate, the team cannot quantify the improvement or justify the rollout. Forfis builds the baseline measurement into the first week of the pilot, even if it means the automation work starts on day four instead of day one.
    • Human-in-the-loop thresholds should be configurable, not hardcoded. The CHF 5,000 approval threshold was a starting point. During the pilot, the team observed that the model’s confidence score was a better predictor of error than the invoice amount. The threshold was adjusted to a hybrid rule: amount above CHF 5,000 OR confidence below 0.92 triggers human review. This reduced unnecessary approvals by 18% without increasing the error rate.
    • Integration through existing APIs keeps the operational surface small. The client did not want a new front-end. The approval workflow lived in Microsoft Teams, the ERP write went through SAP’s REST API, and the ingestion service was a 200-line Python script. The total new infrastructure was one container and one API key. This kept the post-pilot operational overhead low and made the managed-operation retainer straightforward.
    • EU AI Act compliance is a design constraint, not a documentation afterthought. The human-oversight log, the transparency disclosure to affected parties, and the model-output audit trail were built into the workflow from day one. Retrofitting compliance documentation after the pilot is live is more expensive and less defensible than building it in.
  • 8-Week AI Invoice Processing Pilot for German Professional Services Firms

    The Problem: Manual Invoice Entry in a German Professional Services Firm

    You run a 501-2000 employee professional services firm in Germany. Your operations team spends 12-15 hours per week manually entering invoice data from PDFs into your ERP. The error rate is 3-5%, and cycle time from receipt to approval is 5-7 business days. You want to replace this manual work with an AI-native pipeline that extracts data, routes approvals through Slack or Microsoft Teams, and posts to your ERP automatically. The constraint is GDPR: supplier contact details on invoices are personal data under Article 4(1), and you cannot transmit them to a third-party API without a Data Processing Agreement under Article 28. The use case is invoice processing for accounts payable, not customer-facing. The timeline is 8 weeks, and you need a dedicated AI team to deliver a fixed-scope pilot that measures before/after cycle time and error rate.

    Prerequisites: What You Need Before Week 1

    • ERP API access: Your ERP (SAP, Dynamics 365, or similar) must expose a REST or SOAP API for creating vendor invoices. Confirm the API supports field-level mapping for vendor name, invoice number, date, line items, total, and tax. If the API is rate-limited, confirm the limit (e.g., 100 requests/minute) and plan for batching.
    • Invoice repository: A shared folder or document management system where incoming invoices are stored. The pilot will pull from this location. Confirm the format (PDF, image, or both) and the naming convention.
    • Slack or Microsoft Teams workspace: The approval workflow will live here. Confirm you have admin access to create custom apps or bots. If using Teams, confirm you have access to the Teams Developer Portal.
    • GDPR documentation: A Data Processing Agreement template, a records of processing activities entry, and a data flow diagram showing where invoice data resides. If using OpenAI API, confirm the DPA covers EU data residency and zero-data-retention.
    • Baseline metrics: Two weeks of manual processing data: cycle time per invoice, error rate, and cost per invoice. This is your before/after benchmark.
    • Dedicated AI team: A technical lead, data engineer, product manager, QA engineer, and a client-side point of contact. The team works in 2-week sprints.

    Steps: From Audit to Pilot in 8 Weeks

    1. Audit the invoice stream. Pull the last 3 months of AP invoices from your repository. Categorize them by vendor, format (PDF vs. image), and complexity (single-line vs. multi-line). Identify the top 20 vendors that account for 80% of invoice volume. This is your pilot scope. Do not include new vendors or unusual formats.

    2. Define the extraction schema. List the fields you need: vendor name, invoice number, invoice date, due date, line items (description, quantity, unit price, total), tax rate, and total amount. Map each field to the corresponding ERP field. Document the data types and validation rules (e.g., invoice number is alphanumeric, max 20 characters).

    3. Set up the data pipeline. Build a pipeline that pulls invoices from the repository, converts them to text using OCR (Tesseract or Azure Document Intelligence), and sends the text to the extraction model. If using OpenAI API, configure the endpoint with your API key and set the model to gpt-4o for high accuracy. If using an open-weight model, deploy Llama 3 70B on your on-premises GPU server. The pipeline should output a JSON object with the extracted fields and a confidence score per field.

    4. Build the approval workflow. Create a Slack or Teams bot that sends a message to the approver with the extracted data, a link to the original invoice, and approve/reject buttons. The approver clicks approve, and the bot posts the invoice to the ERP via the API. If the approver rejects, the bot flags the invoice for manual review. Log every action with a timestamp and user ID for GDPR audit trails.

    5. Run the pilot. Process 500-1000 invoices over 4 weeks. Track cycle time, extraction accuracy, exception rate, and approver adoption weekly. Compare against your baseline. If the exception rate exceeds 15%, pause and investigate the root cause (e.g., poor OCR quality, ambiguous field labels). If approver adoption is below 80%, investigate workflow friction (e.g., too many clicks, unclear UI).

    6. Validate and document. After 4 weeks, compile a report with before/after metrics, error analysis, and recommendations for rollout. Document the GDPR compliance steps taken: DPA signed, data flow diagram updated, records of processing activities entry created. Present the report to stakeholders and decide on rollout scope.

    Common Pitfalls and How to Detect Them

    • Scope creep: Adding new invoice types or vendors mid-pilot. Detect: track the number of invoice types processed weekly. If it exceeds the pilot scope, pause and re-scope.
    • Poor OCR quality: Low-resolution scans or inconsistent formats cause extraction failures. Detect: track the OCR confidence score. If it falls below 0.8 for more than 10% of invoices, investigate the source documents.
    • Lack of approver buy-in: Approvers bypass the system and process invoices manually. Detect: track the percentage of invoices approved via the bot. If it is below 80%, investigate workflow friction and retrain approvers.
    • Integration failures: ERP API rate limits or authentication issues cause posting failures. Detect: track the API error rate. If it exceeds 5%, investigate the API configuration and implement retry logic with exponential backoff.
    • Over-reliance on the model: No human-in-the-loop for edge cases, leading to incorrect postings. Detect: track the number of invoices posted without approval. If it is greater than zero, investigate the approval workflow and add a mandatory approval step for low-confidence extractions.

    Next Steps: From Pilot to Rollout

    The pilot is complete. You have measured a 30-50% reduction in cycle time and a 20% reduction in error rate compared to baseline. The next logical step is to expand the pilot to additional invoice streams (e.g., AR invoices, expense reports) or to other back-office workflows (e.g., contract extraction, data entry for client onboarding). Before expanding, review the GDPR documentation and confirm that the new data flows are covered by the existing DPA. If the new workflows involve special categories of data (e.g., health data), conduct a Data Protection Impact Assessment under GDPR Article 35. The dedicated AI team can continue to manage the rollout, or you can transition to a managed service model where the team monitors the pipeline, handles exceptions, and iterates on the extraction model based on new invoice formats.

  • RAG Shipment Status Assistant for US Fintech: 12-Item PCI DSS Checklist

    Scope and Baseline

    This checklist applies to a US-based fintech with 2,000+ employees deploying a retrieval-augmented knowledge assistant to cut first-response time on order and shipment status inquiries. The assistant integrates with Slack or Microsoft Teams, uses LangChain and LangGraph for orchestration, and runs on a model-agnostic stack. The pilot is fixed-scope, eight weeks, and measured against a baseline captured in week zero. PCI DSS compliance is a hard constraint: the assistant must never ingest, store, or transmit cardholder data. Every item below is a discrete action you can mark done or not done.

    Data, Compliance, and Scope

    1. Capture the week-zero baseline. Sample 50–100 real shipment status inquiries and record median cycle time and error rate. This baseline is your success metric; without it, you cannot prove the pilot delivered value.

    2. Define the PCI DSS data boundary. Identify which fields in your CRM and ERP are in PCI scope (PAN, CVV, track data) and which are not (order ID, tracking number, status). The RAG vector store must be partitioned so the assistant never retrieves PCI-scope fields.

    3. Select the pilot workflow. Choose one high-volume channel (e.g., a Slack channel for shipment status) and one department. A fixed-scope pilot on a single workflow is deliverable in eight weeks; multi-department rollout is a separate engagement.

    4. Document the approval threshold. Specify which response types trigger human-in-the-loop review (any response touching money, health data, or a contract). This threshold is encoded as a node in the LangGraph pipeline and must be agreed with your compliance team before week one.

    Architecture and Pipeline

    1. Build the extraction pipeline. Ingest shipment status data from your ERP or carrier API using layout-aware OCR and LLM-based field extraction. Validate extracted fields against known formats (e.g., USPS tracking numbers are 20–22 digits) and flag low-confidence extractions for human review.

    2. Partition the vector store. Create a non-PCI partition for shipment status, order metadata, and policy docs. The RAG retrieval query accesses only this partition by default; PCI-scope data is never embedded.

    3. Configure the LangGraph pipeline. Define the stateful graph: parse inbound message → classify intent → query vector store → check PCI scope → route to human if needed → format and send. LangGraph handles branching logic and human-in-the-loop interrupts; LangChain handles LLM calls and vector store interactions.

    4. Select the model stack. Use OpenAI or Anthropic APIs for quality-critical steps (intent classification, response generation) and open-weight models on client hardware if regulated data cannot leave the building. The architecture is model-agnostic; the choice depends on your data residency and compliance constraints.

    Integration, Approval, and Measurement

    1. Integrate with Slack or Microsoft Teams. Use the Events API (Slack) or Bot Framework (Teams) to listen for messages in a designated channel and post responses. The integration layer is a thin adapter that translates between the messaging platform’s format and the LangGraph pipeline’s schema; the core RAG logic is platform-agnostic.

    2. Implement the human-in-the-loop gate. Add a node that pauses the pipeline when the response touches money, health data, or a contract. The gate sends the draft response to a human approver via Slack or Teams and waits for sign-off before delivering to the customer.

    3. Set up monitoring and logging. Log every pipeline execution: input, extracted fields, retrieved documents, generated response, and approval status. This log is your audit trail for PCI DSS and your debugging tool when the assistant misbehaves.

    4. Run the eight-week measurement. Re-measure the same 50–100 inquiries through the automated pipeline and compare cycle time and error rate against the week-zero baseline. The delta is your before/after metric; if the pilot hits its targets, scope the rollout separately with a new SOW.

  • AI Process Audit vs. Direct Contract Review Pilot: A 3-Month Fintech Verdict

    What Is Being Compared

    The two options are not alternatives but sequential phases of the same engagement. Option A is the AI process audit and roadmap: a structured assessment of every finance and accounting workflow in a 2,000+ employee Austrian fintech, scored on volume, error rate, cycle time, and integration complexity, producing a prioritized automation roadmap. Option B is the direct contract review pilot: a fixed-scope, 3-month build that deploys an AI layer for contract clause extraction, data enrichment, and cleanup, integrated into SAP or Microsoft Dynamics ERP, with a measured before/after baseline on cycle time and error rate. The question is whether a company in the “Running Isolated Pilots” maturity stage should spend the first 3 months on the audit or jump straight to the pilot. The answer depends on how many workflows are candidates, how well the ERP integration surface is documented, and whether the finance team can commit senior staff to the audit interviews.

    Criteria for Judgment

    Eight criteria determine which path delivers more value in a 3-month window:

    • Scope clarity: Does the company know which workflows to automate, or is that the unknown?
    • ERP integration readiness: Are SAP BAPI/RFC or Dynamics OData endpoints documented and accessible?
    • Data availability: Can the finance team provide 200+ historical contract samples for model validation?
    • Compliance surface: Does the contract data touch PSD2 payment records or MiFID II client data, requiring on-premise deployment?
    • Senior staff availability: Can 2–3 senior finance or legal reviewers commit 4 hours/week to the human-in-the-loop approval layer?
    • Model accuracy gap: Is the open-weight model’s extraction accuracy within 5% of the frontier API for the specific contract types?
    • Rollout dependency: Does the pilot’s success depend on a roadmap that sequences multiple workflows, or is contract review a standalone win?
    • Budget structure: Is the 3-month budget a fixed pilot fee or an audit-plus-pilot package?

    Comparison Table

    Criterion Option A: AI Process Audit & Roadmap Option B: Direct Contract Review Pilot
    Time to first measurable result 4–6 weeks (audit report) 6–8 weeks (pilot baseline)
    Scope All finance/accounting workflows One workflow: contract review
    ERP integration depth Read-only access for data profiling Write access via SAP BAPI or Dynamics OData
    Data requirement 50–100 sample records per workflow 200+ historical contracts for validation
    Model selection Recommended, not deployed Open-weight model deployed on-premise
    Output Prioritized roadmap with ROI per workflow Measured cycle time and error rate delta
    Senior staff commitment 2–3 reviewers, 4 hrs/week for interviews 2–3 reviewers, 4 hrs/week for approval layer
    Risk of scope creep Low (fixed audit scope) Medium (new contract types discovered mid-pilot)

    Scenario-by-Scenario Verdict

    When Option A wins: The company has not previously run any AI pilot and does not know which of its 15–20 finance workflows are worth automating. The audit prevents the common failure mode of picking a low-volume, high-complexity workflow that looks impressive in a demo but delivers no ROI. For a 2,000+ employee fintech with multiple business units (payments, lending, insurance products), the audit surfaces that contract review is only one of four high-value targets, and sequencing matters. The 3-month audit produces a roadmap that justifies a 9-month rollout budget.

    When Option B wins: The company already knows contract review is the target—perhaps because a prior isolated pilot on invoice processing proved the model-agnostic architecture works. The finance team has 200+ historical contracts, the SAP AP module API is documented, and the project sponsor wants a measurable before/after baseline within 60 days. In this case, the audit adds 4 weeks of delay without changing the pilot scope.

    Hybrid scenario: A 2-week compressed audit (covering only contract review and two adjacent workflows) followed by a 10-week pilot. This fits the 3-month timeline and gives the roadmap context without the full audit cost.

    Recommendation

    For a 2,000+ employee Austrian fintech in the “Running Isolated Pilots” maturity stage, with a 3-month timeline and a specific need to free senior staff from routine contract review, Option B—the direct contract review pilot—is the correct first move, provided two conditions are met: the finance team can supply 200+ historical contract samples within the first two weeks, and the SAP or Dynamics ERP integration surface is documented. The pilot delivers a measurable baseline (cycle time, error rate, throughput) that becomes the business case for the full rollout. The audit is not skipped; it is compressed into the first 10 days of the pilot, covering contract review and two adjacent workflows (invoice data entry, vendor master data cleanup). This hybrid approach respects the 3-month constraint, uses the open-weight model on-premise to keep PSD2 and MiFID II data inside the building, and plugs into the existing ERP via API rather than replacing it. The managed operations retainer begins at pilot completion, ensuring the model stays current as contract templates evolve.

  • 4-Week Invoice Processing Pilot for a 201-500 Employee Firm in Germany

    The Back-Office Bottleneck: Where Senior Hours Go to Die

    A 201-500 employee professional services firm in Germany processes 1,200 to 3,000 vendor invoices per month. Each invoice is received by email, printed or forwarded to a back-office clerk, manually entered into the ERP, and approved by a senior accountant. The average cycle time from receipt to payment entry is 3 to 5 business days. The error rate on data entry sits at 4 to 7%, meaning roughly 50 to 200 invoices per month require rework. Senior staff spend 12 to 18 hours per week on invoice review and correction, time that could go to client work or strategic planning. The pain is not the invoice itself; it is the friction between the document and the system of record, and the human cost of bridging that gap.

    Why Off-the-Shelf OCR and RPA Fall Short

    The first common approach is to buy an OCR tool and hope it works. Most OCR engines handle clean, structured invoices well but fail on the messy 20% that includes handwritten notes, multi-page documents, and vendor-specific layouts. The second approach is to hire more back-office staff. This adds cost without reducing cycle time, and it does not address the root cause: the manual handoff between document and ERP. The third approach is to build a custom RPA bot. RPA works for repetitive, rule-based tasks but breaks when the invoice format changes, and it requires constant maintenance. None of these approaches include a predictive layer that flags high-risk invoices for human review, so the senior accountant still reviews every single entry. The result is a system that is faster than manual entry but still slow, still error-prone, and still dependent on human attention for every transaction.

    The 4-Week Pilot: Extraction, Scoring, and Approval

    The pilot runs for 4 weeks and covers one invoice type, one ERP integration, and one approval channel. Week 1 is the process audit: map the current workflow, measure the baseline cycle time and error rate on a sample of 200 invoices, and identify the fields that the model must extract. Week 2 builds the extraction pipeline using the OpenAI API to parse the invoice and pull out vendor name, amount, tax, due date, and line items. The predictive scoring model is trained on the historical data from that invoice type to assign a risk score to each entry. Week 3 runs the model in shadow mode: it processes invoices in parallel with the human team, and the output is compared against the manual entries. Week 4 flips the switch to human-in-the-loop mode. The AI drafts the entry, the predictive model assigns a risk score, and if the score is below a threshold, the entry is auto-approved and pushed to the ERP. If the score is above the threshold, the entry is sent to a senior accountant via Slack or Microsoft Teams for one-click approval. Every decision is logged with a timestamp, the approver’s name, and the model’s confidence score.

    EU AI Act Compliance: What the Pilot Must Log

    The EU AI Act classifies invoice processing as a limited-risk use case under Article 6. The firm must maintain a record of the model’s intended purpose, document the human-in-the-loop approval step, and ensure the system does not make autonomous financial decisions. For a 201-500 employee firm in Germany, this means logging every AI-drafted invoice entry and the human who approved it, storing those logs for at least six years under the German commercial code, and providing a clear opt-out if a client disputes an automated classification. The predictive scoring model must be explainable: the firm must be able to state why a particular invoice was flagged for manual review. The OpenAI API’s output includes a confidence score for each extracted field, which serves as the basis for the risk score. The Slack or Teams integration provides a natural audit trail: every approval or rejection is timestamped and attributed to a named user. This satisfies the Act’s transparency requirement and gives the firm a defensible position in the event of a regulatory inquiry.

    How to Start: Five Concrete First Steps

    Step 1: Run the process audit. Identify the invoice type with the highest volume and error rate. Measure the baseline cycle time and error rate on a sample of 200 to 500 invoices. Step 2: Define the pilot scope. One invoice type, one ERP integration, one approval channel. Confirm that the ERP API is documented and accessible. Step 3: Build the extraction pipeline. Connect the OpenAI API to the invoice document store. Define the fields to extract and the validation rules. Step 4: Train the predictive scoring model. Use the historical data from the pilot invoice type to train a model that flags high-risk entries. Step 5: Configure the Slack or Teams integration. Set up the approval workflow so that senior accountants receive a notification with the extracted fields and a one-click approve/reject action. Step 6: Run the pilot in shadow mode for one week, then flip to human-in-the-loop mode for the remaining three weeks. Measure the cycle time and error rate at the end of week 4 and compare against the baseline.

  • 3-Month Roadmap: AI Invoice Processing for Austrian Insurers

    The Problem: Manual Back-Office Work Drives Up Support Ticket Costs

    Austrian insurers with 51-200 employees face a specific problem: back-office staff spend 40-60% of their time on manual invoice processing, data entry, and routine customer queries. This drives up the cost per support ticket and delays first-response times, which erodes customer satisfaction. The solution is to integrate AI automation into the systems you already run, starting with a process audit that identifies the workflows worth automating. This article walks you through a 3-month roadmap to implement AI-assisted invoice processing, customer-facing assistants, and Slack/Teams integration, all while staying GDPR-compliant and reducing your cost per support ticket.

    Prerequisites: What You Need Before Step 1

    Before you start, you need:

    • API access to your ERP (e.g., SAP, Microsoft Dynamics) and CRM (e.g., Salesforce, HubSpot) for data extraction and posting.
    • Slack or Microsoft Teams workspace with admin rights to create custom integrations.
    • A designated project owner with authority to approve scope changes and budget.
    • GDPR compliance documentation: Record of Processing Activities (Article 30), Data Protection Impact Assessment (DPIA), and privacy notice updates.
    • A measured baseline on cycle time and error rate for your current invoice processing workflow.
    • Access to OpenAI API or an equivalent model provider for the pilot phase.

    Without these, you will hit blockers in weeks 2-4 that delay the entire timeline.

    Steps 1-3: Audit, Pilot Scope, and AI Extraction Layer

    Step 1: Run a 2-week process audit.
    Identify the highest-volume, highest-error workflows in your back-office. Use a simple spreadsheet to track: workflow name, volume per week, average cycle time, error rate, and staff hours spent. Focus on invoice processing, document extraction, and data entry. This audit tells you which workflows are worth automating and gives you a baseline for measuring ROI.

    Step 2: Define a fixed-scope pilot.
    Pick one workflow (e.g., invoice extraction) and define the scope: input document types, output fields, integration points, and success metrics. Write a one-page pilot charter that includes: scope, timeline (4 weeks), success criteria (e.g., 95% extraction accuracy, 50% reduction in cycle time), and out-of-scope items. This prevents scope creep and keeps the pilot focused.

    Step 3: Build the AI extraction layer.
    Use OpenAI’s GPT-4o or GPT-4 Turbo API to extract data from invoices. Write a Python script that sends the invoice PDF to the API, parses the JSON response, and maps the fields to your ERP schema. Test with 50-100 real invoices from your baseline period. Track accuracy and error rate. If accuracy is below 95%, refine the prompt or add a human-in-the-loop review step.

    Steps 4-6: Slack/Teams Integration, Customer Assistant, and Measurement

    Step 4: Integrate with Slack or Microsoft Teams.
    Create a custom bot in Slack or Teams that receives extracted invoice data and posts it to a channel for human review. Use the Slack API or Teams Bot Framework to send messages with the extracted fields and a link to the original invoice. Add a button for “Approve” and “Reject” so staff can review and approve with one click. This reduces the time from extraction to approval from hours to minutes.

    Step 5: Add a customer-facing assistant.
    Build a retrieval-augmented assistant over your company’s documentation and CRM records. Use OpenAI’s API to generate first-response drafts for common customer queries (e.g., “Where is my claim?”, “How do I file an invoice?”). The assistant drafts the response, and a human approves it before it goes to the customer. This cuts first-response time from hours to minutes and reduces the cost per support ticket.

    Step 6: Measure and refine.
    Track cycle time, error rate, and cost per support ticket weekly. Compare against your baseline. If error rate is above 5%, refine the extraction prompt or add more human review. If first-response time is above 15 minutes, adjust the assistant’s prompt or add more documentation to the retrieval index. Iterate until you hit your success criteria.

    Step 7: Rollout, Managed Operations, and Common Pitfalls

    Step 7: Roll out and transition to managed operations.
    Once the pilot hits its success criteria, roll out to additional workflows (e.g., claims documentation, policy administration). Transition to managed operations: the vendor handles model monitoring, retraining, and integration maintenance. You get an SLA for uptime, accuracy, and response time. The vendor monitors for drift (e.g., if invoice formats change) and retrains the model as needed. This reduces the need for in-house ML expertise and ensures the system stays accurate as your document types evolve.

    Common pitfalls:

    • No baseline: You cannot prove ROI if you do not measure cycle time and error rate before the pilot. Detect this by checking your audit spreadsheet for baseline data.
    • Scope creep: Trying to automate too many workflows at once leads to delays. Detect this by reviewing the pilot charter weekly and rejecting out-of-scope requests.
    • GDPR non-compliance: Ignoring GDPR requirements results in data breaches or regulatory fines. Detect this by reviewing your DPIA and privacy notice before the pilot starts.
    • Low staff adoption: Not training staff on the new system leads to low adoption. Detect this by tracking staff feedback and usage metrics weekly.
  • B2B SaaS Firm in UAE Cuts Contract Review Cycle Time 50% with RAG Assistant

    Background: A 300-Person B2B SaaS Firm in Dubai

    This case study is a composite built from patterns Forfis has observed across multiple engagements. We do not name real customers. The company described here is a 300-person B2B SaaS firm based in Dubai, selling a project-management platform to mid-market clients across the Gulf. Its finance and accounting team of 18 handles contract review, invoice processing, and month-end close. The firm runs on a standard stack: Salesforce for CRM, NetSuite for ERP, Confluence for internal documentation, and Zendesk for customer support. It holds ISO 27001 certification and operates under UAE data residency expectations for client contract data. The team had been using a manual review process where a senior accountant reads every clause in a new contract against a playbook stored in Confluence, flags deviations, and routes the contract to legal for approval. The average cycle time for a standard contract was 4.2 days, and the error rate on clause flags was around 12%.

    Challenge: Contract Review Backlog and ISO 27001 Constraints

    The finance director set a clear goal: reduce the cost per contract review ticket and free the senior team from routine clause checks. The operational pressure was threefold. First, the firm was closing 40-60 new contracts per month, and the review backlog was growing. Second, ISO 27001 required documented controls over how contract data was handled, which limited the options for sending data to external APIs without a clear data processing agreement. Third, the team had a 3-month window before the next quarter’s planning cycle, and the director needed a measurable baseline to justify a larger automation budget. The specific need was not to replace the senior reviewers but to shift them from reading every clause to reviewing only the exceptions the system flagged. The director also wanted the solution to plug into the existing Confluence playbook and Salesforce approval chain, not to replace either tool.

    Approach: RAG Assistant Over Confluence with OpenAI API

    Forfis ran a 2-week process audit that mapped the contract review workflow end to end. The audit confirmed that 70% of the clauses in standard contracts were repetitive checks against the playbook, and that the Confluence space held 200+ pages of precedent and redline history. The pilot scope was fixed: build a retrieval-augmented assistant that ingests the Confluence playbook, retrieves the most relevant precedent for each clause in a new contract, and drafts a flag or approval recommendation. The model layer used the OpenAI API for inference, with a vector store running on the firm’s own AWS account in the UAE region to satisfy data residency. The integration layer connected to Confluence via its REST API and to Salesforce via the standard approval workflow API. The human-in-the-loop design meant the assistant drafted the flag, and a senior reviewer approved or edited it before it went to legal. Every pilot shipped with a measured before/after baseline on cycle time and error rate.

    Outcome: 50% Faster Cycle Time and 4% Error Rate

    The 3-month pilot ran from week 3 to week 13. In month 1, the team ingested the Confluence playbook into the vector store and tuned the retrieval parameters. In month 2, the assistant went into internal testing with 30 real contracts, and the senior reviewers calibrated the flag thresholds. In month 3, the assistant handled live contracts in parallel with the manual process, and the team tracked cycle time and error rate against the pre-pilot baseline. The results: average cycle time for a standard contract dropped from 4.2 days to 2.1 days, a 50% reduction. The error rate on clause flags fell from 12% to 4%, because the assistant caught deviations the manual process had missed. The senior team spent 60% less time on routine clause checks and redirected that time to complex negotiations and month-end close. The cost per contract review ticket dropped by roughly 45% when measured in senior hours. The ISO 27001 audit trail was maintained through the approval log, which recorded every flag, approval, and edit.

    Lessons for Similar Teams

    • Start with the playbook, not the model. The quality of a RAG assistant depends on the quality of the source documents. If the Confluence playbook is stale or inconsistent, the assistant will retrieve the wrong precedent. Spend the first two weeks cleaning and structuring the playbook before building the pipeline.
    • Fix the scope before you build. A 3-month pilot works only if the scope is fixed to one workflow. Trying to automate contract review, invoice processing, and data entry in the same window will stretch the team thin and dilute the baseline measurement.
    • Data residency is a design constraint, not an afterthought. For a firm in the UAE with ISO 27001 certification, the vector store and inference layer must run in a region that satisfies the data residency policy. Planning this in week 1 avoids a rework in week 8.
    • The human-in-the-loop approval log is your audit trail. Every flag, approval, and edit should be logged with a timestamp and reviewer ID. This satisfies ISO 27001 control A.12.4 (logging and monitoring) and gives the team a feedback loop to improve retrieval quality over time.
    • Measure cycle time and error rate from day one. The before/after baseline is the only way to justify the pilot to the board. Without it, the outcome is anecdotal, and the next budget cycle will be harder to win.
  • AI Contract Review for Logistics: Cut Back-Office Errors by 50% in 6 Months

    1. Baseline Measurement Before You Touch a Single Clause

    Logistics firms with 201-500 employees process 500-2,000 carrier agreements, customs declarations, and service contracts monthly. Manual review by legal and compliance staff takes 15-30 minutes per document, with an 8-12% error rate on clause identification. A RAG-based contract assistant reduces this to 3-5 minutes per document with under 2% error rate. The system indexes templates and precedents from Confluence, extracts key clauses, flags deviations from standard terms, and routes exceptions to human reviewers. For a team of 12 legal staff, this saves 15-20 hours weekly, shifting focus from data entry to strategic risk assessment. The 6-month timeline includes a 4-week audit, 6-week pilot on one contract type, and 14-week rollout with measurable checkpoints at each phase.

    2. On-Premise Open-Weight Models Keep Regulated Data In-Building

    Logistics contracts often contain customs declarations, hazardous material certifications, and client NDAs with strict data residency clauses. Sending these to external APIs like OpenAI or Anthropic may violate contractual or regulatory obligations. Open-weight models like Llama 3 or Mistral deployed on the client’s own hardware ensure data sovereignty, reduce latency to under 50ms for local inference, and eliminate per-token API costs at scale. The trade-off is higher initial infrastructure investment and the need for dedicated MLOps support for model updates. For a 201-500 employee firm, on-premise deployment typically requires 2-4 GPU servers and a dedicated MLOps engineer for the 6-month engagement. The model-agnostic architecture allows switching between cloud and on-premise models based on data sensitivity, with the same RAG pipeline and integration layer.

    3. RAG Over Confluence Turns Your Knowledge Base Into a Review Engine

    The RAG pipeline indexes contract templates, past executed agreements, and compliance checklists from Confluence or Notion into a vector database. When a new contract arrives, the system extracts key clauses (liability caps, SLA terms, termination conditions) and retrieves relevant precedents from the knowledge base. The LLM drafts a review summary highlighting deviations from standard terms, flagging clauses that exceed risk thresholds. Human reviewers approve or reject each flag before the contract proceeds to signature. The system logs every decision, creating an audit trail for compliance. Integration with the existing ERP ensures that approved contracts automatically update vendor master data and payment terms. The conversational agent handles initial intake, extracting metadata and routing contracts to appropriate reviewers based on risk classification, reducing ticket volume to legal by 40-60%.

    4. Human-in-the-Loop Approval Is Non-Negotiable for Money and Liability

    The most common failure is treating AI as a replacement for human judgment rather than an augmentation tool. Firms that remove human approval for contracts touching money, liability, or regulatory compliance face significant risk. The second pitfall is insufficient baseline measurement: without pre-implementation data on cycle time and error rate, you cannot prove ROI or identify where the AI is actually helping. The third is poor integration: if the AI assistant doesn’t plug into the existing CRM, ERP, and helpdesk via APIs, it creates a parallel workflow that increases rather than reduces manual work. The fourth is model selection mismatch: using cloud APIs for data that must stay on-premise, or using open-weight models when cloud quality is acceptable and cost-effective. Each pitfall has a measurable cost: unapproved AI decisions can trigger contract disputes, missing baselines make ROI unprovable, poor integration adds 20-30% overhead, and model mismatch increases costs by 40-60%.

    5. Dedicated AI Team Embeds in Your Org for the Full 6 Months

    A dedicated AI team typically includes a technical lead for architecture and model selection, a product designer for workflow mapping and human-in-the-loop UX, two full-stack developers for API integrations with ERP/CRM systems, and an MLOps engineer for on-premise model deployment and monitoring. For a 201-500 employee firm, this team operates as an embedded unit within the client’s organization for the 6-month engagement, with weekly steering meetings and bi-weekly demo cycles. The team size scales with complexity: a single contract type pilot requires 4-5 people, while multi-type rollout may expand to 6-8. Post-engagement, a subset (1-2 people) transitions to managed operation support. The dedicated team model ensures continuity: the same people who built the system understand its failure modes and can respond to edge cases within 4-8 hours, compared to 24-48 hours for external support contracts.

    6. Six-Month Timeline With Measurable Checkpoints at Each Phase

    The 6-month timeline breaks down as: Weeks 1-4 for process audit and baseline measurement of current cycle times and error rates. Weeks 5-10 for pilot development on one contract type (e.g., carrier agreements), including RAG pipeline setup and integration with Confluence/Notion. Weeks 11-16 for pilot validation, error rate measurement, and human-in-the-loop workflow refinement. Weeks 17-24 for rollout to additional contract types, team training, and managed operation handoff. Each phase includes measurable checkpoints: the pilot must demonstrate at least 30% cycle time reduction and 50% error rate improvement before rollout proceeds. The final deliverable is a fully operational AI contract review system integrated with existing ERP, CRM, and helpdesk, with a documented runbook for the internal team to manage day-to-day operations. The system is model-agnostic, allowing future migration to newer models without re-architecting the pipeline.