Tag: Austria

  • Austrian Medtech Firm Cuts Invoice Reporting Cycle Time 61% in a 4-Week Pilot

    Background: A 1,200-Person Medtech Firm in Graz

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements in the healthcare and medtech sector. It does not describe a single named client. The company, the metrics, and the timeline are representative of what we see in the field when a mid-sized European healthcare organization moves from isolated AI pilots to a compliance-safe, managed rollout of invoice-processing automation.

    The company is a 1,200-person medtech firm based in Graz, Austria, manufacturing surgical instruments and diagnostic kits. It operates in 14 EU markets and reports under Austrian GAAP with quarterly IFRS reconciliation. The finance and accounting team is 42 people, of whom 11 handle accounts payable and monthly reporting. Their ERP is SAP S/4HANA, their document management system is a legacy on-prem archive, and their internal communication runs on Microsoft Teams. They had run two prior AI pilots — one for email triage, one for contract clause extraction — but neither had moved past the pilot stage. The finance director’s mandate was clear: automate the monthly invoice-to-reporting cycle without introducing a new compliance surface, and do it within a 4-week pilot window before the Q3 close.

    The Challenge: 3,400 Invoices, 11 Staff, a 4-Week Window

    The monthly reporting cycle ran from the 1st to the 10th of each month. During that window, 11 finance staff manually processed roughly 3,400 vendor invoices, extracted line items, matched them to purchase orders, flagged discrepancies, and posted entries to SAP. The cycle time from invoice receipt to ERP posting averaged 6.2 days, and the error rate — measured as manual corrections per 100 invoices — sat at 14.3. The finance director had a hard deadline: the Q3 close was in 11 weeks, and the board had asked for a visible efficiency gain by year-end. Headcount was not the constraint; the constraint was that the 11-person team could not absorb the 14% error rate without a second review pass, which doubled the cycle time. The prior two AI pilots had stalled because they were scoped as “AI projects” rather than as workflow replacements with a measured baseline. The finance director wanted a fixed-scope pilot with a before/after metric, not a proof of concept.

    Approach: Audit, Then a Fixed-Scope Pilot on LangChain and LangGraph

    Forfis ran a two-week AI Automation Audit before the pilot began. The audit mapped the invoice-to-reporting workflow end-to-end, sampled 200 invoices from the prior month, and established the baseline: 6.2-day cycle time, 14.3% error rate, 11 FTEs. The audit identified three automation points: (1) invoice ingestion and OCR extraction from the legacy archive, (2) line-item classification and PO matching, and (3) discrepancy flagging with a human approval step before ERP posting.

    The pilot architecture used LangChain for the orchestration layer and LangGraph for the stateful workflow graph that tracked each invoice through ingestion, extraction, classification, review, and posting. The model layer was split: OpenAI’s GPT-4o handled the ambiguous line-item classification (where quality mattered), and an open-weight Llama 3 70B model ran on the client’s own GPU server for the regulated data extraction step, so no invoice data left the Graz data center. The integration points were SAP S/4HANA (via its OData API for posting approved entries), Microsoft Teams (for reviewer notifications and the approval workflow), and the legacy document archive (via a file-watcher ingestion pipeline). Every entry that touched money required a human approval in Teams before it hit SAP. The pilot ran for four weeks: Week 1 integration, Week 2 model tuning on the client’s actual invoice samples, Week 3 live operation with human-in-the-loop review, Week 4 measurement and the go/no-go report.

    Outcome: 61% Cycle-Time Reduction, 3.8% Error Rate

    By the end of Week 4, the pilot had processed 3,100 live invoices. The measured results against the audit baseline:

    • Cycle time dropped from 6.2 days to 2.4 days (a 61% reduction). The bottleneck shifted from manual extraction to the human approval step, which the finance team chose to keep as a compliance control.
    • Error rate fell from 14.3% to 3.8% (manual corrections per 100 invoices). The remaining errors were concentrated in two vendor categories with non-standard invoice formats, which the team flagged for a follow-up prompt-tuning pass.
    • FTE allocation: the 11-person team redirected 4 FTEs from manual extraction to exception handling and vendor relationship management. The finance director did not reduce headcount; the freed capacity was absorbed into the Q3 close workload.
    • Compliance surface: no new data left the building. The open-weight model ran on the client’s GPU server; the OpenAI API calls were limited to the classification step, which operated on anonymized line-item text, not on invoice metadata or vendor names.

    The go/no-go report recommended a phased rollout to the remaining 12 EU markets over two quarters, with the same human-in-the-loop approval step retained for all money-touching entries.

    Lessons for Teams Running Isolated Pilots

    • Baseline before pilot, not after. The audit’s 200-invoice sample established the 6.2-day / 14.3% baseline before any model was tuned. Without that, the pilot’s results would have been unmeasurable. Teams that skip the baseline step cannot distinguish model improvement from natural variance.
    • Split the model layer by data sensitivity, not by convenience. The open-weight model on the client’s hardware handled the regulated extraction step; the cloud API handled the classification step on anonymized text. This split is what made the rollout compliance-safe without requiring a full on-prem LLM deployment.
    • Human-in-the-loop is a design constraint, not a fallback. The Teams approval step was in the LangGraph state machine from day one, not added after the pilot showed errors. Removing it post-hoc would have broken the workflow graph and required a re-architecture.
    • Fixed-scope pilot, not open-ended POC. The 4-week window with a defined go/no-go report forced the team to ship a measurable result rather than iterate indefinitely. The finance director’s mandate — “show me a number by week 4” — was the single most important constraint in the engagement.
    • Integration through existing APIs, not replacement. The system plugged into SAP, Teams, and the legacy archive through their native APIs. No system was replaced, which kept the rollout risk low and the change-management burden minimal.
  • AI Ticket Triage in Austrian Insurance: A 14-Term Glossary for Pilot Teams

    Scope and Conventions

    This glossary defines the operational and regulatory vocabulary that appears when an insurance or insurtech company with 2,000+ employees in Austria runs an isolated pilot for AI-assisted ticket triage and routing. The terms are alphabetized and each entry gives a definition followed by a contextual example tied to the scenario: a dedicated AI team integrating the OpenAI API into an existing helpdesk via custom REST API and webhooks, with a 4-week fixed-scope pilot and human-in-the-loop approval as the default. Where a term carries competing definitions in the industry, both are named and the one used here is indicated. The glossary assumes the reader is an operator or technical lead who has already completed a process audit and is scoping the pilot.

    A–C: AI Maturity, Automation Type, Baseline Metrics

    AI Maturity: Running Isolated Pilots — A stage in an organization’s AI adoption curve where the company has completed a process audit, selected one or two workflows for automation, and is executing a fixed-scope pilot with measurable baselines before committing to broader rollout. The pilot is “isolated” because it runs in parallel with existing processes, does not replace them, and ships with a before/after comparison on cycle time and error rate. In this scenario, the isolated pilot covers ticket triage and routing for a 2,000+ employee insurer in Austria, using the OpenAI API through a dedicated AI team over a 4-week timeline. The pilot’s output is a measured error-rate reduction in the back office, not a full system replacement.

    D–F: Customer-Facing AI, Dedicated AI Team, EU AI Act

    Customer-Facing AI Assistant — A software agent that interacts directly with end customers through a support channel (chat, email, voice) to answer questions, draft first responses, or route tickets. In this glossary the term refers specifically to the triage-and-routing layer, not a fully autonomous agent. Dedicated AI Team — A fixed-scope delivery unit (typically 3–5 specialists) assigned to a single client for the duration of the pilot and rollout, as opposed to a fractional or on-call resource. The team owns technical planning, prompt engineering, integration, and managed operation. EU AI Act — Regulation (EU) 2024/1689, which classifies AI systems by risk level. A ticket-triage system that only sorts and routes is generally not high-risk, but if it drafts policy terms or calculates premiums it may cross into high-risk territory. Forfis applies human-in-the-loop approval for any output touching money, health data, or contracts, satisfying the Act’s transparency and accountability requirements under Articles 13 and 14.

    H–O: Human-in-the-Loop, OpenAI API, Process Audit

    Human-in-the-Loop (HITL) — An architectural pattern where the AI model drafts, classifies, or routes, and a human agent reviews and approves before the output reaches the customer or triggers a financial transaction. HITL is the default configuration in Forfis engagements; it is not an optional add-on. OpenAI API — The hosted inference endpoint (e.g., GPT-4o, GPT-4o-mini) accessed via HTTPS with a client-provided API key. In this scenario it handles general triage classification and first-response drafting. Under a zero-data-retention agreement, OpenAI does not store or train on the client’s prompts. Process Audit — The initial engagement phase where Forfis maps existing workflows, measures baseline cycle time and error rate, and identifies which processes are worth automating. The audit output is a prioritized list; the pilot then targets the highest-ROI item, here ticket triage and routing.

    R–W: Round-the-Clock Response, Ticket Triage, Workflow Orchestration

    Round-the-Clock Customer Response — The operational requirement that customer support channels (email, chat, phone) are staffed or automated 24/7, 365 days a year. For an insurer in Austria, this means handling policy inquiries, claim status checks, and document requests outside business hours without a human agent. The AI triage layer addresses this by classifying and drafting responses for routine tickets at 03:00 CET, while flagging complex or regulated tickets for the next business-day human review. Ticket Triage and Routing — The process of classifying an incoming support ticket by category (claim, policy change, billing, technical) and assigning it to the correct team or queue. In this scenario, the AI performs the classification via the OpenAI API and pushes the routed ticket back into the helpdesk through a custom REST API call. Workflow Orchestration — The software layer that sequences the steps of a multi-system process: receive webhook → call AI API → validate output → push to helpdesk → log for audit. The orchestration layer is model-agnostic, so it can route to OpenAI for quality or to an on-prem open-weight model for regulated data.

  • AI Invoice Processing in Healthcare: A Glossary of 15 Key Terms

    Before/After Baseline

    A before/after baseline is a set of metrics measured before and after the AI system is deployed to quantify its impact. For a healthcare organization, this includes cycle time (the time from invoice receipt to payment), error rate (the percentage of invoices requiring manual correction), and cost per invoice. The baseline is established during the process audit and used to measure the ROI of the AI system after the fixed-scope pilot and rollout. In a 2,000+ employee organization, even a 10% reduction in cycle time can save thousands of hours annually, making the baseline a critical tool for justifying the investment in AI automation.

    Document Extraction Pipeline

    A document extraction pipeline is a series of steps that convert unstructured or semi-structured documents, such as invoices, into structured data. For a healthcare organization, this pipeline includes steps like OCR (optical character recognition), layout analysis, field extraction, and data validation. The pipeline is built using LangChain and LangGraph, with human-in-the-loop checks for any fields that fall below a confidence threshold. In a HIPAA-regulated environment, the pipeline must ensure that patient-identifiable information is not exposed to cloud-based models, requiring the use of open-weight models on the client’s own hardware for sensitive data.

    Data Enrichment and Cleanup

    Data enrichment and cleanup in this context refers to the automated process of standardizing, validating, and augmenting raw invoice data before it enters the ERP. This includes mapping vendor names to master data, converting currency to the reporting currency, and flagging discrepancies in tax codes. For a 2,000+ employee organization, this step reduces manual data entry errors and ensures that monthly reporting is based on clean, consistent data. In a healthcare setting, data enrichment also involves mapping billing codes to the correct regulatory categories, ensuring that the data is compliant with HIPAA and other relevant regulations.

    HIPAA Compliance

    HIPAA compliance in this context means that the AI system must protect patient-identifiable information and ensure that data is not stored or processed in ways that violate the Health Insurance Portability and Accountability Act. For a healthcare organization, this requires using open-weight models on the client’s own hardware for any data that contains patient information, while using cloud-based models for non-sensitive data. The system must also include audit logs and access controls to track who accessed what data and when. In Austria, where data protection laws are strict, HIPAA compliance is often supplemented by GDPR requirements, making the compliance landscape even more complex.

    Human-in-the-Loop Workflow

    A human-in-the-loop workflow means the AI model drafts or classifies the data, but a human operator reviews and approves any output that touches financial records, patient-identifiable information, or contractual terms. For a 2,000+ employee healthcare organization, this ensures that while the system processes 90% of invoices automatically, the remaining 10% containing complex billing codes or HIPAA-sensitive data are routed to a finance team member for final sign-off before posting to the ERP. This approach balances the speed of AI automation with the accuracy and compliance required in a regulated environment.

    LangChain and LangGraph

    LangChain provides the foundational abstractions for connecting large language models to external tools and data sources, while LangGraph extends this by allowing developers to define stateful, multi-step workflows with explicit control flow. In a document extraction pipeline, LangChain handles the initial parsing and vector retrieval, whereas LangGraph manages the conditional logic that determines whether a parsed invoice requires human review or can be auto-approved based on confidence thresholds. This combination allows the system to handle complex workflows with precision, ensuring that each step is auditable and that the system can adapt to changes in invoice formats or regulatory requirements.

    Managed AI Operations

    Managed AI operations is a delivery model where the vendor not only builds the AI system but also monitors, maintains, and optimizes it after deployment. For a healthcare company, this includes tracking model performance, updating prompts as invoice formats change, and ensuring that the human-in-the-loop workflow remains efficient. This model is critical for scaling operations without new hires, as it shifts the burden of AI maintenance from the client’s IT team to the vendor. In a 2,000+ employee organization, managed operations ensure that the AI system continues to perform at a high level as the volume of invoices and the complexity of the data increase.

  • On-Premise AI Invoice Processing for Austrian Healthcare: A 2-Week Pilot

    The Problem: Manual Data Entry in Austrian Healthcare Finance

    Finance teams in Austrian healthcare and medtech companies face a persistent bottleneck: manual data entry from invoices. For a company of 201-500 employees, this means dozens of hours per week spent transcribing vendor details, line items, and tax codes into the ERP. The risk is not just cost; it is error. A single misclassified VAT code can trigger an audit finding under Austrian tax law. The goal is to replace this manual process with an AI workflow that extracts data, enriches it with vendor master data, and posts it to the ledger. This must be done on-premise to comply with GDPR, ensuring patient data on invoices never leaves the building. The timeline is tight: two weeks to a working pilot.

    Prerequisites for a 2-Week Pilot

    • On-premise GPU server: Minimum 24 GB VRAM (e.g., NVIDIA A5000 or RTX 4090) for running 7B-13B parameter open-weight models.
    • ERP API access: A stable REST API or webhook endpoint for your accounting system (SAP, Dynamics, or Lexware).
    • Baseline data: At least 500 historical invoices with their correct ledger entries to measure accuracy.
    • Legal review: A DPO or legal counsel to approve the GDPR Article 30 record of processing activities.
    • Network isolation: A dedicated VLAN for the AI server to prevent data exfiltration.
    • Human-in-the-loop workflow: A defined process for finance staff to review and approve AI-extracted data.

    Steps 1-3: Deployment, Preprocessing, and Fine-Tuning

    Step 1: Deploy the open-weight model on-premise.
    Install Ollama or vLLM on your GPU server. Pull a 7B or 13B parameter model (e.g., Llama 3 8B or Mistral 7B). Configure the model to run in a secure, isolated container. Ensure the server is on a dedicated VLAN with no internet access except for model updates. Test the inference speed; it should process an invoice in under 5 seconds.

    Step 2: Build the invoice preprocessing pipeline.
    Use a library like PyMuPDF to extract text from PDF invoices. Implement a rule-based filter to strip personal data (names, addresses) that is not required for the ledger entry. This satisfies GDPR data minimization. Store the cleaned text in a local database.

    Step 3: Fine-tune the model on your invoice data.
    Use your 500 historical invoices to fine-tune the model. Focus on the specific fields you need: vendor name, invoice number, line items, total, and VAT rate. Use a low learning rate (1e-5) to avoid overfitting. Evaluate the model on a holdout set of 50 invoices. Aim for 95% accuracy on key fields.

    Steps 4-6: ERP Integration, Human-in-the-Loop, and Pilot

    Step 4: Integrate with the ERP via REST API.
    Build a Python service that takes the extracted data and sends it to your ERP’s REST API. Use OAuth 2.0 for authentication. The payload should include the invoice ID, vendor, line items, and tax breakdown. Implement a webhook to notify the finance team when an invoice is processed. If the API fails, queue the data and retry with exponential backoff. Log all API calls for audit purposes.

    Step 5: Implement the human-in-the-loop workflow.
    Configure the system to route invoices with a confidence score below 95% to a human reviewer. Use a simple web interface for finance staff to approve or correct the data. Ensure the interface clearly shows the AI’s confidence score and the original invoice image. This step is critical for GDPR compliance and error prevention.

    Step 6: Run the pilot with 10-20% of invoice volume.
    Start with a small subset of invoices to validate the pipeline. Monitor the accuracy, speed, and rejection rate. Collect feedback from the finance team. Adjust the model or preprocessing pipeline based on the feedback. Do not scale to 100% volume until the error rate is below 2%.

    Common Pitfalls and How to Detect Them

    • Hallucination in vendor details: The model invents a vendor name or misclassifies a tax code. Detect this by monitoring the confidence score. If the score for a field drops below 95%, route the invoice to a human reviewer.
    • Data leakage: Personal data is not stripped before processing. Detect this by auditing the logs for any personal data in the model’s context window. Ensure the preprocessing pipeline is working correctly.
    • ERP API downtime: The ERP API is down, and the system drops invoices. Detect this by monitoring the API health and implementing a queue with exponential backoff. Ensure the system does not lose data during outages.
    • Model drift: The model’s accuracy degrades over time as invoice formats change. Detect this by tracking the rejection rate. If the rate increases, retrain the model with new data.

    Conclusion: From Pilot to Managed Operations

    The 2-week pilot is a validation, not a full rollout. Once the pilot is successful, the next step is to scale to 100% of invoice volume and add new invoice types. This should take 2-4 weeks. After that, move to managed AI operations, where a partner handles monitoring, retraining, and updates. The goal is to reduce manual data entry by 80-90% and cut cycle time from days to hours. The on-premise architecture ensures GDPR compliance, and the human-in-the-loop workflow ensures accuracy. The next logical step is to extend the AI workflow to other finance processes, such as expense reports or purchase orders.

  • Document Extraction Pilot for E-Commerce Operations in Austria

    The Operational Bottleneck: Manual Order and Shipment Data Entry

    E-commerce and retail operations teams in Austria face a persistent bottleneck: order and shipment status updates from suppliers arrive in inconsistent formats—PDFs, scanned images, email attachments, and portal exports. Manual extraction and data entry into SAP or Microsoft Dynamics consumes 30-45 minutes per batch, with error rates averaging 2-4% that cascade into delayed customer notifications and reconciliation headaches.

    A fixed-scope pilot addresses this by automating one specific workflow within a four-week window. The engagement starts with a process audit that maps your current document flow, measures baseline cycle time and error rate, and identifies the highest-ROI extraction targets. From there, the team builds a document extraction pipeline using the OpenAI API for its strong performance on varied layouts, integrates it with your existing ERP via native APIs, and validates results against your baseline metrics.

    The deliverable is not a new system but a faster, more accurate version of the workflow you already run. Senior operations staff move from data entry to exception handling and supplier relationship management, while the AI layer handles the repetitive extraction and mapping work.

    Four-Week Pilot Structure: From Audit to Validated Pipeline

    The four-week timeline follows a structured sequence. Week one covers the process audit: the team reviews 50-100 sample documents from your supplier base, maps data fields to your ERP schema, and establishes the baseline metrics—current cycle time per batch, error rate, and staff hours consumed. This phase also confirms compliance requirements under the EU AI Act, including transparency logging and human oversight protocols for data that affects financial records.

    Weeks two and three handle model configuration and integration. The OpenAI API is tuned for your specific document types, with prompt engineering and post-processing rules to handle edge cases like merged invoices or multi-page shipments. The extraction pipeline connects to SAP or Microsoft Dynamics through their standard APIs, writing validated data directly to the relevant tables. Human-in-the-loop review queues are configured so that low-confidence extractions route to staff for approval before ERP sync.

    Week four focuses on validation and handover. The team processes a full week’s worth of live documents, compares results against the baseline, and documents the error rate, cycle time improvement, and any remaining edge cases. The handover package includes runbooks, model version records, and escalation procedures for ongoing managed operation.

    Model-Agnostic Architecture: OpenAI API and Open-Weight Options

    The architecture is deliberately model-agnostic, but the OpenAI API serves as the default for quality-critical extraction tasks. Its strength lies in handling varied document formats—scanned PDFs with mixed layouts, email attachments with inconsistent headers, and portal exports with variable column structures—without requiring custom OCR preprocessing for each format.

    For regulated data that cannot leave the building, the same pipeline runs on open-weight models deployed on your own hardware. This configuration maintains the same integration points and human-in-the-loop workflows while ensuring data sovereignty. The trade-off is higher initial setup effort and potentially lower accuracy on edge cases, which the human review queue compensates for.

    The pipeline plugs into your existing SAP or Microsoft Dynamics ERP through their standard APIs rather than replacing them. Extracted data maps to your existing data structures: order numbers to sales order tables, shipment dates to delivery schedule lines, status codes to your internal workflow states. No ERP migration or reconfiguration is required. The AI layer sits alongside your current systems, handling the extraction and mapping work while your ERP continues to manage the downstream business logic.

    EU AI Act Compliance: Transparency and Human Oversight

    The EU AI Act classifies document extraction systems as limited-risk AI, requiring transparency about AI involvement and human oversight for decisions that affect financial records or customer commitments. For e-commerce operations in Austria, this means the system must log its actions, maintain records of model versions and training data, and allow human review before extracted data syncs to the ERP.

    The pilot ships with compliance documentation built in: action logs showing which documents were processed, confidence scores for each extraction, and a review trail for any human approvals. Model version records track which API version or open-weight model was used for each batch, supporting audit requirements. The human-in-the-loop workflow ensures that anything touching money, health data, or contracts requires explicit staff approval before ERP sync.

    For a 501-2000 employee company, this compliance layer adds minimal overhead to the four-week timeline. The documentation and logging are configured during the integration phase, and the review queue is part of the standard human-in-the-loop setup. The result is a system that meets EU AI Act requirements without requiring a separate compliance project or legal review cycle.

    Measuring Success: Cycle Time, Error Rate, and Staff Hours

    The pilot’s success is measured against the baseline established in week one. Typical targets for order and shipment status extraction include reducing cycle time from 30-45 minutes per batch to under 10 minutes, cutting error rates from 2-4% to under 0.5%, and freeing 60-80% of the staff hours previously consumed by manual data entry.

    The before/after comparison uses the same document samples processed through both the manual and AI-assisted workflows. Cycle time measures the elapsed time from document receipt to ERP sync. Error rate counts the number of fields requiring correction after initial extraction, divided by total fields processed. Staff hours are tracked through time-stamped review queues, showing how much time staff spend on exception handling versus routine data entry.

    The handover package includes a validation report with these metrics, a runbook for daily operations, and escalation procedures for edge cases. The managed operation phase continues with monthly performance reviews, model updates as supplier document formats change, and support for new document types as your supplier base evolves. The goal is not a one-time automation but a continuously improving AI-native operations layer that scales with your business.

  • 3-Month AI Contract Review Pilot for a 51-200 Person B2B SaaS Firm in Austria

    The Problem: Manual Contract Review Bottlenecks in Mid-Sized B2B SaaS

    Your legal and compliance team spends 12-15 hours per week manually extracting key terms from vendor contracts, flagging non-standard language, and drafting review notes. For a 51-200 person B2B SaaS firm in Austria, this manual work creates a bottleneck: contracts sit in review queues for 3-5 days, and data entry errors propagate into your CRM and ERP. The problem is not a lack of legal expertise but a lack of automation for repetitive extraction and classification tasks. A retrieval-augmented knowledge assistant, powered by Anthropic Claude API and integrated with your existing Notion or Confluence workspace, can reduce this cycle time to under 2 hours per contract while maintaining human approval for all final decisions. This article walks you through a 3-month pilot that replaces manual data entry with an AI-assisted workflow, delivered by a dedicated AI team.

    Prerequisites: What You Need Before Step 1

    • Document inventory: A complete list of active contracts, SLAs, and compliance checklists stored in Notion or Confluence. You need at least 200 documents to build a meaningful retrieval index.
    • Baseline metrics: Measure current cycle time (from contract receipt to approved review) and error rate (percentage of contracts requiring rework due to missed terms). Record these numbers before the pilot starts.
    • API access: Valid API keys for Anthropic Claude, Notion, and Confluence. For Notion, use the internal integration token; for Confluence, use the personal access token with read permissions on your contract spaces.
    • Human approval workflow: Define which decisions require human sign-off. For contract review, this includes any clause that touches payment terms, liability, termination, or data handling. Document this in a one-page policy.
    • Dedicated team: A technical lead, a prompt engineer, and a product designer who will work with your legal and compliance staff throughout the 3-month pilot.

    Step 1: Audit Your Contract Review Workflow

    Map every contract review task your team performs today. For a B2B SaaS firm, this typically includes: receiving a vendor contract, extracting key terms (payment schedule, termination clause, liability cap, data handling provisions), comparing against your standard template, flagging non-standard language, drafting review notes, and entering data into your CRM. Time each task. Identify which tasks are repetitive and rule-based, suitable for automation. For this pilot, focus on extraction and flagging, not final legal judgment. The output is a one-page process map with task durations and error rates. This map becomes the baseline for measuring ROI after the pilot.

    Step 2: Build the Retrieval Layer Over Your Document Store

    Build a vector database of your contract documents. Use Notion or Confluence APIs to pull all contract documents into a staging area. Chunk each document into 500-800 token passages, preserving section headers as metadata. Embed these passages using Anthropic’s embedding model or a compatible open-weight model. Store the embeddings in a vector database like Pinecone, Weaviate, or Qdrant. For a 51-200 person firm, this typically means indexing 200-500 contracts, which takes 2-3 hours of compute time. The retrieval layer should return the top 5 most relevant passages for any query, with a similarity threshold of 0.75 or higher to avoid low-confidence matches.

    Step 3: Configure the Anthropic Claude API for Contract Review

    Configure Anthropic Claude API as the reasoning engine. Use the Claude 3.5 Sonnet model for contract review tasks, as it balances quality and cost. Set the system prompt to instruct the model to answer only using the retrieved passages, to cite the source document and section for every claim, and to flag any clause that deviates from your standard template. Set the temperature to 0.1 for deterministic outputs. For high-stakes decisions, such as liability caps or termination clauses, the model should output a structured JSON object with fields for clause text, risk level, and suggested redline. This structure makes it easy for your legal team to review and approve.

    Step 4: Integrate with Notion or Confluence for Human-in-the-Loop Review

    Build a chat interface that sits on top of your Notion or Confluence workspace. For Notion, use the Notion API to create a database view that displays contract metadata (client name, contract value, renewal date) alongside the assistant’s review notes. For Confluence, create a page template that includes a chat widget powered by the assistant. The interface should allow your legal team to ask questions like “What is the termination clause in the Acme Corp contract?” and receive an answer with citations. It should also allow them to approve or reject the assistant’s suggested redlines. Every human decision should be logged in a separate audit table, capturing the user, timestamp, and decision.

    Step 5: Run the Pilot and Measure Before/After Metrics

    Run the pilot for 4-6 weeks, targeting one contract review workflow. Measure cycle time and error rate weekly. Compare against your baseline. For a 51-200 person firm, you should see cycle time drop from 3-5 days to under 2 hours per contract, and error rate drop by 30-50%. Collect feedback from your legal and compliance team on the quality of the assistant’s suggestions. Adjust the retrieval parameters, system prompt, and chunking strategy based on this feedback. If the assistant misses a specific type of clause, add that clause type to the retrieval index and re-test. The goal is to reach a 90% accuracy rate on extraction tasks before scaling to other workflows.

  • AI Ticket Triage for Austrian Medtech: n8n, Zendesk, and GDPR in 4 Weeks

    The Problem: Triage Overhead in a Small Medtech Support Team

    A 51-200 employee medtech company in Austria typically runs its customer support on Zendesk or Intercom, with 3-8 agents handling 200-800 tickets per month. The tickets span billing inquiries, device technical issues, regulatory questions, and patient-related communications. The problem is not volume alone; it is the cognitive overhead of triage. Every agent reads each ticket, decides its category, assigns priority, and routes it to the right team. This manual classification takes 4-7 minutes per ticket, and error rates on misrouting hover around 8-12% in small teams without formalized playbooks.

    The AI maturity here is one process automated: the company has likely experimented with a chatbot or a basic keyword filter, but has not yet built a structured, measurable automation layer. The goal of this deep dive is to design a compliance-safe AI rollout that fits within a 4-week integration sprint, uses n8n orchestration to connect the AI model to the existing helpdesk, and handles multilingual support coverage in German, English, and secondary languages relevant to the Austrian market.

    The constraint that shapes every decision: GDPR. Patient data, device serial numbers linked to patients, and adverse event reports cannot be processed by a model whose training data or inference infrastructure is outside the company’s control. This is not a theoretical concern; it is the difference between a pilot that ships and one that stalls in legal review for three months.

    The Mechanism: n8n Orchestration with a Dual-Path Model Layer

    The architecture has three layers. The orchestration layer is n8n, self-hosted on the client’s infrastructure. n8n receives a webhook from Zendesk or Intercom when a new ticket is created, passes the ticket body to the AI model, receives a structured JSON response, and calls the helpdesk API to update the ticket’s tags, assignee, and priority. The entire round trip completes in 2-5 seconds.

    The model layer is deliberately model-agnostic. For ticket classification and routing, the quality bar is high enough to justify a frontier API: OpenAI GPT-4o or Anthropic Claude 3.5 Sonnet handle multilingual classification with strong accuracy on structured tasks. The prompt returns a JSON object with category, priority, suggested_assignee, and language_detected. If the ticket contains patient-identifiable data, the n8n workflow routes it to a locally hosted open-weight model (e.g., Llama 3 70B on the client’s GPU server) so that no patient data leaves the building. This dual-path design is the core of the compliance-safe approach.

    The integration layer uses the Zendesk or Intercom REST API. The n8n workflow calls PATCH /api/v2/tickets/{id} to update tags and assignee, and POST /api/v2/tickets/{id}/comments to post a first-response draft. All API calls use TLS 1.3, and n8n’s execution history is configured to exclude ticket body content from logs, satisfying GDPR Article 5(1)(f) integrity and confidentiality requirements.

    Zendesk/Intercom Webhook
            |
            v
       n8n Workflow (self-hosted)
            |
            +---> Language Detection (langdetect / model output)
            |
            +---> Sensitive Data Check (regex + model flag)
            |         |
            |         +-- No PII --> OpenAI / Anthropic API
            |         +-- PII present --> Local Llama 3 70B
            |
            v
       JSON: {category, priority, assignee, language}
            |
            v
       Zendesk/Intercom API (PATCH ticket, POST comment)
            |
            v
       Human-in-the-loop approval (if PII or high-risk category)
    
    ## Trade-offs: Model Choice, Human Oversight, and Multilingual Cost
    
    The first trade-off is **model quality versus data residency**. Using GPT-4o or Claude 3.5 Sonnet gives the highest classification accuracy (92-95% on structured ticket categorization), but it requires sending ticket text to a third-party API. For a medtech company, this is acceptable for non-patient tickets (billing, order status, general technical questions) but not for tickets containing patient names, device serial numbers linked to patients, or adverse event descriptions. The dual-path design resolves this: the n8n workflow runs a lightweight PII detection step (regex for Austrian ID formats, device serial patterns, and a model-based flag for health-related language) and routes sensitive tickets to the local model. The cost is a 15-20% accuracy drop on the local model for nuanced classification, which is mitigated by the human-in-the-loop approval step.
    
    The second trade-off is **automation depth versus human oversight**. Full automation (AI classifies, routes, and drafts the response without human review) would save the most time, but it violates GDPR Article 22 for any ticket with legal or significant effects. The compromise: the AI handles classification, routing, and first-response drafting for all tickets, but a human agent must approve any ticket flagged as containing PII, involving adverse events, or touching contractual terms. This adds 30-60 seconds of human review per sensitive ticket, but it is the price of compliance.
    
    The third trade-off is **multilingual coverage versus model cost**. Running a separate model per language is expensive and operationally complex. Instead, the workflow uses a single multilingual model for classification and language detection, then branches to language-specific response templates. This keeps the model call to one per ticket and avoids maintaining parallel rule sets.
    
    ## Recommendation: A 4-Week Sprint for Billing and Order Status Triage
    
    For a 51-200 employee medtech company in Austria, the recommendation is to start with **billing and order status tickets** as the first automation target. These typically account for 40-60% of ticket volume, carry minimal GDPR risk (no patient data), and have a clear, low-risk routing taxonomy. The 4-week sprint breaks down as follows:
    
    - **Week 1: Process audit and baseline.** Sample 200-300 historical tickets. Measure current cycle time (target: 4-7 min per ticket) and misrouting error rate (target: 8-12%). Define the ticket taxonomy: billing, order status, technical, regulatory, patient inquiry.
    - **Week 2: n8n workflow build.** Set up the self-hosted n8n instance. Build the webhook receiver, PII detection step, dual-path model routing, and JSON response parser. Test with synthetic tickets.
    - **Week 3: Helpdesk integration.** Connect the n8n workflow to Zendesk or Intercom via API. Implement the `PATCH` and `POST` calls. Build the human-in-the-loop approval flow: sensitive tickets are queued for agent review before the AI's routing action is applied.
    - **Week 4: Shadow-mode testing and go-live.** Run the AI in shadow mode for 5 business days: it classifies and routes tickets, but the human agent's action is the one that actually updates the ticket. Compare AI routing against human routing. If agreement is above 85%, go live with the AI handling routing and the human approving sensitive tickets.
    
    The measured outcome should be a 30-40% reduction in average cycle time for the automated category and a misrouting error rate below 5%. The pilot ships with a before/after baseline report that the client can use to justify the next automation phase.
  • AI Process Audit vs. Single-Process RAG Pilot: A Healthcare Company in Austria

    What Is Being Compared

    The two options are not alternatives in a vacuum; they are different scopes of the same engagement. Option A is a full AI process audit and roadmap: Forfis maps every back-office and customer-facing workflow, measures baseline cycle time and error rate on each, and produces a prioritised automation roadmap across the company. Option B is a single-process pilot: one workflow — here, an internal knowledge search assistant built on retrieval-augmented generation over the company’s Google Workspace documents — is scoped, built, and measured in a fixed three-month window. Both use the OpenAI API as the model layer, both integrate through existing APIs rather than replacing tools, and both ship with a human-in-the-loop approval gate. The difference is breadth: Option A covers the whole operation; Option B covers one process and proves the pattern before scaling.

    Criteria for the Comparison

    The judgment rests on seven criteria that matter to a 201-500 person healthcare company in Austria with no specific compliance mandate and a three-month timeline:

    • Time to first measurable value — how many weeks until a workflow runs with a before/after baseline.
    • Upfront cost — the fixed-scope fee for the audit or the pilot, before managed operation.
    • Breadth of coverage — how many workflows are mapped or automated by the end of the engagement.
    • Integration surface — which existing systems (Google Workspace, CRM, helpdesk) the AI layer touches.
    • Model-agnostic flexibility — whether the architecture can swap OpenAI for an open-weight model on client hardware if data-residency needs emerge.
    • Human-in-the-loop overhead — how many approval steps a support agent must complete per query.
    • Scalability path — how the engagement extends from one process to the next without re-scoping.

    Side-by-Side Comparison

    Criterion Option A: Full Audit + Roadmap Option B: Single-Process RAG Pilot
    Time to first measurable value 8-10 weeks (audit) + 4-6 weeks (first pilot) 3 weeks (audit slice) + 4-6 weeks (pilot)
    Upfront cost Higher: covers all workflows, multiple integrations Lower: one workflow, one integration (Google Workspace)
    Breadth of coverage All back-office and customer-facing workflows mapped One workflow: internal knowledge search
    Integration surface CRM, ERP, helpdesk, Google Workspace, messaging Google Workspace (Gmail, Drive, Calendar)
    Model-agnostic flexibility Full: per-workflow model selection Full: OpenAI API default, swappable
    Human-in-the-loop overhead Varies by workflow; set during audit Light: internal search, no money/health/contract decisions
    Scalability path Roadmap already built; next process is a scheduling decision Must re-scope for the second process

    When Each Option Wins

    Option B wins when the company’s immediate pain is concentrated in one workflow and the three-month timeline is a hard constraint. A 201-500 person healthcare company whose support team spends 25-40 minutes per ticket searching through Drive documents and Gmail threads will see a measurable cycle-time reduction within six weeks of the pilot starting. The RAG assistant indexes the existing Google Workspace content, retrieves the relevant SOP or device manual passage, and returns a grounded answer with a citation. The support agent approves the answer before sending it to the requester. No new hires are needed; the senior staff who previously handled routine knowledge lookups are freed to work on complex cases. The before/after baseline on time-to-answer and accuracy is captured in the first two weeks and compared at the end of the pilot.

    Option A wins when the company has multiple workflows with similar automation potential — invoice processing, document extraction, ticket triage, data entry — and the leadership team wants a single prioritised roadmap rather than a sequence of ad-hoc pilots. The audit maps all of them, measures baselines on each, and ranks them by expected cycle-time reduction and error-rate improvement. The cost is higher, but the company avoids the re-scoping overhead of going back to Forfis for every second process. For a company that has already automated one process and is now asking “what next?”, the audit is the natural next step.

    Recommendation for This Scenario

    For the scenario as specified — a 201-500 person healthcare and medtech company in Austria, no compliance mandate, three-month timeline, one process already automated, need to free senior staff from routine work, and a Google Workspace integration — Option B is the correct starting point. The company has already proven the pattern with one automated process; the next step is to apply the same pattern to internal knowledge search, not to commission a full audit that would extend the timeline beyond three months. The RAG pilot on Google Workspace is the highest-leverage single workflow for a support-heavy operation: it directly reduces the time senior staff spend on routine lookups, it integrates with the tools the team already uses, and it ships with a measured baseline that justifies the next investment. Once the pilot is live and the before/after numbers are in hand, the company can decide whether to commission the full audit (Option A) to map the remaining workflows, or to run a second pilot on a different process. The model-agnostic architecture means that if data-residency requirements emerge later, the OpenAI API layer can be swapped for an open-weight model on the company’s own hardware without re-architecting the integration.

  • EU AI Act-Compliant Invoice Processing Pilot for Austrian Logistics

    The Problem: Manual Invoice Processing in Austrian Logistics

    A 501-2000 employee logistics and supply chain company in Austria processes supplier invoices across German, Hungarian, and Polish. Each invoice passes through manual data entry, cross-checking against purchase orders, and approval in the ERP. Cycle time averages 4.2 days from receipt to payment-ready status, with a 3.1 percent field-level error rate that triggers payment delays and supplier disputes. The company has run isolated AI pilots on document extraction but has not connected them to the approval workflow or measured the operational impact. The EU AI Act, in force since August 2024, now requires transparency and human oversight for AI systems handling financial data. You need a compliance-safe rollout that integrates with existing Slack or Microsoft Teams channels, supports multilingual invoices, and ships with a measured before/after baseline within 8 weeks.

    Prerequisites Before You Start

    Before step 1, confirm the following are in place:

    • Historical invoice dataset: at least 500 invoices in each target language (German, Hungarian, Polish) with ground-truth field values for validation.
    • ERP API access: read and write credentials for your accounting system (SAP, Microsoft Dynamics 365, or similar) to post approved invoices.
    • Slack or Microsoft Teams workspace: a dedicated channel where the AI will post extraction results and request approvals.
    • Named approvers: at least two human approvers per invoice stream, with defined escalation paths.
    • Anthropic Claude API key: provisioned and scoped to the pilot project, with usage limits set to prevent cost overruns.
    • Baseline metrics: current cycle time (days) and error rate (percent) measured over the last 90 days, documented in a one-page report.

    Step 1: Audit the Invoice Stream and Set the Baseline

    Run a 2-week process audit on the invoice stream you will automate. Map every step from invoice receipt to payment-ready status in the ERP. Record the average cycle time, the number of manual touchpoints, and the error rate by field type (vendor name, amount, tax ID, line items). Use the historical dataset to label 100 invoices per language with correct field values. This becomes your validation set. The audit output is a one-page document with the baseline numbers and the specific fields the AI must extract. You are not building a system yet; you are defining the problem precisely so the pilot has a measurable target.

    Step 2: Build the Extraction Pipeline with Claude API

    Build the extraction pipeline using the Anthropic Claude API. Configure the model to extract vendor name, invoice number, date, line items, total amount, and tax ID from the invoice PDF or image. Set the temperature to 0 for deterministic output. Use structured output (JSON schema) so the response is parseable without regex. For multilingual support, include the language code in the prompt and validate that the model handles Hungarian and Polish field labels correctly. Test on 50 invoices per language from your validation set. Target: field-level accuracy above 95 percent. If any language falls below threshold, adjust the prompt or add few-shot examples before proceeding.

    Step 3: Wire the Approval Workflow into Slack or Teams

    Integrate the pipeline with your Slack or Microsoft Teams workspace. When an invoice is processed, the AI posts a card to the dedicated channel showing the extracted fields, confidence scores, and a link to the ERP record. For exceptions (confidence below 80 percent or mismatch with the purchase order), the AI sends a direct message to the approver with approve/reject buttons. The approver’s action triggers the ERP update via the API. Log every interaction with timestamp, user ID, and model version. This log is your EU AI Act audit trail under Article 50. The integration uses the platform’s webhook and message API, not a custom chatbot framework.

    Step 4: Run the Parallel Operation and Measure

    Run the AI pipeline in parallel with the manual process for 2 weeks. Every invoice goes through both paths. Compare the AI’s extraction against the manual entry and the ground-truth data. Track cycle time from receipt to approval and the error rate by field type. The pilot succeeds if the AI reduces cycle time by at least 40 percent and keeps the error rate below 2 percent. Document the results in a before/after report with specific numbers: for example, cycle time drops from 4.2 days to 2.1 days, and error rate drops from 3.1 percent to 1.4 percent. This report is the deliverable of the fixed-scope pilot.

    Common Pitfalls and How to Detect Them

    Three failure modes appear consistently in invoice processing pilots:

    • Language drift: the model handles German well but misreads Hungarian tax fields. Detect it by running the validation set weekly and alerting if any language’s accuracy drops below 95 percent.
    • Approval bottleneck: approvers do not respond to Slack messages within 24 hours, negating the cycle-time gain. Detect it by tracking the median approval latency and setting a 4-hour SLA.
    • ERP sync failure: the AI posts to Slack but the ERP update fails silently. Detect it by adding a reconciliation job that compares the number of approved invoices in Slack against the ERP records every 6 hours.
  • How an Austrian Medtech Firm Cut First-Response Time to 38 Minutes in Four Weeks

    Background: A 2,400-Person Medtech Firm in Austria

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements in healthcare and medtech. We do not name real clients. The company described here is a mid-sized Austrian medtech firm with roughly 2,400 employees, operating in the DACH region and serving hospital networks in Austria, Germany, and parts of the UK. It sells diagnostic equipment and consumables, and its customer support team handles order confirmations, shipment tracking, and return requests. The support stack is a mix of a legacy helpdesk, an ERP for order management, and a CRM for account records. The company is not a digital-native; its IT team maintains the existing systems but has no in-house AI capability. The trigger for change was a 22 percent year-over-year increase in support ticket volume, driven by a new product line and a shift toward direct-to-hospital sales. The support team of 34 agents was already at capacity, and first-response times had drifted past the 4-hour internal target.

    The Challenge: 4.2-Hour First Responses and a HIPAA Constraint

    The core problem was not a lack of agents but a lack of speed in the first step: reading the inbound document, extracting the relevant fields, and drafting a response. Each ticket arrived as a PDF or scanned image, often a mix of an order confirmation, a shipping label, and a handwritten note from the hospital’s procurement office. An agent had to open the file, read it, cross-reference the order number in the ERP, check the shipment status, and type a reply. The average cycle time from receipt to first response was 4.2 hours, with a peak of 9 hours during Monday mornings. The error rate on manual extraction was 11 percent, mostly misread order numbers or confused shipment references. The compliance constraint was non-negotiable: the company serves US-based hospital partners and is subject to HIPAA. Any document containing patient-identifiable information, even indirectly through a hospital’s internal reference number, had to stay on the client’s own infrastructure. The deadline was four weeks, aligned to the start of the next fiscal quarter, when the support team would be restructured.

    Approach: A Four-Week Pilot with a Dedicated AI Team

    Forfis deployed a dedicated AI team of four: two backend engineers, one product designer, and one engineer focused on the integration layer. The first week was a process audit. The team sampled 800 tickets from the prior quarter, categorized them by document type, and measured the baseline cycle time and error rate. The audit identified three document types worth automating: order confirmations, shipment status requests, and return authorizations. The pilot scope was fixed to the first two: order confirmations and shipment status. The architecture used a two-tier model setup. Open-weight models, fine-tuned on the client’s historical documents, ran on the client’s own GPU server for all extraction tasks involving PHI. A commercial API model handled the drafting of the first-response text, but only after the PHI fields had been stripped by the on-premises layer. The pgvector index stored embeddings of the client’s order history and shipment records, enabling the system to match an extracted order number to the correct ERP record in under 18 milliseconds. The integration layer was a set of custom REST API endpoints and webhooks that wrote back to the helpdesk and ERP without replacing either system.

    Outcome: 38-Minute First Responses and a 3.4 Percent Error Rate

    By the end of week four, the pilot was in production for the two in-scope document types. First-response time dropped from 4.2 hours to a median of 38 minutes, with the 95th percentile at 2 minutes 14 seconds. The extraction error rate fell from 11 percent to 3.4 percent, with the remaining errors concentrated in handwritten notes, which the system correctly flagged for human review rather than guessing. The human-in-the-loop layer caught 14 percent of documents in the first week, dropping to 4.8 percent by week four as the model adapted to the client’s document formats. The support team reported that agents spent 60 percent less time on data entry and cross-referencing, redirecting that time to complex cases. The cost per ticket, measured as fully loaded labor cost divided by tickets handled, fell by an estimated 31 percent. The client’s compliance officer confirmed that no PHI left the on-premises environment during the pilot. The system handled 1,200 tickets per week at peak, a 40 percent increase over the pre-pilot volume, without adding headcount.

    Lessons for Teams in Regulated, Document-Heavy Support

    • Fix the baseline before you build. The two-week pre-pilot measurement of cycle time and error rate is not optional. Without it, the post-pilot comparison is anecdotal, and the client cannot justify the rollout to the board. Forfis treats the baseline as a deliverable in its own right.
    • Scope the pilot to one or two document types, not a whole department. A four-week timeline is realistic only if the scope is narrow. Expanding to return authorizations, warranty claims, and invoice disputes in the same window would have pushed the timeline to ten weeks and muddied the metrics.
    • Put the PHI boundary in the architecture, not in the policy. The on-premises model for PHI and the API model for non-PHI text are separated at the routing layer. A policy document saying “do not send PHI to the API” is not a control. The code enforces it.
    • Human-in-the-loop is a tuning parameter, not a fallback. The confidence threshold for routing to a human is adjusted weekly during the pilot. Starting too high (routing 40 percent of documents to humans) defeats the purpose; starting too low (routing 2 percent) risks errors. The 12-to-5 percent drop over four weeks reflects this tuning.
    • The integration layer is the real product. The LLM is a commodity. The REST API adapters, webhook handlers, and pgvector index that connect the model to the client’s existing helpdesk and ERP are what make the system work in production. Budget engineering time accordingly.