Tag: Austria

  • AI Process Audit vs. Triage Pilot: A Two-Week Comparison for Austrian Logistics

    What Is Being Compared

    The two options under comparison are not competing products but two distinct automation workstreams that a mid-size logistics firm in Austria would typically sequence within a single AI maturity roadmap. Option A is an AI process audit and roadmap engagement: a structured assessment of existing back-office and support workflows that identifies which processes have the highest volume, error rate, and cycle time, then produces a prioritized automation sequence. Option B is a round-the-clock customer response pilot: a fixed-scope, two-week deployment of an AI triage layer on the firm’s existing helpdesk, integrated with Slack or Microsoft Teams, using the Anthropic Claude API to classify and route inbound tickets and draft first responses. The firm operates in logistics and supply chain, employs 51–200 people, has no specific regulatory compliance mandate, and its primary need is to cut first-response time on customer support tickets. The audit (Option A) is the prerequisite that determines whether the triage pilot (Option B) is the correct first deployment, or whether document extraction on carrier invoices should come first.

    Criteria for Judgment

    Eight criteria determine which option delivers measurable value first in a two-week window:

    • Time-to-first-measurable-result: how many days from kickoff to a quantified before/after metric.
    • Baseline dependency: whether the option requires a pre-existing measurement of cycle time and error rate to demonstrate improvement.
    • Integration surface: number of existing systems (helpdesk, CRM, Slack/Teams, ERP) that must be connected via API.
    • Model dependency: whether the option is tied to a specific LLM provider or is model-agnostic.
    • Human-in-the-loop threshold: the minimum error rate below which auto-approval is safe.
    • Scalability across departments: how easily the output extends from customer support to claims, carrier coordination, or back-office.
    • Cost structure: fixed fee versus usage-based API cost, and the engineering hours required for integration.
    • Rollout risk: the probability that the pilot’s success does not translate to a full deployment without rework.

    Side-by-Side Comparison

    Criterion Option A: AI Process Audit & Roadmap Option B: Round-the-Clock Triage Pilot
    Time-to-first-measurable-result 10–14 days (audit report + prioritized sequence) 5–7 days (shadow-mode baseline vs. AI-assisted response)
    Baseline dependency Produces the baseline; does not consume one Consumes the baseline; requires 3-day pre-pilot measurement
    Integration surface Read-only access to helpdesk, CRM, Slack/Teams logs Write access to helpdesk API + Slack/Teams webhook; 2–3 system connections
    Model dependency None (analytical, not generative) Anthropic Claude API (claude-sonnet-4-20250514 or claude-3-5-sonnet)
    HITL threshold N/A Error rate < 5% on 200-ticket sample before auto-approve
    Scalability across departments Directly maps to multi-department rollout sequence Extends via parameterized prompts; requires new baseline per department
    Cost structure Fixed fee, EUR 6,000–10,000 for 2 weeks Fixed fee EUR 8,000–15,000 + API usage (~EUR 200–300/month at 500 tickets/day)
    Rollout risk Low; output is a document, not a live system Medium; live integration must survive API changes and volume spikes

    Scenario-by-Scenario Verdict

    When Option A wins first. If the firm has never measured its support workflow, the audit is the correct starting point. A logistics company handling 400–800 inbound tickets per week across shipment status, delivery exceptions, and billing disputes cannot demonstrate a first-response-time improvement without a baseline. The audit captures that baseline in days 1–3, identifies which ticket categories have the highest volume and error rate, and determines whether triage or document extraction on carrier invoices should be piloted first. In this scenario, the audit also reveals whether the existing helpdesk has a clean REST API or whether a Slack/Teams bridge is needed—information that directly affects the pilot’s integration scope and timeline. Without the audit, the two-week pilot risks measuring against a baseline that does not reflect steady-state workload.

    When Option B wins first. If the firm already has a documented baseline—average first-response time of 4.2 hours, routing error rate of 12%—the triage pilot can start immediately. The Claude API triage layer, integrated with the helpdesk and Slack/Teams, can be in shadow mode by day 5. For a 51–200 employee firm where the support team of 6–10 agents is the bottleneck, cutting first-response time from 4.2 hours to under 30 minutes for the top three ticket categories (status inquiries, delivery confirmations, tracking lookups) is the highest-impact single change. The pilot’s fixed scope means the firm commits to two weeks and a defined deliverable, not an open-ended engagement.

    Recommendation

    The sequencing recommendation. For a logistics firm in Austria with no compliance mandate and a two-week timeline, the correct sequence is: audit in week 1, triage pilot in week 2, compressed into a single fixed-scope engagement. The audit occupies days 1–3 and produces the baseline and the prioritized workflow list. The triage pilot occupies days 4–14, with shadow-mode testing on days 4–10, HITL validation on days 11–13, and the go/no-go review on day 14. This sequencing is feasible because the audit’s output (the baseline and the top-three ticket categories) is exactly the input the pilot needs. Attempting to run both in parallel would dilute measurement quality; running the audit alone would waste the two-week window without producing a live system.

    The explicit recommendation. Option B—the round-the-clock triage pilot using the Anthropic Claude API—is the correct primary deliverable for this scenario, but it is contingent on Option A’s audit output. The firm should contract a single fixed-scope engagement that bundles both: the audit as the first three days, the triage pilot as the remaining eleven. The pilot’s success criterion is a measured reduction in first-response time for the top three ticket categories, with a routing error rate below 5% on a 200-ticket validation sample. The integration targets the existing helpdesk and Slack or Microsoft Teams; no system is replaced. The model-agnostic architecture means that if the firm later moves to an open-weight model on its own hardware for a different workflow, the triage layer’s integration points remain unchanged.

  • Cut Compliance First-Response Time in 4 Weeks with n8n and Open-Weight Models

    The Problem: Compliance Queries Eat Hours You Cannot Afford to Lose

    Your legal and compliance team in a 201-500 person Austrian logistics firm spends an average of 4.2 hours per query answering the same 20 questions about customs clearance, carrier contracts, and GDPR data handling. You cannot hire more compliance staff without breaking your operating margin, and you cannot keep scaling operations by adding headcount. The problem is not a lack of knowledge; it is a lack of retrieval. The answers exist in your SharePoint folders, Confluence pages, and CRM records, but finding them requires a human to search, read, and synthesize. AI workflow automation with n8n orchestration solves this by building a retrieval-augmented search layer that sits on top of your existing documentation and posts answers directly into Slack or Microsoft Teams. The pilot runs in 4 weeks, uses open-weight models on your own hardware to keep GDPR-sensitive data inside your Austrian data center, and ships with a measured before/after baseline on cycle time and error rate. You do not replace your CRM, ERP, or helpdesk; you plug into them through their APIs.

    Prerequisites: What You Need Before Week 1

    Before you build the n8n workflow, you need five things in place. First, a knowledge corpus with at least 500 documents (SOPs, contracts, compliance checklists, FAQ pages) exported from SharePoint, Confluence, or a shared drive into a flat directory structure. Second, a vector database running on your own infrastructure: Weaviate, Qdrant, or pgvector on a PostgreSQL instance with at least 16 GB of RAM. Third, an inference endpoint for an open-weight model: Ollama or vLLM running Llama 3 8B or Mistral 7B on a GPU with 24 GB of VRAM (an NVIDIA A100 or a cloud instance with equivalent specs). Fourth, a Slack or Microsoft Teams workspace where the bot will post, with a dedicated channel (e.g., #compliance-questions) and a named owner for the human-in-the-loop review. Fifth, a GDPR compliance file: a Data Protection Impact Assessment (DPIA) drafted under Article 35 of the GDPR, a data processing agreement (DPA) if you use any third-party service, and a record of processing activities (ROPA) updated to include the new AI system. Without these five items, the pilot will stall in week 1.

    Step 1: Build the Retrieval Pipeline in n8n

    Export your knowledge corpus into a flat directory: one folder per document type (customs, contracts, GDPR, carrier agreements). Use a script to split each document into 512-token chunks with a 64-token overlap. Embed each chunk using a sentence-transformers model (e.g., all-MiniLM-L6-v2) and load the embeddings into your vector database. In n8n, create a new workflow and add a Slack Trigger node set to listen for messages in #compliance-questions. Add a Vector Store Search node (or an HTTP Request node to your Weaviate/Qdrant endpoint) with a similarity threshold of 0.80. Add an HTTP Request node that calls your local Ollama endpoint (http://localhost:11434/api/generate) with the retrieved chunks as context and the user’s question as the prompt. Add a Slack Post node that formats the answer with a citation to the source document. Test the workflow with 10 known questions before moving to the next step.

    Step 2: Add the Human-in-the-Loop Approval Gate

    In the n8n workflow, add an IF node after the LLM response that checks whether the answer touches money, health data, or a contract. If yes, route the message to a Slack Approval node that tags the compliance owner and waits for a @channel approve or @channel reject response. If no, post the answer directly. This is your human-in-the-loop gate. For the pilot, define three categories that always require approval: (1) any answer referencing a specific contract clause, (2) any answer involving personal data of a client or employee, (3) any answer about customs duties or tariff codes. Log every approval decision in a spreadsheet or a lightweight database (Postgres table approval_log with columns timestamp, question, answer, approver, decision). This log is your audit trail for GDPR Article 30 and your evidence for the before/after baseline.

    Step 3: Measure the Before/After Baseline

    Before you go live, measure the baseline. Pull 100 historical questions from your Slack or Teams archive from the last 90 days. For each question, record the time from the question being posted to the first verified answer being posted. Calculate the median and the 90th percentile. In a typical Austrian logistics firm, the median is 3.8 hours and the 90th percentile is 11.2 hours. Now run the n8n workflow on the same 100 questions in a test channel. Record the time from question to model output, and the time from model output to human approval (if applicable). Calculate the median and 90th percentile for the automated path. Your target: reduce the median from 3.8 hours to under 1.5 hours and the 90th percentile from 11.2 hours to under 4 hours. If the automated path does not beat the baseline on at least 70% of the 100 questions, your retrieval layer is not working. Tighten the similarity threshold, add metadata filters, or re-chunk the documents.

    Step 4: Deploy to Production and Monitor

    Deploy the n8n workflow to the production #compliance-questions channel. Set the workflow to run continuously (n8n’s built-in scheduler or a Docker container with restart: always). Enable n8n’s execution log and export it to a monitoring dashboard (Grafana or a simple Postgres view). Track three metrics daily: (1) cycle time from question to final answer, (2) error rate (percentage of answers flagged as incorrect by the compliance owner), (3) approval latency (time from model output to human approval). Alert if the error rate exceeds 10% over a rolling 7-day window or if the approval latency exceeds 30 minutes. In week 2, review the error log and retrain the retrieval layer: if a specific document type (e.g., carrier contracts) has a high error rate, re-chunk those documents with a smaller overlap (32 tokens instead of 64) and re-embed. In week 3, expand the knowledge corpus to include any new SOPs published during the pilot. In week 4, run the final baseline measurement and document the results.

    Common Pitfalls: Where the Pilot Breaks

    The most common failure is a hallucination loop: the model generates a confident answer that cites a document that does not exist or misstates a clause. You detect this by tracking the error rate on a weekly sample of 20 answers. If more than 10% are factually wrong, your retrieval threshold is too loose. Tighten it from 0.80 to 0.85 and add a metadata filter (e.g., only retrieve from the customs/ folder for customs questions). A second failure is knowledge staleness: your SOPs change but the vector index is not updated. You detect this by spot-checking 5 answers per week against the current SOPs. If an answer references a procedure that was updated in the last 30 days, re-embed the affected documents. A third failure is approval bottleneck: the human-in-the-loop review takes longer than the original manual process. You detect this by measuring the time from model output to approval, not just the time from question to model output. If approval latency exceeds 30 minutes, you have not actually cut response time. Reduce the number of questions that require approval by tightening the IF condition in Step 2.

  • LangGraph AI Agent for HR Workflow Orchestration in an Austrian Fintech

    The Problem: Fragmented HR Data Entry in a 30-Person Austrian Fintech

    A 30-person fintech in Vienna processes 40-60 onboarding documents per month: contracts, bank details, compliance attestations, and internal policy acknowledgments. Each document requires a human to extract fields, cross-reference against the HR system, and log the data into three separate tools. The median cycle time is 72 minutes per document, and the error rate on manual data entry sits at 4-6%, triggering rework and compliance risk under GDPR Article 5(1)(d) (accuracy of personal data). The problem is not volume but fragmentation: the data lives in PDFs, email threads, and a legacy HR system, and no single tool connects them. The automation target is not to replace the HR team but to eliminate the 12-15 hours per week of manual data entry and document routing that currently consume senior staff time. The constraint is strict: personal data cannot leave Austrian or EU jurisdiction, and any automated action affecting a candidate or employee requires human approval under GDPR Article 22.

    Mechanism: LangGraph State Machine and RAG Pipeline

    The architecture uses LangGraph as the orchestration layer and LangChain for LLM and vector store abstractions. LangGraph models the workflow as a stateful directed graph with nodes for intake, classification, RAG retrieval, draft generation, human approval, and dispatch. Each node is a Python function that receives and returns a state object. The graph supports conditional edges: if the classifier flags a document as high-risk (e.g., a contract amendment), the path routes to a senior reviewer; if it is a routine bank-detail update, it routes to a junior approver. The state persists in PostgreSQL via LangGraph’s checkpoint store, so the workflow survives process restarts. The RAG pipeline ingests internal policy docs, onboarding checklists, and HR system exports. Documents are chunked at 512 tokens with 64-token overlap, embedded using BGE-M3 (multilingual, supports German and English), and stored in pgvector. At query time, the agent retrieves the top-5 chunks, constructs a context-augmented prompt, and generates a structured JSON response with extracted fields and a confidence score. The LLM layer is model-agnostic: OpenAI GPT-4o handles general knowledge queries where no personal data is in the prompt, while Llama 3 70B running on the client’s own GPU server handles any task involving personal data, ensuring GDPR data residency.

    Trade-offs: Model Choice, Approval Granularity, and Integration Depth

    Three architectural choices dominate the trade-off space. First, model selection: using OpenAI or Anthropic APIs reduces infrastructure cost and improves quality on complex reasoning, but personal data in the prompt violates GDPR data residency for an Austrian company. The cost of using open-weight models on client hardware is a 15-20% drop in classification accuracy on edge cases and a one-time GPU server cost of EUR 8,000-12,000. Second, human-in-the-loop granularity: inserting an approval node after every agent action maximizes compliance but adds 5-10 minutes of latency per document. A tiered approach, where routine documents auto-approve after a 24-hour window and high-risk documents require immediate human review, reduces latency by 40% but requires a well-defined risk taxonomy. Third, integration depth: building a custom UI for HR staff gives full control but adds 2-3 weeks of development. Integrating with Slack or Microsoft Teams via their existing APIs (Slack Block Kit, Teams Adaptive Cards) reuses the tools the team already uses, cuts development time by 60%, and keeps the approval workflow in the channel where the document was originally shared. The Teams integration uses the Bot Framework with a webhook endpoint; the Slack integration uses a slash command that triggers the LangGraph agent via a REST API.

    Recommendation: 8-Week Integration Sprint for One Process

    For a 30-person Austrian fintech, the 8-week sprint follows a fixed sequence. Weeks 1-2: process audit. Map every HR document type, identify the three highest-volume workflows (typically onboarding data entry, policy acknowledgment tracking, and candidate status updates), and measure baseline cycle time and error rate. Weeks 3-4: build the LangGraph agent. Scaffold the state machine, implement the RAG pipeline, and connect to the HR system API. Deploy the open-weight model on the client’s hardware. Weeks 5-6: integrate with Slack or Teams. Build the interactive approval cards, test the webhook flow, and configure the checkpoint store. Weeks 7-8: pilot and measure. Run the agent on one workflow (e.g., onboarding document processing) for two weeks, with a human approving every action. Measure cycle time, error rate, and manual hours saved against the baseline. The pilot ships with a before/after report. The recommendation is to start with the workflow that has the highest volume and the lowest compliance risk, not the most complex one. For a fintech, that is usually routine onboarding data entry, not contract amendment review. The agent should be scoped to extract and classify, not to make decisions. Every output that touches a candidate’s or employee’s data must pass through a human approval node before it is written to the HR system or sent to the individual.

  • Austrian Insurtech Cuts Support Cycle Time 50% with Voice Agent and RAG Pilot

    Background: A 300-Person Austrian Insurtech

    This case study is a composite based on patterns observed across multiple engagements. We do not name real customers. The company described here matches the profile of a mid-sized Austrian insurtech: 300 employees, 12 years in operation, serving private and small-business customers across Austria and Germany. The stack includes a legacy CRM (Salesforce), an ERP (SAP), and a helpdesk (Zendesk). Internal documentation lives in Confluence, with some policy procedures in Notion. The company had been using basic rule-based chatbots for two years but had not moved to generative AI. The operations team was under pressure to scale support without adding headcount, as the Austrian labor market for customer support specialists was tight and salaries had risen 12% year-over-year.

    Challenge: Scaling Support Without New Hires

    The operations director identified three specific pain points. First, 45% of inbound support tickets involved repetitive data entry: policy number lookups, claim status updates, and address changes. Second, agents spent an average of 14 minutes per ticket searching internal documentation for policy details and claim procedures. Third, the company faced a compliance deadline under the EU AI Act, which required transparency and human oversight for customer-facing AI systems. The deadline was 18 months out, but the company wanted to be ahead of the curve. The operations team had 12 full-time support agents, and the director was told by HR that hiring two more would cost EUR 120,000 annually. The goal was to replace manual data entry and reduce documentation search time without adding headcount.

    Approach: Process Audit and Fixed-Scope Pilot

    The engagement began with a two-week process audit. We mapped every step of the top 20 support workflows, measured cycle time and error rate for each, and identified where manual data entry occurred. The audit revealed that 60% of the top 20 workflows involved repetitive data entry that could be automated. We then built a fixed-scope pilot targeting one workflow: first-response triage for policy status inquiries. The pilot used Anthropic Claude API for the voice agent, with a RAG assistant indexing Confluence and Notion documentation. The architecture was model-agnostic, so we could switch to an open-weight model on the client’s hardware if data residency became an issue. The pilot integrated with Salesforce and Zendesk through their APIs, not by replacing them. Human-in-the-loop approval was built in: the voice agent drafted responses and extracted data fields, but a human approved anything that touched money, health data, or a contract.

    Outcome: Measured Baseline and Rollout Decision

    The 8-week pilot delivered measurable results. Average ticket resolution time for policy status inquiries dropped from 14 minutes to 7 minutes, a 50% reduction. Manual data entry errors fell from 8% to 2%, a 75% reduction. The voice agent handled first-response triage for 70% of policy status inquiries, reducing the need for human escalation. The RAG assistant cut documentation search time from 14 minutes to 3 minutes per ticket. The human-in-the-loop approval process added 2 minutes to each ticket, but the net effect was a 5-minute reduction in cycle time. The pilot met the EU AI Act transparency requirements: all interactions were logged, and the voice agent disclosed its AI nature to customers. The operations director approved a rollout to the remaining 19 workflows, with a target of 12 months for full deployment.

    Lessons for Similar Teams

    • Start with the process audit, not the model. The audit revealed that 60% of the top 20 workflows were automatable, but the model choice was secondary. Teams that skip the audit and jump to model selection often automate the wrong workflows.
    • Fixed-scope pilots reduce risk. The 8-week timeline and defined success metrics gave the operations director confidence to approve the rollout. Without the pilot, the rollout would have been a 6-month project with no baseline to measure against.
    • Human-in-the-loop is not optional. The EU AI Act requires human oversight for customer-facing AI systems. Building it in from the start avoids rework and reduces liability risk.
    • Model-agnostic architecture future-proofs the investment. The ability to switch between Anthropic Claude and open-weight models on the client’s hardware means the company can adapt to changes in cost, latency, and compliance requirements without rebuilding the system.
    • Integrate with existing systems, not replace them. The pilot plugged into Salesforce, Zendesk, and Confluence through their APIs. This reduced integration risk and allowed the operations team to continue using the tools they already knew.
  • LLM Document Extraction with n8n: EU AI Act Compliance for B2B SaaS in Austria

    EU AI Act

    The EU AI Act (Regulation (EU) 2024/1689) is the first comprehensive AI regulation in the world, entering into force on 1 August 2024. It classifies AI systems by risk level and imposes obligations on providers and deployers. For a document extraction pipeline that processes order and shipment data, the system is generally not high-risk, but if it touches personal data or feeds automated decisions, it may trigger transparency and logging obligations under Articles 13 and 14. The Act’s Article 4 requires AI literacy for staff operating the system, which Forfis addresses through the pilot’s training module. In this scenario, the compliance checklist maps each pipeline step to the relevant Act articles, ensuring the client can demonstrate conformity during audits.

    Document Extraction

    Document extraction is the process of converting unstructured or semi-structured documents (PDFs, emails, scanned images) into structured data (JSON, CSV, database records). In this scenario, the LLM reads order confirmations and shipment notifications from Gmail, extracts fields like order ID, shipment ID, carrier, and tracking number, and outputs them as JSON. The extraction accuracy depends on the document format and the LLM’s training data; Forfis measures accuracy per field during the pilot and reports it in the baseline. The human-in-the-loop review step catches extraction errors before the data is written to the SaaS platform, reducing the error rate to below 0.5% in Forfis’s measured baselines.

    Human-in-the-loop (HITL)

    Human-in-the-loop (HITL) means a person reviews and approves the AI’s output before it affects downstream systems. In this pipeline, the LLM extracts order and shipment data, but a human operator confirms the extracted fields before the data is written to the B2B SaaS platform. This is mandatory under Forfis’s default delivery model for anything touching financial records or customer commitments. The HITL step adds roughly 30–60 seconds per document but reduces error rates to below 0.5% in Forfis’s measured baselines. The EU AI Act’s Article 14 requires human oversight for high-risk systems, and the HITL review step satisfies this requirement by allowing the operator to reject, correct, or escalate the extracted data.

    n8n Orchestration

    n8n is an open-source workflow automation platform that uses a visual node-based editor to connect APIs, databases, and services. In this scenario, n8n acts as the orchestration layer: it receives a new email from Google Workspace, triggers the LLM extraction node, validates the output against a schema, and pushes the structured data into the B2B SaaS platform’s order management API. n8n’s self-hosted deployment option keeps data within the client’s Austrian infrastructure, satisfying data residency requirements. The platform’s node-based architecture means the pipeline can be modified without code changes, and the model-agnostic design allows swapping between OpenAI, Anthropic, or open-weight models by changing a single configuration parameter.

    LLM Integration

    LLM integration refers to embedding a large language model into an existing system to perform a specific task, such as document extraction or text classification. In this scenario, the LLM is integrated into the n8n pipeline to read order and shipment emails and extract structured data. The integration is model-agnostic: Forfis uses OpenAI’s GPT-4o or Anthropic’s Claude 3.5 Sonnet for cloud-based processing, or an open-weight model like Llama 3 70B on the client’s own GPU server for regulated data. The n8n orchestration layer abstracts the model choice, so switching providers requires only a configuration change, not a code rewrite. The LLM’s output is validated against a JSON schema before being pushed to the SaaS platform.

    Fixed-Scope Pilot

    Fixed-scope pilot is a bounded engagement with a defined deliverable, timeline, and success metric. Here, the pilot runs for two weeks, targets one specific workflow (order and shipment status updates), and ships with a measured before/after baseline on cycle time and error rate. The scope excludes multi-language support, voice interfaces, or integration with systems outside the agreed API list. This structure limits risk for the client and gives Forfis a clear acceptance criterion. The pilot report compares the baseline metrics from the first three days (manual process) with the metrics from the remaining nine days (automated pipeline), quantifying the reduction in cycle time and error rate as the business case for full rollout.

    Process Audit

    Process audit is the first phase of Forfis’s delivery model, typically taking two to three days. Forfis interviews the operations team, observes the current manual workflow, and maps every step from email receipt to data entry completion. The audit identifies which fields are extracted, which systems are involved, where errors occur, and how long each step takes. The output is a process map and a recommendation on which workflow to automate first. In this scenario, the audit confirmed that order and shipment status updates were the highest-volume, most error-prone workflow, making it the ideal pilot candidate. The audit also identifies compliance requirements under the EU AI Act and data residency constraints that shape the architecture.

  • AI Automation Glossary for Austrian Insurance: 12 Terms from Pilot to Scale

    Process Audit

    A process audit is the first step in any AI automation engagement. It maps existing workflows, measures current cycle times and error rates, and identifies which tasks are repetitive, rule-based, and suitable for automation. For a 51-200 person insurance firm in Austria, this typically involves reviewing 10-20 back-office processes across claims, underwriting, and customer support. The audit produces a prioritized list with estimated ROI, complexity, and compliance risk for each candidate workflow. This baseline is critical because it defines the success metrics for the subsequent pilot and ensures the automation targets the highest-impact processes rather than the easiest ones.

    Fixed-Scope Pilot

    A fixed-scope pilot is a bounded engagement where the deliverable, success metrics, and timeline are agreed before work begins. For an Austrian insurer, this typically means automating one specific workflow—like extracting data from claims forms or triaging support tickets—within 3 to 6 weeks. The scope is deliberately narrow: one process, one team, one set of success criteria. The pilot ships with a measured before/after baseline on cycle time and error rate, providing a clear go/no-go decision for full rollout. This approach reduces risk for both the insurer and the vendor, as the cost and effort are capped, and the outcome is objectively measurable rather than subjective.

    Human-in-the-Loop

    Human-in-the-loop (HITL) is a design pattern where AI systems draft or classify information, but a human reviews and approves actions that have financial, legal, or health implications. In insurance, this means the AI can extract data from invoices, triage support tickets, or draft response emails, but a human must approve any claim payment, policy change, or contract modification before it proceeds. HITL is not optional in regulated industries; it is a compliance requirement under ISO 27001 and GDPR. The design ensures that the AI handles the volume and speed, while humans retain accountability for decisions that affect customers or the company’s financial position.

    Retrieval-Augmented Generation

    Retrieval-augmented generation (RAG) is a technique where an AI model retrieves relevant documents from a knowledge base before generating a response. For an insurer, this means the assistant pulls from policy documents, claims history, and internal procedures stored in Confluence or Notion, ensuring answers are grounded in the company’s actual records rather than general training data. RAG is critical for customer support, where accuracy and consistency matter. Without it, the AI might generate plausible but incorrect answers about coverage details or claim status. With RAG, the model cites the specific policy clause or internal procedure it is referencing, making the response auditable and verifiable.

    Voice Agent

    A voice agent is an AI system that handles inbound or outbound phone calls using speech-to-text, natural language processing, and text-to-speech. In insurance, it can answer routine queries about policy status, claim progress, or payment schedules. The agent is integrated with the CRM and claims system, so it can pull real-time data and provide accurate answers. Human-in-the-loop design ensures that if the caller asks about coverage details, disputes, or complex claims, the call transfers to a human agent within 30 seconds. For a 51-200 person insurer, a voice agent can reduce call handling time by 40-60% for routine queries, freeing senior staff to focus on high-value interactions.

    ISO 27001 Compliance

    ISO 27001 is an international standard for information security management systems. For AI projects in insurance, it requires documented risk assessments, access controls, and audit trails. When using external APIs like Anthropic Claude, the insurer must ensure data processing agreements comply with ISO 27001 Annex A controls, particularly A.13 (communications security) and A.14 (system acquisition, development and maintenance). For regulated data that cannot leave the building, the architecture uses open-weight models on the client’s own hardware. This model-agnostic approach allows the insurer to use the best model for each task while maintaining compliance with ISO 27001 and GDPR requirements.

    Document Extraction Pipeline

    Document extraction pipelines use AI to pull structured data from unstructured documents like invoices, claims forms, and policy documents. For an Austrian insurer, this might involve extracting policyholder names, claim amounts, and dates from scanned PDFs, then validating the data against the CRM before entering it into the ERP system. The pipeline includes multiple stages: document ingestion, OCR (if scanned), data extraction, validation, and human review for edge cases. Error rates are typically measured against a human-verified sample of 100-200 documents, with a target of less than 2% error rate for high-volume processes. This reduces manual data entry by 70-80%, freeing back-office staff to focus on exception handling and customer interaction.

  • B2B SaaS in Austria Cuts First-Response Time 94% with RAG Ticket Triage

    Background: A 30-Person B2B SaaS Firm in Vienna

    This case study is a composite based on patterns observed in the field. We do not fake named customers. The company is a 30-person B2B SaaS vendor based in Vienna, selling a project-management tool to mid-market clients across DACH. The stack runs on AWS, with a custom helpdesk built on top of a commercial ticketing platform. Google Workspace handles email, calendar, and document storage. The team is lean: four engineers, two product managers, one operations lead, and a part-time compliance officer. The company holds ISO 27001 certification, which constrains where customer data can be processed and stored. The operations team handles roughly 180 support tickets per week, with a median first-response time of 4 hours and a 12% mis-routing rate. The CEO had set a target: cut first-response time below 30 minutes within a quarter, without adding headcount.

    Challenge: 4-Hour First-Response Time and ISO 27001 Constraints

    The operations team was drowning in repetitive triage work. Every incoming ticket required a human to read it, classify it by product area, assign it to the right engineer, and draft a first response. The 12% mis-routing rate meant tickets bounced between teams, adding 2-3 hours of dead time per mis-routed ticket. The compliance officer flagged that any AI solution had to respect ISO 27001 controls: customer data could not be sent to unvetted third-party processors, and the data-processing agreement had to be in place before any model touched production data. The deadline was tight: the CEO wanted a measurable improvement within two weeks, not a six-month transformation. The team had no in-house ML expertise. They needed a partner who could audit the process, build a working pilot, and hand over a managed operation without requiring the client to hire a data-science team.

    Approach: RAG Assistant on OpenAI API with Human-in-the-Loop

    Forfis started with a process audit that mapped the ticket lifecycle from intake to resolution. The audit identified three high-leverage automation points: ticket classification, routing, and first-response drafting. The pilot scope was fixed: a retrieval-augmented knowledge assistant that ingested the company’s product documentation, past resolved tickets, and Google Workspace emails. The assistant used the OpenAI API for classification and drafting, with a human-in-the-loop approval step for any ticket touching billing, data deletion, or contract terms. The integration plugged into the existing helpdesk and Google Workspace through their APIs, not a replacement. The architecture was model-agnostic: if the compliance officer later required on-premises processing, the stack could swap to an open-weight model without re-architecting the integration layer. The pilot ran for two weeks, with a measured before/after baseline on first-response time, mis-routing rate, and escalation rate.

    Outcome: 94% Faster First Response in Two Weeks

    After two weeks, the pilot showed a 94% reduction in median first-response time, from 4 hours to 22 minutes. The mis-routing rate dropped from 12% to 3%. Agent escalation rate fell by 40%, because the assistant handled routine queries without human intervention. The human-in-the-loop approval step caught 14 tickets that required manual review, all of which were billing or data-deletion requests. The compliance officer confirmed that no customer data left the approved processing boundary. The operations lead reported that the team could now focus on complex escalations instead of triage. The CEO approved full rollout to all product lines. The engagement moved to managed AI operations, with Forfis monitoring model performance, updating the knowledge base, and handling API changes. The client did not hire a data-science team; the managed operation absorbed that responsibility.

    Lessons for Similar Teams

    • Fix the process before the model. The audit identified that 60% of mis-routes came from ambiguous ticket categories, not from model error. Renaming three categories cut mis-routes by half before the model even ran.
    • Human-in-the-loop is not optional for compliance. The approval step for billing and data-deletion tickets was the difference between a compliant pilot and a liability. ISO 27001 auditors accepted the design because the human approval was logged and auditable.
    • Model-agnostic architecture protects you from regulatory shifts. The client could swap from OpenAI to an on-premises open-weight model if a regulator required it, without rewriting the integration layer. This flexibility was a selling point in the compliance review.
    • Two weeks is enough for a pilot if the scope is fixed. The team resisted the urge to expand the pilot to include voice or email drafting. Staying on ticket triage and routing kept the timeline realistic and the metrics clean.
    • Managed operations beat one-off delivery. The client did not have the in-house capacity to maintain the model, update the knowledge base, or handle API deprecations. The managed operation model removed that burden and kept the system running at pilot-level performance.
  • Voice Agent for Order Status in Austrian Fintech: Two-Week Pilot with pgvector

    The Problem: Routine Inquiries Consuming Senior Staff Time

    Your support team handles 300-500 calls per week, 60% of which are routine inquiries about order status or shipment tracking. Senior staff spend 12-15 hours weekly on these repetitive tasks, delaying complex escalations and fraud reviews. The goal is to free senior staff from routine work by deploying a voice agent that handles 24/7 customer response for order and shipment status updates. The agent must integrate with your existing CRM and ERP, comply with the EU AI Act, and operate within a two-week pilot window. The architecture uses pgvector embeddings search to retrieve relevant records from your own database, keeping regulated data on-premises. The pilot ships with a human-in-the-loop approval gate for any action that touches money or modifies a contract.

    Prerequisites: What You Need Before Step 1

    • Access to your CRM or ERP API with read permissions for order and shipment records.
    • A sample of 50-100 historical customer inquiries, anonymized, to train the intent classifier.
    • A designated human approver with authority to approve or reject transactional actions.
    • Slack or Microsoft Teams workspace where your support team already operates.
    • A PostgreSQL database with pgvector extension enabled, or a plan to deploy it.
    • A clear definition of the pilot scope: one workflow (order/shipment status), one channel (voice), two weeks.
    • Compliance sign-off from your legal team on the EU AI Act requirements for financial services AI.

    Steps 1-3: Audit, Embeddings, and Agent Configuration

    1. Audit the workflow. Map the current process for order status inquiries: average call duration, number of escalations, error rate, and the specific data points customers ask for. Document the before/after baseline: cycle time from inquiry to resolution, and the percentage of inquiries that require human intervention. This baseline becomes the success metric for the pilot.

    2. Set up pgvector embeddings. Install the pgvector extension in your PostgreSQL database. Create a table for embeddings with a vector column of dimension 1536 (matching OpenAI’s text-embedding-3-small). Ingest your order and shipment records, generating embeddings for each record. This allows the voice agent to retrieve relevant records via semantic search rather than exact keyword matching.

    3. Configure the voice agent. Use a model-agnostic architecture: OpenAI or Anthropic APIs for quality-critical tasks like intent classification and response generation, and an open-weight model on your own hardware for any task involving regulated data. Configure the agent to query pgvector for order and shipment records, then generate a response. Set the human-in-the-loop gate: any action that modifies a customer’s financial state requires approval from a human in Slack or Microsoft Teams.

    Steps 4-6: Integration, Pilot, and Go-Live

    1. Integrate with Slack or Microsoft Teams. Configure the agent to post notifications to your support team’s channel when a case requires human approval. The notification includes a summary of the customer’s inquiry, the retrieved records, and the proposed action. The human approver reviews the case, clicks approve or reject, and the agent executes the approved response. This keeps the workflow within your existing communication tools, reducing friction.

    2. Run the pilot in parallel. For two weeks, the voice agent handles incoming calls in parallel with your existing support process. Measure the after/after metrics: cycle time, error rate, and the percentage of inquiries resolved without human intervention. Compare these to the baseline from Step 1. Identify any misclassifications or retrieval errors, and feed them back into the embeddings and intent classifier.

    3. Go-live and hand off to managed operations. After two weeks, if the pilot meets the success criteria, transition the voice agent to production. Forfis takes over managed AI operations: monitoring model performance, handling drift, updating embeddings as new records are added, and maintaining the human-in-the-loop workflow. Your staff focuses on reviewing flagged cases and expanding the agent’s scope to new workflows.

    Common Pitfalls and How to Detect Them

    • Over-scoping the pilot. Trying to automate multiple workflows or channels in two weeks leads to a rushed build with insufficient testing. Stick to one workflow and one channel. Detect this by reviewing the pilot scope document: if it lists more than one workflow or channel, cut the scope.
    • Skipping the baseline measurement. Without a clear before/after metric on cycle time and error rate, you cannot prove the pilot’s value to stakeholders. Detect this by checking whether the audit in Step 1 produced a documented baseline with specific numbers.
    • Untrained human approvers. If your approvers are not trained on the approval workflow, the human-in-the-loop gate becomes a bottleneck, negating the time savings. Detect this by measuring the average time from notification to approval during the pilot. If it exceeds 10 minutes, retrain the approvers.
    • Embedding drift. As new order and shipment records are added, the embeddings may become stale, leading to retrieval errors. Detect this by monitoring the retrieval accuracy metric during the pilot. If it drops below 90%, re-ingest the embeddings.

    Conclusion: What Comes After the Pilot

    The pilot proves whether a voice agent can handle routine order and shipment status inquiries in an Austrian fintech within a two-week window. If the success criteria are met, the next logical step is to expand the agent’s scope to additional workflows, such as payment disputes or account changes. This requires a deeper integration with your ERP and a more complex human-in-the-loop approval workflow. The managed operations model ensures that the technical side of this expansion is handled by Forfis, while your staff focuses on the business side: defining the new workflows, training the approvers, and measuring the impact on senior staff time. The architecture remains model-agnostic and data-resident, satisfying the EU AI Act and GDPR requirements throughout the scaling process.

  • Retrieval-Augmented Candidate Screening: A 4-Week Pilot for Austrian Healthcare

    1. Replace Manual Data Entry First

    Most companies that automate candidate screening start by replacing the manual data entry step. Recruiters spend 2-3 hours per week copying data from resumes into their ATS. A retrieval-augmented assistant built on pgvector can extract structured fields (name, experience, certifications) and classify candidates against your job description in under 18 seconds per application. The human-in-the-loop design means a recruiter approves or rejects each classification before it touches the hiring pipeline. This single process automation reduces cycle time by 40-60% and eliminates transcription errors, giving you a measurable baseline before you consider expanding to other workflows.

    2. Build EU AI Act Compliance Into the Pilot

    The EU AI Act, which entered into force in August 2024, classifies AI systems that make decisions affecting individuals as high-risk. Candidate screening tools that process personal data and influence hiring decisions fall squarely into this category. Article 10 requires data governance, Article 13 mandates transparency, and Article 14 demands human oversight. Forfis builds these controls into the pilot from day one: every classification is logged, every decision is auditable, and no candidate is screened out without human review. This is not a compliance checkbox added at the end; it is the architecture of the system.

    3. Use pgvector for Grounded Answers

    pgvector is a PostgreSQL extension that stores vector embeddings and performs similarity search. For a 501-2000 employee company, this means you can run your RAG pipeline on the same database as your transactional data, avoiding the cost and complexity of a dedicated vector database. The assistant embeds your job descriptions, screening criteria, and past hiring decisions into pgvector. When a new application arrives, the system retrieves the most relevant chunks and feeds them to an LLM, which generates a classification grounded in your data. This reduces hallucinations and keeps answers current as your criteria change.

    4. Integrate With Your Existing ATS via REST APIs

    The assistant connects to your ATS, HRIS, or recruitment platform via their REST APIs. Webhooks trigger the screening workflow when a new application arrives. The system extracts structured data from resumes, classifies candidates, and writes results back to your existing system. No replacement of your current tools is required. The architecture is deliberately model-agnostic: OpenAI or Anthropic APIs where quality matters, open-weight models on your own hardware where regulated data cannot leave the building. This means you can switch models without rebuilding the pipeline, and you can keep candidate data within your infrastructure if required.

    5. Ship a Measurable Result in 4 Weeks

    A 4-week timeline is realistic for a single-process pilot. Week 1: process audit and baseline measurement. Week 2: build the RAG pipeline and API integration. Week 3: test with real data and tune the model. Week 4: measure results, document findings, and hand over. This assumes your APIs are accessible and your data is in a usable format. The pilot ships with a report showing whether the automation meets the agreed thresholds on cycle time and error rate before you commit to rollout. This fixed-scope approach protects you from scope creep and ensures you have a measurable result before expanding to other workflows.

    6. Keep Humans in the Loop for High-Risk Decisions

    The assistant drafts a shortlist of candidates based on your job description and screening criteria. A recruiter reviews each draft, approves or rejects the classification, and the system logs the decision. This human-in-the-loop design ensures no candidate is screened out without human review, satisfying EU AI Act requirements for high-risk AI systems. The model classifies, the person decides. This is not a limitation; it is the correct architecture for a regulated environment. Every pilot ships with a measured before/after baseline on cycle time and error rate, so you know exactly what the automation achieved and where human judgment still adds value.

  • Cutting Contract Review Errors by 60% in a Two-Week B2B SaaS Pilot

    1. Baseline Error Rate Is the Real KPI

    The finance team at a 2,000+ employee B2B SaaS company in Vienna processes roughly 1,200 contracts per month. Each one passes through a manual review queue where an analyst extracts termination clauses, liability caps, and auto-renewal flags into the ERP. The baseline error rate sits at 5.2%: a missed auto-renewal date or a misread liability cap ends up in the system and surfaces three months later during a renewal dispute. A two-week pilot with a dedicated AI team replaced the manual extraction step with a LangGraph pipeline that parses PDFs, extracts 14 structured fields, and writes the result to a staging table via a custom REST API. The measured error rate dropped to 1.8% on the pilot’s 300-contract sample, and cycle time per contract fell from 11 minutes to 90 seconds of model time plus 4 minutes of human approval. The pilot did not touch the production ERP; it ran on a read-only copy of the contract repository and output to a sandbox workspace in the CRM.

    2. LangGraph Handles the Multi-Step Extraction

    The extraction pipeline runs on LangGraph, not a single LLM call. The graph has five nodes: PDF ingestion (PyMuPDF for text-layer PDFs, Tesseract OCR fallback for scanned documents), clause segmentation (a fine-tuned classifier that splits the document into 8–12 logical sections), field extraction (GPT-4o for high-accuracy fields like liability caps, Llama 3 70B on the client’s own GPU for fields containing personal data), confidence scoring, and output formatting. The REST API endpoint POST /v1/extract accepts a multipart PDF upload and returns a JSON object with 14 fields, each carrying a confidence score between 0 and 1. Fields below 0.90 route to a human reviewer in the existing helpdesk queue; fields at or above 0.90 auto-populate the staging table. Webhooks fire on completion so the finance team’s dashboard updates without polling. The entire pipeline runs on the client’s AWS eu-central-1 region, keeping data within Austria’s borders.

    3. Two Weeks Is Enough for a Measured Pilot

    The pilot ran for exactly 14 calendar days. Days 1–3: process audit. The AI team shadowed three finance analysts, logged every manual step, and identified the 14 fields that caused the most downstream errors. Days 4–6: data preparation. The team pulled 300 historical contracts from the repository, had two analysts independently annotate the 14 fields, and resolved disagreements to build a gold-standard test set. Days 7–10: pipeline build and tuning. The LangGraph workflow was assembled, the extraction prompt was iterated four times, and the confidence threshold was calibrated so that the false-negative rate (a wrong value auto-approved) stayed below 0.5%. Days 11–14: measurement. The pipeline ran on the 300-contract set, and the team compared field-level accuracy against the gold set, measured cycle time, and produced a before/after report. The report included a cost model: at 1,200 contracts per month, the pilot’s error reduction translated to an estimated EUR 18,400 in avoided dispute costs per quarter.

    4. Human-in-the-Loop Is Non-Negotiable

    The model does not replace the analyst; it removes the 11 minutes of copy-paste and field-mapping that precede the actual judgment call. The human-in-the-loop design is explicit: the model drafts the 14 extracted fields, the analyst reviews them in a purpose-built UI that highlights low-confidence fields in amber, and the analyst approves or corrects before the record writes to the ERP. For a B2B SaaS company, the highest-risk fields are termination notice periods and liability caps, because a wrong value here has direct financial consequences. The pilot’s measurement showed that 78% of fields required no human correction, 19% needed a single-field edit, and 3% required a full re-extraction. The analyst’s role shifted from data entry to exception handling, which freed roughly 6.5 hours per analyst per week. The dedicated AI team operated the pipeline during the pilot, monitored confidence drift, and tuned the prompt when a new contract template appeared in the sample.

    5. The Integration Is a Thin REST Layer

    The pilot’s REST API and webhook architecture was designed to plug into the client’s existing stack without replacing it. The extraction service exposes a stateless POST /v1/extract endpoint that the finance team’s internal tool calls via a simple HTTP request. On completion, a webhook POSTs the result to the client’s CRM (Salesforce) and ERP (SAP S/4HANA) through their respective API endpoints. No middleware, no new database, no replacement of the existing document management system. The client’s IT team reviewed the API contract in day 2 of the pilot and approved the integration scope. The model-agnostic design meant the team could swap GPT-4o for Llama 3 on the client’s GPU for any field that contained personal data, without changing the API contract or the downstream integration. This matters for a 2,000+ employee firm where IT governance requires that no new SaaS dependency is introduced for a pilot that may not scale.

    6. What the Pilot Does Not Cover

    The pilot’s 1.8% error rate is not the end state. The team’s rollout plan, presented in the final pilot report, targets a 0.9% error rate within 90 days of production deployment. The path: expand the gold-standard test set from 300 to 2,000 contracts, add a second extraction pass for fields with confidence between 0.80 and 0.90, and introduce a feedback loop where analyst corrections are logged and used to fine-tune the clause-segmentation classifier. The dedicated AI team continues to operate the pipeline in production, monitoring a dashboard that tracks field-level accuracy, confidence distribution, and cycle time per contract. The B2B SaaS firm’s finance director approved the rollout on the basis of the pilot’s measured numbers, not a projection. The two-week window was sufficient because the scope was narrow: one document type, 14 fields, one team, one measurement. Expanding to multi-party agreements or adding a second document type (e.g., purchase orders) would require a second pilot of similar duration.