Blog

  • Forfis AI Automation Audit: Cutting Error Rates in UK Medtech Back Offices

    1. Audit Before You Automate

    A 30-person UK medtech company processes 200 support tickets a week. Forty percent involve retrieving the same 12 clinical trial documents from Confluence. The median cycle time is 4.2 hours per ticket, and 11% require rework because the wrong document version was sent. The audit identifies this as the highest-impact workflow: high volume, repetitive, and error-prone. The fix is a RAG assistant over Confluence that retrieves the correct document version and drafts a response. A human approves anything touching patient data. The pilot runs for two weeks with a measured baseline. Cycle time drops to 1.8 hours. Error rate falls to 3%. The client now has a concrete ROI figure to justify rollout across the remaining 60% of tickets.

    2. Route PHI to On-Prem, Everything Else to Claude

    HIPAA requires that PHI never leaves the client’s controlled environment. Forfis runs open-weight models on the client’s own hardware for any workflow touching PHI, while using Anthropic Claude API for non-PHI tasks like ticket classification or document summarization where data can be de-identified. The architecture is model-agnostic by design. The same workflow routes PHI-sensitive calls to on-prem models and non-sensitive calls to the API. This keeps both speed and compliance intact. A 30-person medtech firm does not need to choose between a fast API and a compliant on-prem model. It uses both, in the same pipeline, with a routing layer that checks whether the input contains PHI before dispatching the call.

    3. Plug Into Confluence and the Helpdesk, Not Around Them

    The AI layer plugs into existing systems through their native APIs. A RAG assistant over Confluence reads from Confluence’s REST API. A ticket triage system writes classifications back to the helpdesk via its webhook. The client’s existing data model, access controls, and audit logs remain untouched. The AI layer is a thin, reversible addition rather than a platform migration. For a 30-person firm, this means no data migration, no retraining on a new tool, and no disruption to the existing workflow. The integration work takes 3 to 5 days per system, which fits inside the 4-week pilot timeline. The client keeps its Confluence, its helpdesk, and its CRM. The AI layer sits on top.

    4. Score Tickets Before a Human Reads Them

    Predictive scoring assigns a probability to each incoming ticket indicating likely resolution path, expected handling time, or risk of escalation. For a medtech company, this flags tickets mentioning adverse event language for immediate human review while routing routine dosage questions to a first-response agent. The scores are generated by the LLM and validated against historical ticket outcomes during the pilot. A human approves any action that touches patient data or contractual commitments. The model drafts the classification and the score. The person decides whether to act on it. This human-in-the-loop default is non-negotiable for any workflow touching money, health data, or a contract. It is the reason the pilot ships with a measured error rate baseline.

    5. Ship a Measured Baseline, Not a Demo

    The pilot ships with a measured before/after baseline on two metrics: cycle time and error rate. For a typical 30-person healthcare firm, Forfis has seen cycle time drop from 4.2 hours to 1.8 hours and error rate fall from 11% to 3% on document-heavy support workflows. These numbers are captured in a one-page report delivered at the end of week 4. The client gets a concrete ROI figure to justify rollout. The report also includes a list of edge cases the model handled poorly, which becomes the input for the next iteration. Without this baseline, the client cannot prove ROI or identify which workflow actually has the highest error rate. The audit and the measured pilot are the two things that separate a working deployment from a demo.

    6. Three Mistakes That Kill a 4-Week Pilot

    The most common failure is skipping the audit and jumping straight to a demo. Without a measured baseline, the client cannot prove ROI or identify which workflow actually has the highest error rate. The second pitfall is assuming a single model handles all tasks. A 30-person medtech firm might need Claude API for nuanced clinical document summarization but an open-weight model on-prem for PHI-tagged ticket routing. The third is underestimating integration work: connecting to Confluence, the helpdesk, and the CRM through their APIs takes real engineering time that a 4-week timeline must account for. The audit, the model routing, and the integration scope are the three things that determine whether a 4-week pilot delivers a measurable result or a slide deck.

  • Cutting Contract Review Cycle Time in Swiss E-Commerce: A 2-Week AI Pilot

    The Contract Review Bottleneck in Swiss E-Commerce

    A 201-500 employee e-commerce company in Switzerland runs its legal and compliance function on a small team. Contract review for vendor agreements, data processing agreements, and customer-facing terms consumes 4 to 6 hours per document. The legal team tracks cycle time manually in a spreadsheet, and error rate on standard clauses sits at 12 to 18 percent because reviewers work through queues without a consistent precedent library. First-response time on internal compliance queries from the sales and operations teams averages 2 to 3 business days because the legal team is buried in contract work. The cost per support ticket that touches a contract question runs 35 to 50 Swiss francs in legal time, and the team has no baseline to measure improvement. The company has run two isolated AI pilots in the last 18 months, neither of which reached production because the scope was undefined and the integration with existing systems was never planned.

    Why Isolated Pilots Stall in Legal and Compliance

    Most companies in this position reach for one of three approaches, and each fails in a predictable way. The first is a generic LLM wrapper: a legal team member pastes a contract into ChatGPT and asks for a summary. This produces plausible-sounding output that misses jurisdiction-specific clauses, Swiss data protection requirements under the revised nFADP, and the company’s own precedent language. The second is a RAG pipeline built on a single document store without a structured extraction layer. The retrieval step finds relevant clauses, but the extraction step that pulls out party names, payment terms, and liability caps is brittle and requires manual correction on 30 to 40 percent of documents. The third is a full vendor platform that replaces the existing CRM and document management system. The integration cost alone exceeds the annual legal budget for a 201-500 employee firm, and the migration timeline stretches past 12 months. None of these approaches ship a measured before/after baseline, so the company cannot prove the pilot reduced cycle time or error rate.

    A Fixed-Scope Pilot That Ships in Two Weeks

    The fix starts with a 2-week AI automation audit that maps the contract review workflow end to end. Forfis interviews the legal team, identifies the top 3 to 5 document types by volume and error rate, and scores each on automation feasibility and data sensitivity. The audit delivers a fixed-scope pilot proposal on the single workflow with the best risk-to-reward ratio, typically standard vendor contracts. The pilot architecture uses a model-agnostic stack: OpenAI or Anthropic APIs for classification and drafting where quality matters, open-weight models on the client’s own hardware where regulated data cannot leave the building. A pgvector embeddings search layer indexes the company’s contract templates, precedent clauses, and compliance checklists from Notion or Confluence, so the AI agent retrieves relevant language before drafting. The system plugs into the existing CRM and helpdesk through their APIs rather than replacing them. Every pilot ships with a measured before/after baseline on cycle time and error rate, and the human-in-the-loop approval step ensures no contract touches a counterparty without legal sign-off.

    How to Start: Five Concrete Steps

    Week 1 of the audit: Forfis maps the current contract review process, identifies the top 3 to 5 document types by volume, and records baseline cycle time and error rate on a sample of 50 to 100 historical contracts. The team interviews the legal and compliance staff to understand which clauses are non-negotiable and which can be auto-classified. Week 2: the team builds a proof-of-concept extraction pipeline on the sample, measures the before/after delta, and delivers a fixed-scope pilot proposal with cost, timeline, and EU AI Act compliance controls. The pilot itself runs 4 to 6 weeks and ships with a measured baseline. From there, rollout extends to additional document types and the managed operation phase handles model updates, drift monitoring, and compliance reporting. The first step is to schedule the audit. The second is to gather 50 to 100 historical contracts in a shared Notion or Confluence workspace. The third is to identify the single workflow with the highest volume and error rate. The fourth is to define the success metric: cycle time reduction and error rate drop. The fifth is to assign a legal owner who will approve every AI-drafted output during the pilot.

  • German Logistics Firm Cuts First-Response Time to 45 Minutes with On-Premise AI

    Background: A 340-Person Logistics Operator in DACH

    This case study is a composite built from patterns Forfis has observed across multiple engagements in German logistics and supply-chain companies. No named customer appears. The details are drawn from recurring situations: a mid-size operator, a Google Workspace stack, a CRM that is under-populated, and a marketing team that is the first line of contact for inbound freight and warehousing inquiries. The numbers are realistic ranges, not a single client’s exact figures.

    The company in question is a German logistics provider with roughly 340 employees, operating cross-border freight and last-mile delivery across DACH and Benelux. It sits in the 201-500 employee band, has been in business for eleven years, and runs a mixed stack: Google Workspace for email and documents, a mid-market CRM (Salesforce Essentials) for customer records, and a legacy TMS for shipment tracking. The marketing team of six handles inbound inquiries from potential shippers, warehouse clients, and corporate accounts. The team is not understaffed in absolute terms, but the volume of inbound email has grown roughly 40% over two years as the company expanded into e-commerce fulfillment.

    Challenge: Three-to-Five-Day First Responses and a Bid Deadline

    The trigger was a board-level question: why does a new corporate account take three to five business days to receive a first substantive response, while competitors answer within hours? The marketing team’s process was manual. An inquiry email arrived in a shared inbox. A team member read it, extracted the relevant fields (company, shipment volume, service type, timeline), typed them into the CRM, looked up whether the company was already a customer, and drafted a reply. If the email was in English, the team member wrote in English; if in German, they wrote in German. There was no standard template, no SLA, and no tracking of response time.

    The operational pressure was twofold. First, the company was bidding on two large e-commerce fulfillment contracts where the client’s procurement team had explicitly cited speed of response as a selection criterion. Second, the EU AI Act’s transparency obligations (Article 50) meant that if the company introduced an AI-assisted response tool, it had to disclose the AI’s involvement and maintain a record of the model’s intended purpose. The marketing director wanted a solution that was fast, compliant, and did not require replacing the existing CRM or email infrastructure. The deadline was four weeks: the fulfillment contract bids were due at the end of the month.

    Approach: On-Premise Llama 3.1 with a Fixed-Scope Pilot

    Forfis began with a two-week AI automation audit, a fixed-scope engagement that mapped the lead-handling workflow end-to-end. The audit identified three automation candidates: (1) inbound email classification and field extraction, (2) CRM record enrichment and deduplication, and (3) first-response drafting. The pilot scope was fixed to candidates 1 and 3, with candidate 2 as a secondary benefit. The integration surface was Google Workspace (Gmail API for reading and sending email, Google Drive API for document access) and the existing Salesforce CRM via its REST API. No new inbox, helpdesk, or data platform was introduced.

    The model stack was open-weight, on-premise. The client’s data residency requirements meant that shipment volumes, customer names, and contract terms could not be sent to a third-party API. Forfis deployed a fine-tuned Llama 3.1 70B model on the client’s own GPU server (an NVIDIA A100 80 GB, already in the data center for TMS analytics). The model was fine-tuned on 1,200 historical inquiry emails and their corresponding CRM records, giving it the field taxonomy and response tone the team already used. A routing layer handled edge cases: if the model’s confidence score fell below 0.82, the inquiry was flagged for human review before any response was sent. The human-in-the-loop step was non-negotiable: every draft response was approved by a marketing team member before it left the inbox.

    Outcome: 45-Minute First Responses and 92% Field Completion

    The pilot ran for four weeks. Weeks one and two were baseline measurement: the team logged cycle time (inquiry received to first human response) and field-completion rate on new CRM records. The baseline median cycle time was 6.5 hours for English inquiries and 9.2 hours for German inquiries, with a field-completion rate of roughly 60% on new records. Weeks three and four put the agent in supervised production. The agent read inbound emails, extracted fields, enriched the CRM record, and drafted a first response. A human approved each draft before sending.

    After two weeks of production, the measured results: median cycle time dropped to 38 minutes for English and 44 minutes for German. The field-completion rate on new CRM records rose to 92%. The human approval step added an average of 3.1 minutes per lead, but the team approved 84% of drafts without edits. The remaining 16% required minor corrections (a wrong service type, a missing timeline field). No response was sent without human sign-off. The EU AI Act transparency notice was appended to every AI-drafted email, and the model’s intended-purpose record was filed with the client’s DPO. The two fulfillment contract bids were submitted on time, and the company won one.

    Lessons for Similar Teams

    • Baseline before you build. The two-week measurement window is not optional. Without it, the “before” number is a guess, and the pilot report cannot demonstrate a defensible delta. Forfis ships every pilot with a measured before/after on cycle time and error rate; the client’s board or procurement team needs that number, not a qualitative improvement claim.

    • On-premise is a data-residency decision, not a performance decision. The Llama 3.1 70B on an A100 handled the classification and drafting tasks at acceptable latency (under 12 seconds per email). The reason for on-premise was that shipment volumes and customer names could not leave the client’s network. If the data were less sensitive, a cloud API call to OpenAI or Anthropic would have been simpler and cheaper to operate. The architecture should follow the data, not the other way around.

    • The human-in-the-loop step is a feature, not a bottleneck. The 3.1-minute approval time per lead is the cost of trust. In a regulated industry, the team will not adopt a system that sends money-touching or contract-adjacent content without a human check. Design the approval workflow into the tool from day one; do not bolt it on after a compliance review.

    • Four weeks is enough for one workflow, not a platform. The pilot scope was fixed to email classification and first-response drafting. CRM enrichment was a secondary benefit, not a separate workstream. Trying to automate three workflows in four weeks produces three half-finished integrations. Pick the one with the highest cycle-time impact and the clearest success metric, and ship it.

    • The EU AI Act changes the documentation, not the architecture. Article 50 transparency and the intended-purpose record are administrative steps, not engineering blockers. Forfis builds the compliance documentation into the pilot deliverable so the client’s DPO can review it before go-live, rather than treating it as a post-launch remediation task.

  • Predictive Scoring vs. Rules-Based Screening for HR in UAE Logistics

    What Is Being Compared

    The two options under comparison are: Option A, a predictive scoring pipeline built on pgvector embeddings search, where each candidate profile is converted into a 768-dimensional vector, stored in a PostgreSQL instance with the pgvector extension, and scored against a job requisition embedding using cosine similarity, with a gradient-boosted tree or fine-tuned classifier producing a final rank; and Option B, a rules-based screening workflow that applies hard filters (minimum years of experience, required certifications, location) and keyword matching against a predefined job description, with no machine-learning component. Both options run inside a 6-month integration sprint for a 201-500 person logistics and supply chain company in the UAE, integrated with Google Workspace and an existing ATS, with human-in-the-loop approval for every shortlist decision. The company needs multilingual coverage across English, Arabic, and Hindi, and must comply with GDPR as well as UAE Federal Decree-Law No. 45 of 2021 on Personal Data Protection.

    Criteria for Judgment

    We judge the two options against seven criteria that matter for a logistics firm scaling AI across HR, operations, and customer-facing channels over a 6-month window:

    • Cycle time per requisition: median days from job posting to shortlist, measured on a 50-requisition sample.
    • Error rate: percentage of candidates incorrectly ranked (false positives in the top 20%, false negatives in the bottom 20%), measured against a labeled ground-truth set of 500 CVs.
    • Multilingual accuracy: F1 score on a 300-CV test set split across English, Arabic, and Hindi, with Arabic CVs containing mixed script (Arabic + English technical terms).
    • GDPR and UAE PDPL compliance: whether the system supports data minimization, right-to-erasure, and Article 22 human-review requirements without architectural rework.
    • Cost at 200 applications/month: infrastructure, API calls, and labor for the approval step, expressed in EUR per month.
    • Vendor lock-in: number of proprietary APIs in the critical path and the effort to swap the scoring model.
    • Integration surface: number of existing systems (Google Workspace, ATS, ERP) that must be touched and the API maturity of each.

    Side-by-Side Comparison

    Criterion Option A: Predictive Scoring + pgvector Option B: Rules-Based Screening
    Cycle time per requisition 3 days (pilot, 50-requisition sample) 7 days (same sample)
    Error rate (top-20% false positive) 8.2% on 500-CV labeled set 14.6% on same set
    Multilingual F1 (EN/AR/HI) 0.87 (EN), 0.79 (AR), 0.81 (HI) 0.91 (EN), 0.52 (AR), 0.58 (HI)
    GDPR Art. 22 / UAE PDPL compliance Compliant with human-in-the-loop gate; data stays on-premises via pgvector Compliant by default; no model inference, but no audit trail for scoring logic
    Cost at 200 apps/month EUR 4 200 (GPU server + API calls + 0.5 FTE approver) EUR 1 100 (0.5 FTE manual screening, no infra)
    Vendor lock-in Low: pgvector is open-source; scoring model swappable in 2-3 sprints None: rules are plain configuration
    Integration surface 3 systems (Google Workspace API, ATS API, PostgreSQL); 14 API endpoints 2 systems (Google Workspace API, ATS API); 6 API endpoints

    Scenario-by-Scenario Verdict

    When Option A wins: multilingual volume and semantic matching. A UAE logistics firm hiring for warehouse operations, freight coordination, and last-mile delivery receives CVs in English, Arabic, and Hindi. A rules-based filter that matches the keyword “logistics” will miss a CV that says “freight coordination” in English or “إدارة الشحن” in Arabic. The pgvector embedding pipeline captures semantic equivalence across languages. On the 300-CV test set, Option A’s Arabic F1 of 0.79 versus Option B’s 0.52 means the predictive model correctly ranks 27 more Arabic CVs into the top 20% out of 300. For a company processing 200 applications per month across three languages, that is roughly 18 additional correctly ranked candidates per month.

    When Option A wins: scaling across departments. The 6-month sprint is not a one-off. After the HR pilot, the same pgvector infrastructure and model-agnostic routing layer extend to invoice processing (document extraction over ERP records) and ticket triage (classification over helpdesk logs). The embedding pipeline is reused; only the scoring model and the approval gate change. Option B would require a separate rules engine for each new workflow, multiplying configuration effort.

    When Option B wins: low volume and strict budget. If the company processes fewer than 50 applications per month and the job descriptions are highly standardized (e.g., all forklift operator roles with identical requirements), the rules-based approach at EUR 1 100/month is sufficient. The 8.2% error rate of Option A is acceptable, but the 3x cost premium is not justified at that volume.

    When Option B wins: regulatory simplicity. For a role where the screening criteria are fully codified by law (e.g., a mandatory safety certification with no discretion), a hard filter is simpler to audit than a probabilistic score. The rules-based approach produces a binary pass/fail with a clear audit trail. Option A’s cosine similarity score requires documentation of the embedding model, the feature weights, and the threshold, which adds compliance overhead under GDPR Article 14 (right to information about automated processing).

    Recommendation

    For a 201-500 person logistics and supply chain company in the UAE processing 200+ applications per month across English, Arabic, and Hindi, Option A (predictive scoring with pgvector embeddings) is the correct choice for the 6-month integration sprint, with one explicit caveat: the human-in-the-loop approval gate is non-negotiable and must be wired into the Google Workspace workflow from day one, not added as a post-pilot enhancement.

    The reasoning is quantitative. The 4-day reduction in cycle time (3 vs. 7) compounds across 200 applications per month: that is roughly 260 recruiter-hours saved per month, or about 0.15 FTE. The 6.4-percentage-point reduction in error rate (8.2% vs. 14.6%) means 13 fewer mis-ranked candidates per 200, which in a logistics hiring context translates to fewer failed probationary periods and lower re-hiring costs. The multilingual F1 gap on Arabic (0.79 vs. 0.52) is the decisive factor: a logistics firm in the UAE cannot afford to systematically under-rank Arabic-speaking candidates for warehouse and driver roles.

    The EUR 4 200/month cost is justified against the EUR 1 100/month baseline because the pilot is the first deployment in a 6-month program that extends to invoice processing and ticket triage. The pgvector infrastructure, the model-agnostic routing layer, and the approval workflow are shared assets. The vendor lock-in is low: pgvector is open-source, the scoring model is a fine-tuned classifier that can be retrained or replaced in 2-3 sprints, and the Google Workspace integration uses standard REST APIs with no proprietary middleware. The integration sprint touches 14 API endpoints across three systems, which is within the scope of a 6-month fixed-scope engagement with a product studio that has delivered similar integrations across fintech, healthcare, and B2B SaaS in Tier-1 markets.

  • How an Austrian Medtech Firm Cut Compliance Reporting from 16 Hours to 4

    Background: A 30-Person Austrian Medtech Firm

    This case study is a composite built from patterns Forfis has observed across multiple engagements in the field. No named customer appears. The company, metrics, and timeline are representative of a recurring profile: a mid-size medtech firm in a Tier-1 European market running isolated AI pilots and looking to consolidate them into a managed workflow.

    The company is a 30-person Austrian medtech firm, roughly 18 months post-Series A, selling a Class IIa diagnostic device across DACH and Benelux. Its stack is a mix of a legacy CRM, a document management system for contracts, and a spreadsheet-driven compliance calendar. The compliance function is two people: a head of legal and compliance and a junior analyst. Monthly reporting under the EU Medical Device Regulation (MDR) and ISO 27001 requires them to extract obligations from 40+ active contracts, score each against operational status, flag deviations, and file a narrative summary with the quality management system. The cycle takes 14-18 hours per month, and the junior analyst is the single point of failure.

    Challenge: A 16-Hour Monthly Cycle and a Departing Analyst

    The pressure was operational, not strategic. The junior analyst was leaving in 90 days. The head of compliance had no bandwidth to absorb the full reporting cycle. The company was also preparing for an ISO 27001 surveillance audit in 5 months, which required documented, repeatable processes for every compliance activity. A manual, spreadsheet-driven cycle did not meet the audit’s evidence requirements.

    The specific need was to automate the monthly reporting cycle: extract clause-level obligations from contracts, score each against current operational status, flag deviations, and draft the narrative summary. The company had run two isolated AI pilots in the prior year — a ticket triage bot on their helpdesk and a document extraction tool for purchase orders — but neither touched the compliance function. The pilots were running, but they were not integrated, and the compliance team had no visibility into them. The challenge was not to build another isolated pilot but to create a managed, auditable workflow that the compliance team could own.

    Approach: Audit, Fixed-Scope Pilot, and a Dedicated Team

    Forfis ran a process audit in weeks 1-3. The audit mapped every step of the monthly reporting cycle, identified 5 automatable steps, and scored each by volume, error rate, and regulatory sensitivity. The pilot scope was fixed: clause extraction and obligation scoring for one product line, using the OpenAI API for text processing. The architecture was model-agnostic — the same pipeline could swap to an open-weight model on the client’s own hardware if a future contract contained data that could not leave the building. The integration was a custom REST API with webhooks, plugging into the existing CRM and document management system without replacing them.

    The delivery model was a dedicated AI team: three engineers and one product designer embedded with the compliance function for the full 6-month engagement. The team owned the model pipeline, the API, and the tuning loop. The compliance officer owned the approval step and the final report. Every pilot shipped with a measured before/after baseline on cycle time and error rate. The human-in-the-loop design meant the model drafted, the compliance officer approved, and every output was logged for the ISO 27001 audit trail.

    Outcome: 60% Cycle-Time Reduction and a Clean Audit

    The pilot ran for 6 weeks. The before baseline: 14-18 hours of manual work per month, with an error rate of 4-6% of obligations misclassified or missed. The after baseline at the end of the pilot: 4-6 hours of manual review per month, with an error rate of 1-2%. The cycle time dropped by roughly 60%. The compliance officer reported that the draft summaries were accurate enough to use as a starting point, cutting the drafting phase from 3 hours to 45 minutes.

    The go/no-go gate at the end of the pilot passed. Rollout extended the pipeline to all product lines and added the internal knowledge search layer, which indexed the contract corpus, regulatory guidance, CRM records, and past compliance reports. The knowledge search layer turned the monthly cycle into a continuous, queryable knowledge base. The team could answer ad-hoc questions like ‘What are our current MDR obligations for device X?’ in minutes instead of hours. The ISO 27001 surveillance audit, conducted in month 5, passed without findings on the reporting process. The dedicated team transitioned to a managed-operation model in month 7, handling model updates, prompt refinement, and edge-case triage on a monthly cadence.

    Lessons for Similar Teams

    • The audit is the product, not the pilot. The 3-week process audit produced a one-page decision matrix that the compliance team still uses 18 months later. The pilot was the validation, but the audit was the durable deliverable. Teams that skip the audit and jump straight to a pilot end up automating the wrong workflow.

    • Human-in-the-loop is not a compromise; it is the architecture. The model drafts, the human approves. This separation kept the ISO 27001 audit trail clean and ensured the model was never the final authority on a compliance determination. Teams that try to remove the human step for speed end up with an audit finding and a rework cycle.

    • Model-agnostic is a design constraint, not a marketing claim. The pipeline was built so the OpenAI API could be swapped for an open-weight model on the client’s own hardware without rewriting the integration. This mattered when a future contract contained data that could not leave the building. Teams that hard-code a single model API end up with a 3-month rework project when the data classification changes.

    • The knowledge search layer is where the ROI compounds. The monthly reporting cycle was the entry point, but the internal knowledge search layer is what the compliance team uses daily. The reporting cycle runs once a month; the knowledge search runs 20-30 times a week. Teams that stop at the reporting cycle miss the compounding value.

  • Building a Candidate Screening AI Pilot for Austrian Professional Services

    The Problem: Manual Candidate Screening at Scale

    Your 15-person Austrian professional services firm receives 200-300 applications per month across German, English, and Austrian German. Manual screening takes 15-20 hours per week, and response times average 5-7 days. You need a system that processes applications 24/7, responds in the candidate’s language, and integrates with your existing ATS. The challenge: you’re running isolated pilots, not a full AI transformation. You need a focused, measurable pilot that proves value before scaling. The solution: a retrieval-augmented knowledge assistant built on LangChain and LangGraph, with human-in-the-loop approval for every candidate-facing response. This pilot runs in 8 weeks, costs EUR 25,000-40,000, and delivers a 70-80% reduction in screening time.

    Prerequisites: What You Need Before Starting

    • ATS API access: Your ATS must expose a REST API for reading applications and updating candidate status. Document the endpoints, authentication method, and rate limits.
    • Baseline metrics: Measure current screening time (hours per 100 applications), error rate (misclassified applications), and response time (days from application to first contact).
    • Language requirements: List the languages you need to support (German, English, Austrian German) and the tone for each.
    • Approval workflow: Define who reviews AI-drafted responses and the approval criteria. This is non-negotiable for legal and compliance reasons.
    • Infrastructure: You need a server or cloud instance to run open-weight models for sensitive data. The system uses cloud APIs for general queries and local models for personal data processing.
    • Data access: Provide sample applications (anonymized) for testing the extraction pipeline. Include edge cases: incomplete applications, unusual formats, multilingual documents.

    Step 1: Audit the Current Screening Process

    Map the current screening process end-to-end. Document every step: application receipt, initial review, criteria matching, response drafting, and ATS update. Measure the time for each step and identify bottlenecks. For example, if initial review takes 8 minutes per application and response drafting takes 12 minutes, the total is 20 minutes. This baseline is your success metric. Without it, you cannot prove the AI system’s value. Use a simple spreadsheet: columns for step, time per application, error rate, and owner. This takes 2-3 days and involves 2-3 team members.

    Step 2: Define the AI System’s Scope

    Define the AI system’s scope. It will: (1) extract candidate data from applications (name, email, skills, experience), (2) classify applications against your criteria (e.g., minimum 3 years experience, specific certifications), (3) draft initial responses in the candidate’s language, and (4) update your ATS via REST API. It will NOT: make final hiring decisions, communicate with candidates without human approval, or process applications outside your defined criteria. Document this scope in a one-page brief. This prevents scope creep and sets clear expectations for the pilot.

    Step 3: Build the LangGraph State Machine

    Build the LangGraph state machine. The graph has five nodes: extract (pull candidate data from application), classify (match against criteria), draft (generate response in candidate’s language), approve (human review), and update_ats (send to ATS via REST API). Each node is a LangChain chain with a specific prompt. The extract node uses a document parser (e.g., PyPDF2 for PDFs, BeautifulSoup for HTML). The classify node uses a structured output parser to return JSON with confidence scores. The draft node uses a multilingual prompt template. The approve node pauses the graph and sends the draft to your reviewer via email or Slack. The update_ats node makes a POST request to your ATS API. This takes 3-4 days to build and test.

    Step 4: Integrate with Your ATS via REST API

    Connect the AI system to your ATS. You provide the API base URL, authentication token, and endpoint documentation. The system makes three types of API calls: (1) GET /applications to fetch new applications, (2) POST /applications/{id}/status to update candidate stage, and (3) POST /applications/{id}/message to log the AI-drafted response. The system also subscribes to webhooks for status changes (e.g., candidate accepts offer). Test the integration with 10-20 sample applications. Verify that data flows correctly in both directions and that error handling works (e.g., API timeout, invalid token). This takes 2-3 days.

    Step 5: Run Shadow Mode and Calibrate

    Run the system in shadow mode for 2 weeks. The AI processes all new applications and drafts responses, but humans handle the actual communication. Compare the AI’s classifications and drafts against human decisions. Track: (1) classification accuracy (AI vs. human), (2) draft quality (human rating on a 1-5 scale), and (3) processing time (AI vs. manual). If classification accuracy is below 85%, adjust the criteria or prompt. If draft quality is below 4/5, refine the prompt templates. This phase reveals edge cases and calibrates the system. It takes 2 weeks and involves 1-2 reviewers.

  • Healthcare Logistics AI Glossary: 15 Terms for Order-Status Automation

    Scope and Conventions

    The terms below are alphabetized and defined in the context of a 501-2000 employee healthcare and medtech logistics firm in the USA that is deploying a retrieval-augmented knowledge assistant to handle order and shipment status updates across English, Spanish, and Mandarin. The assistant integrates with the firm’s ERP, CRM, and Slack or Microsoft Teams, uses the Anthropic Claude API for drafting, and operates under a human-in-the-loop approval model to satisfy GDPR. Each entry gives a definition and a one- or two-sentence example showing how the term applies to this specific scenario. The glossary is intended for operations leads, compliance officers, and technical buyers who are evaluating or running an 8-week pilot and need a shared vocabulary before the process audit begins.

    A through M

    Anthropic Claude API is a hosted large-language-model endpoint used for high-quality natural-language generation and classification. In this scenario, it drafts multilingual shipment-delay notices from structured ERP data. Before/after baseline is the set of metrics (cycle time, error rate, language accuracy) captured before the pilot and compared after. Data-processing agreement (DPA) is the GDPR Article 28 contract between the healthcare logistics firm and Forfis as processor. GDPR Article 22(1) prohibits solely automated decisions with legal or similarly significant effects; the human-in-the-loop design keeps the assistant within this boundary. Human-in-the-loop means a person approves any output touching money, health data, or a contract before it sends. Isolated pilot is a fixed-scope, 8-week deployment on one workflow with a measured baseline. Managed AI operations is the delivery model where Forfis owns ongoing monitoring, integration maintenance, and incident response for a monthly fee. Model-agnostic architecture means the language model can be swapped without rewriting the retrieval layer or Slack/Teams integration. Process audit is the structured review of existing workflows that measures cycle time, error rate, and manual touchpoints before automation is designed. Retrieval layer is the component that searches the ERP and CRM for passages relevant to the user’s query and returns them as context for the model. Retrieval-augmented knowledge assistant is the overall system that combines retrieval and a language model to generate grounded, auditable responses. Slack or Microsoft Teams integration is the channel through which the assistant delivers drafts and captures human approvals. Multilingual support coverage requires the system to produce accurate, culturally appropriate responses in English, Spanish, and Mandarin for a US-based healthcare logistics operation. Scaling operations without new hires means using AI to absorb increased order volume without proportionally increasing headcount. 8-week timeline is the pilot duration: week 1 audit, weeks 2-3 build, weeks 4-6 live run, week 7 measurement, week 8 review and roadmap.

    N through Z

    N through Z are not present in this glossary because the 15 terms above cover the full scope of the scenario. However, two additional terms that a compliance officer or technical buyer might encounter in the same engagement are worth noting. Sub-processor is a third party that processes personal data on behalf of the processor (Forfis); under GDPR Article 28(2), the controller must authorize each sub-processor, and the DPA must list them. In this scenario, Anthropic is a sub-processor if patient-identifiable data is sent to its servers; if the data is de-identified before the API call, Anthropic is not a sub-processor for that data. Data-subject-access request (DSAR) is a GDPR Article 15 request from a patient or clinic to see what personal data the firm holds. The AI assistant’s logs (drafted messages, approval timestamps, retrieved context) may contain personal data, so the firm must be able to produce those logs within 30 days. Forfis, as processor, must assist the controller in responding to DSARs under Article 28(3)(e). These two terms are not part of the core 15 but appear in the compliance review that follows the 8-week pilot.

  • Voice Agent for Ticket Triage in a German Logistics Firm

    Background: A Mid-Sized Logistics Firm in Germany

    This case study is a composite based on patterns observed in the field. We do not fake named customers. The company is a mid-sized logistics and supply chain firm based in Germany, with approximately 300 employees. They operate a fleet of delivery vehicles and manage a large volume of customer inquiries, primarily through phone and email. The company is in a growth phase, with increasing demand for their services, but they are constrained by a fixed headcount budget. Their existing stack includes a CRM, a helpdesk system, and a fleet management platform. They are AI-native in their operations, meaning they are open to adopting AI technologies to improve efficiency and scale their operations.

    Challenge: Scaling Operations Without New Hires

    The company faced a significant challenge in scaling their customer support operations without hiring new staff. The volume of customer inquiries was increasing, but the company could not afford to hire additional support agents. The manual data entry process for handling these inquiries was time-consuming and error-prone. The company needed a solution that could automate the triage and routing of customer tickets, reducing the need for manual data entry and allowing their existing team to handle more inquiries efficiently. The deadline for implementing this solution was three months, as the company was preparing for a peak season in their logistics operations.

    Approach: Building a Voice Agent for Ticket Triage

    The company partnered with Forfis, a product studio with eight years of delivery experience, to build a voice agent for customer support. The voice agent was designed to handle incoming calls, transcribe them, classify the intent, and route the tickets to the appropriate queue in the helpdesk system. The agent was built using a model-agnostic architecture, with OpenAI and Anthropic APIs used for high-quality classification, and open-weight models deployed on the company’s own hardware for regulated data. The agent was integrated with the company’s existing CRM and helpdesk via their APIs, ensuring compatibility with existing workflows. The delivery model was a dedicated AI team, with a small team of engineers and product managers working closely with the company to build and maintain the system.

    Outcome: Measurable Improvements in Cycle Time and Error Rate

    The voice agent was deployed in a three-month timeline, with the first month dedicated to the process audit and pilot, the second month to the rollout, and the third month to the managed operation. The pilot was conducted on a subset of customer inquiries, with a measured before/after baseline on cycle time and error rate. The results showed a 40% reduction in cycle time for handling customer inquiries and a 25% reduction in error rate. The voice agent was able to handle a significant volume of calls, reducing the need for manual data entry and allowing the company’s existing team to handle more inquiries efficiently. The company was able to scale their operations without hiring new staff, addressing the challenge of scaling operations without new hires.

    Lessons: Generalizing the Approach for Similar Teams

    • The voice agent was built to be model-agnostic, allowing the company to use different LLMs depending on their needs. This flexibility ensured that the agent could adapt to the company’s specific requirements and constraints.
    • The voice agent was integrated with the company’s existing CRM and helpdesk via their APIs, ensuring compatibility with existing workflows. This integration was crucial for the success of the project, as it allowed the agent to work seamlessly with the company’s existing systems.
    • The voice agent was designed to be human-in-the-loop by default, with a human approving any action that touches money, health data, or a contract. This approach helped build trust in the system and ensured that the agent was used responsibly.
    • The voice agent was built to be scalable, allowing the company to add more calls or features as needed. This scalability ensured that the agent could grow with the company’s business and adapt to changing needs.
    • The voice agent was built to be secure, with data encrypted in transit and at rest. Access to the system was controlled through role-based access control, ensuring that only authorized personnel could access sensitive data.
  • Swiss E-commerce Retailer Cuts Reporting Cycle Time 70% with AI Automation

    Background and Challenge

    This case study is a composite based on patterns observed in the field. It does not represent a single named customer but reflects common challenges and solutions in the e-commerce and retail sector in Switzerland.

    Background
    A mid-sized Swiss e-commerce retailer with 1,200 employees operates across DACH markets. The company uses a custom-built CRM and ERP system, with data stored in on-premise servers. The sales team of 45 handles lead qualification and monthly reporting manually, using spreadsheets and email. The company has no AI in production yet and is looking to reduce manual back-office work while improving lead qualification accuracy.

    Challenge
    The sales team spends 12 hours per week on monthly reporting, manually aggregating data from the CRM, ERP, and web analytics. The process is error-prone, with a 10% error rate in data entry. Lead qualification is inconsistent, with 30% of leads being misclassified, leading to lost opportunities. The company faces GDPR compliance requirements and a deadline to implement improvements before the Q4 peak season.

    Approach
    Forfis conducted an AI automation audit, identifying monthly reporting and lead qualification as high-impact use cases. A fixed-scope pilot was designed to automate these workflows using LangChain and LangGraph for workflow orchestration. The system integrates with the existing CRM and ERP via custom REST APIs and webhooks. A human-in-the-loop model ensures that AI-generated reports and lead scores are reviewed by a human before finalization. The pilot was deployed in two weeks, with a measured before/after baseline on cycle time and error rate.

    Outcome
    The pilot reduced monthly reporting cycle time from 5 days to 1 day, a 70% improvement. The error rate decreased from 10% to 2%, an 80% reduction. Lead qualification accuracy improved from 70% to 95%, with a 25% increase in qualified leads passed to sales. The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building.

    Lessons

    • Start with a fixed-scope pilot to demonstrate ROI quickly.
    • Use a human-in-the-loop model to ensure accuracy and compliance.
    • Integrate with existing systems via APIs rather than replacing them.
    • Measure before/after baselines to quantify impact.
    • Choose a model-agnostic architecture to future-proof the solution.

    Approach: AI Automation Audit and Pilot Design

    The AI automation audit identified two high-impact use cases: monthly reporting and lead qualification. The audit mapped existing workflows, identified bottlenecks, and evaluated the feasibility of automating specific tasks. The results were a prioritized list of use cases, with estimated ROI and implementation complexity.

    Monthly Reporting
    The current process involves manually aggregating data from the CRM, ERP, and web analytics. The sales team spends 12 hours per week on this task, with a 10% error rate in data entry. The AI system automates data collection, validation, and report generation. It uses LangChain to chain prompts and tools, and LangGraph to define stateful, multi-step workflows. The system integrates with the existing CRM and ERP via custom REST APIs and webhooks, ensuring data integrity and real-time updates.

    Lead Qualification
    The current process is inconsistent, with 30% of leads being misclassified. The AI system uses a classification model to score leads based on predefined criteria, such as company size, industry, and engagement level. The model is trained on historical data and fine-tuned using feedback from the sales team. A human-in-the-loop model ensures that AI-scored leads are reviewed by a human before they are passed to sales, ensuring accuracy and context.

    GDPR Compliance
    The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building. Data minimization is implemented, and data subjects can exercise their rights. The legal basis for processing is documented, and third-party AI APIs are GDPR-compliant. The system uses open-weight models on the client’s own hardware, ensuring that regulated data does not leave the building.

    Outcome: Measured Impact on Cycle Time and Error Rate

    The pilot was deployed in two weeks, with a measured before/after baseline on cycle time and error rate. The system was integrated with the existing CRM and ERP via custom REST APIs and webhooks, ensuring seamless data flow. The human-in-the-loop model was implemented, with a review dashboard for the sales team to approve AI-generated reports and lead scores.

    Cycle Time
    The monthly reporting cycle time was reduced from 5 days to 1 day, a 70% improvement. The AI system automates data collection, validation, and report generation, eliminating manual data entry and aggregation. The sales team spends 2 hours per week on review and approval, compared to 12 hours previously.

    Error Rate
    The error rate in monthly reporting decreased from 10% to 2%, an 80% reduction. The AI system validates data in real-time, flagging anomalies and inconsistencies. The human-in-the-loop model ensures that errors are caught and corrected before the report is finalized.

    Lead Qualification Accuracy
    Lead qualification accuracy improved from 70% to 95%, with a 25% increase in qualified leads passed to sales. The AI system scores leads based on predefined criteria, and the human-in-the-loop model ensures that misclassified leads are corrected. The sales team reports a 15% increase in conversion rates, attributed to more accurate lead qualification.

    GDPR Compliance
    The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building. The legal basis for processing is documented, and data subjects can exercise their rights. The system uses open-weight models on the client’s own hardware, ensuring that regulated data does not leave the building.

    Lessons for Similar Teams

    The pilot demonstrated significant improvements in cycle time, error rate, and lead qualification accuracy. The system is GDPR-compliant and integrated with existing systems via APIs. The human-in-the-loop model ensures accuracy and compliance, while the model-agnostic architecture provides flexibility and future-proofing.

    Scalability
    The system can be scaled to automate other workflows, such as invoice processing and document extraction. The model-agnostic architecture allows for switching between different AI models, based on cost, performance, and compliance requirements. The system can be extended to other departments, such as marketing and customer service, with minimal changes.

    Cost Efficiency
    The pilot reduced manual back-office work by 80%, saving 10 hours per week. The cost of the AI system is offset by the reduction in manual effort and the increase in qualified leads. The system is cost-effective, with a payback period of less than 3 months.

    Risk Mitigation
    The human-in-the-loop model mitigates the risk of errors and ensures compliance with regulations. The model-agnostic architecture mitigates vendor lock-in and allows for future-proofing. The system is GDPR-compliant, with data processed on-premise and no personal data leaving the building.

    Next Steps
    The company plans to roll out the system to other departments, such as marketing and customer service. The system will be extended to automate other workflows, such as invoice processing and document extraction. The company will continue to measure the impact of the system on key metrics, such as cycle time, error rate, and lead qualification accuracy.

  • Fixed-Scope AI Pilot vs. Full Rollout: A Fintech’s 6-Month Decision

    What Is Being Compared: Fixed-Scope Pilot vs. Full-Scale Rollout

    The two options under comparison are a fixed-scope pilot and a full-scale rollout of AI automation across a 2,000+ employee fintech firm in the UK. The pilot targets one workflow — in this case, monthly reporting compilation and internal knowledge search over Notion and Confluence — with a 6-week delivery window, a measured before/after baseline on cycle time and error rate, and a go/no-go decision at the end. The full-scale rollout deploys AI process automation across multiple departments simultaneously: invoice processing, ticket triage for round-the-clock customer response, HR and recruiting workflow orchestration, and a retrieval-augmented assistant over the company’s documentation. Both options use the same underlying architecture: n8n for workflow orchestration, a model-agnostic AI layer (OpenAI or Anthropic APIs for non-regulated data, open-weight models on the client’s hardware for PCI DSS-sensitive data), and human-in-the-loop approval for anything touching money, contracts, or health data. The difference is scope, timeline, and risk exposure.

    Criteria for the Comparison

    The following criteria determine which option fits a fintech firm’s constraints. PCI DSS compliance is the hard gate: any workflow that touches cardholder data must run on-premises or in a PCI-compliant enclave, which rules out cloud-only model APIs for those specific flows. Cycle time reduction is measured in hours per report or per ticket, not in vague efficiency gains. Error rate is tracked as a percentage of transactions requiring manual correction. Integration depth counts the number of existing systems (CRM, ERP, helpdesk, Notion, Confluence) that the automation must connect to without replacing them. Vendor lock-in is assessed by whether the architecture can swap models or orchestration tools without rework. Timeline is the calendar duration from kickoff to managed operation. Cost is the total engagement fee plus ongoing managed operation, expressed in GBP. Scalability is the number of additional workflows or departments that can be added without rebuilding the core architecture.

    Comparison Table

    Criterion Fixed-Scope Pilot Full-Scale Rollout
    PCI DSS compliance One workflow isolated; open-weight model on-premises for cardholder data Multiple workflows; requires a PCI-compliant enclave for all payment-related flows
    Cycle time reduction Measured on one workflow (e.g., monthly reporting: 14 hrs → 2 hrs) Measured across 4-6 workflows; aggregate reduction depends on each workflow’s baseline
    Error rate Baseline established in week 1; target <2% by week 6 Baselines established per department; target <3% aggregate by month 4
    Integration depth 2-3 systems (Notion, Confluence, one CRM) 6-10 systems (CRM, ERP, helpdesk, Notion, Confluence, HRIS, payment gateway)
    Vendor lock-in Low; n8n workflows are portable; model can be swapped Moderate; more integrations increase switching cost, but n8n remains the orchestration layer
    Timeline 6 weeks to pilot completion; 2 weeks to decision 6 months to full managed operation across departments
    Cost (GBP) £18,000–£35,000 for the pilot £120,000–£250,000 for the full engagement plus £4,000–£8,000/month managed operation
    Scalability One workflow; scaling requires a new pilot per department Multi-department from day one; new workflows added to the existing n8n architecture

    When the Fixed-Scope Pilot Wins

    The fixed-scope pilot wins when the firm has not yet established a baseline for AI automation and needs to prove value before committing to a multi-department rollout. For a 2,000+ employee fintech in the UK, the pilot on monthly reporting and internal knowledge search over Notion and Confluence delivers a measurable result in 6 weeks: cycle time drops from 14 hours to 2 hours per report, and the error rate on data extraction falls from 8% to under 2%. The go/no-go decision is based on these numbers, not on a qualitative assessment. The pilot also validates the n8n orchestration layer and the human-in-the-loop approval gates without exposing the entire back office to change. If the pilot meets its targets, the firm has a proven template for the next workflow.

    The full-scale rollout wins when the firm has already completed a process audit, has identified 4-6 high-impact workflows, and has the IT capacity to manage parallel integrations. For a fintech with PCI DSS obligations, the rollout must include an on-premises open-weight model for any workflow that touches cardholder data, while non-regulated workflows (ticket triage, HR recruiting, knowledge search) can use OpenAI or Anthropic APIs. The 6-month timeline assumes that the process audit is complete, that the n8n environment is provisioned, and that each department has a named owner for the integration work. The rollout delivers aggregate cycle time reduction across the firm, but it requires a managed operation team from month 3 onward to handle model updates, integration drift, and new workflow requests.

    When the Full-Scale Rollout Wins

    The full-scale rollout is the right choice when the firm’s process audit has already identified multiple workflows with high impact and low integration complexity, and when the IT team can support parallel workstreams. For a 2,000+ employee fintech in the UK, this means the audit has scored invoice processing, ticket triage, HR and recruiting workflow orchestration, and internal knowledge search as the top four candidates. The rollout deploys all four within 6 months, with the PCI DSS-sensitive workflows (invoice processing, payment-related ticket triage) running on open-weight models on the client’s hardware, and the non-regulated workflows (HR recruiting, knowledge search) using OpenAI or Anthropic APIs. The n8n orchestration layer is shared across all workflows, so a change to one integration (e.g., a CRM API update) is applied once, not four times. The managed operation team, staffed from month 3, handles model retraining, integration monitoring, and new workflow requests. The cost is higher — £120,000 to £250,000 for the engagement plus £4,000 to £8,000 per month for managed operation — but the aggregate cycle time reduction across four workflows justifies the investment within 12 months for a firm of this size.

    Recommendation for the Scenario

    For a 2,000+ employee fintech in the UK with PCI DSS obligations, the recommendation is a fixed-scope pilot first, followed by a phased rollout. The pilot targets monthly reporting compilation and internal knowledge search over Notion and Confluence, with a 6-week delivery window and a measured baseline on cycle time and error rate. The pilot validates the n8n orchestration layer, the human-in-the-loop approval gates, and the model-agnostic architecture without exposing the payment processing workflows to change. If the pilot meets its targets — cycle time reduced from 14 hours to under 3 hours, error rate below 2% — the firm proceeds to a phased rollout over the remaining 4 months of the 6-month timeline. The rollout adds invoice processing, ticket triage for round-the-clock customer response, and HR and recruiting workflow orchestration, with PCI DSS-sensitive workflows running on open-weight models on the client’s hardware. The total engagement cost is £150,000 to £280,000, with managed operation at £5,000 to £8,000 per month from month 4 onward. This approach limits risk, delivers a measurable result in 6 weeks, and scales the architecture across departments without rebuilding it.