Tag: UK

  • UK B2B SaaS Firm Cuts Invoice Cycle Time 74% With On-Premise AI Pilot

    Background: A 30-Person B2B SaaS Firm in Manchester

    This case study is a composite drawn from patterns observed across multiple engagements. No named customer is represented. The firm described here is a 30-person B2B SaaS company based in Manchester, selling a project-management platform to mid-market clients across the UK and Ireland. The operations team of four handles supplier invoices, delivery notes, and credit notes for a mix of cloud hosting, office supplies, and professional services vendors. The existing stack is a standard ERP (Xero for accounting, a lightweight project-management tool for internal tracking) and Slack as the primary communication channel. No AI system is in production anywhere in the company. The trigger for change is not a technology initiative but a headcount constraint: the operations lead has been absorbing invoice processing work that was previously split across two part-time staff, and the founder has set a deadline to reduce the manual workload before the next hiring cycle in Q3.

    Challenge: Four-Day Cycle Time and a GDPR Gap

    The operations lead processes roughly 180 supplier invoices per month, each requiring manual data entry into Xero: vendor name, line items, tax codes, and total amount. The median cycle time from invoice receipt to payment approval is four business days, with a long tail of invoices taking nine to twelve days when the operations lead is pulled into client escalations. The error rate on manual data entry is 18 percent, measured over a two-week sample in the audit phase. Each error triggers a correction cycle that adds 20 to 35 minutes of senior staff time. The compliance pressure is GDPR: the invoices contain personal data (vendor contact names and email addresses), and the firm’s data protection officer has flagged that the current manual process, which involves forwarding PDFs between personal email accounts and the operations lead’s inbox, does not meet the Article 5(1)(f) integrity and confidentiality requirement. The deadline is eight weeks: the founder wants a working pilot before the Q3 hiring decision, and the data protection officer wants a documented DPIA before any new system touches the invoice data.

    Approach: Two-Week Audit, Fixed-Scope Pilot, On-Premise Inference

    The engagement starts with a two-week AI automation audit. The team maps every document that enters the operations workflow, measures the current cycle time and error rate, and scores each workflow on volume, error cost, and automation feasibility. Invoice processing wins the composite score: 180 documents per month, a 18 percent error rate with a 20-to-35-minute correction cost per error, and a document format that maps cleanly to a structured extraction task. The pilot scope is fixed: extract vendor name, line items, tax codes, and total amount from PDF invoices, write the data to Xero via the API, and route flagged fields to the operations lead in Slack for approval. The architecture is model-agnostic: the orchestration service routes inference to an on-premise vLLM endpoint running a 7B-parameter open-weight model, because the GDPR review confirms that the invoice data cannot be sent to a cloud API. The Slack integration is built with the Slack Bolt framework, posting flagged items to a dedicated channel with approve and reject buttons. The human-in-the-loop gate is hard-coded: any field with a confidence score below 0.92 is flagged for human review.

    Outcome: 74 Percent Cycle-Time Reduction in Six Weeks

    The pilot runs for six weeks after the audit, with a two-week shadow period at the end where the AI drafts and the operations lead approves every output. The before/after baseline is measured over the final two weeks of the shadow run. The median cycle time drops from 4.2 days to 1.1 days, a 74 percent reduction. The manual correction rate falls from 18 percent to 4 percent. The operations lead reviews 22 flagged items per day in week one, dropping to 8 per day by week six as the model’s confidence improves on the firm’s specific vendor set. The senior operations lead, who had been spending roughly 14 hours per week on invoice processing, reports spending 3 hours per week on the approval queue and 2 hours per week on exception handling. The GDPR DPIA is completed in week three, documenting the data flows, the retention policy (invoices retained for seven years per UK tax law, extracted data retained for 12 months), and the human-in-the-loop approval gate. The on-premise hardware is a single workstation with an NVIDIA L40S 48 GB GPU, provisioned in week one and running the vLLM inference server for the duration of the pilot.

    Lessons for Similar Teams

    • The audit is the product, not the pilot. The two-week process audit produced a one-page baseline report that the client retained for internal reporting and the GDPR accountability record. The pilot was the validation, but the audit was the deliverable that justified the investment. Teams that skip the audit and jump straight to a pilot often discover mid-engagement that the workflow they chose is not the highest-impact one. – On-premise hardware is a procurement decision, not a technical one. The L40S workstation was ordered in week one, before the audit was complete. The lead time for GPU hardware in the UK is four to six weeks. Teams that order the hardware after the audit is done lose two to three weeks of the pilot timeline. – The Slack integration is the adoption lever. The operations lead approved 22 items per day in week one without any training, because the interface was the tool she already used. A separate dashboard would have added friction and likely reduced the approval rate below the threshold needed for the baseline comparison. – The confidence threshold is a tuning parameter, not a fixed constant. The 0.92 threshold for monetary fields was set in week one and adjusted to 0.95 in week four after the model’s performance on the firm’s specific vendor set improved. Teams that treat the threshold as a fixed constant either over-flag (wasting senior time) or under-flag (letting errors through). – The GDPR DPIA is a two-week task, not a one-day checkbox. The data protection officer spent three hours in week two reviewing the data flow diagram and two hours in week three reviewing the retention policy. The DPIA was completed in week three, not week one, because the model’s training data provenance had to be documented before the review could be signed off.
  • 8-Week AI Integration Sprint Checklist for UK Professional Services Firms

    1. Audit workflows and pick one pilot task

    Before writing a single line of code, map every manual workflow in sales, finance, and operations. Score each on volume, error rate, and cycle time. Pick the workflow with the highest volume and lowest complexity for the pilot. For a 201-500 employee firm, this is usually invoice processing, document extraction from client contracts, or lead qualification from inbound forms. The pilot should replace one specific task, not an entire department. Measure baseline cycle time and error rate before the pilot starts, then compare after 4 weeks of operation. This baseline becomes your proof of value when you scale across departments.

    2. Measure baseline cycle time and error rate

    Record the current cycle time and error rate for the chosen workflow before any automation. For document extraction, time how long a person takes to parse a typical invoice or contract and count how many fields they get wrong. For lead qualification, measure how long it takes to respond to an inbound lead and what percentage of leads are misclassified. Use a simple spreadsheet or your existing CRM’s audit log. This baseline is your control group. Without it, you cannot prove the AI improved anything, and you cannot justify scaling the solution to other departments later.

    3. Choose the model stack for GDPR compliance

    Run the document extraction pipeline on open-weight models deployed in the firm’s VPC or on-premises server. This keeps regulated client data local and satisfies GDPR data residency requirements. Use OpenAI API for the customer-facing assistant that drafts responses to client queries in Slack or Microsoft Teams, since the data in those channels is less sensitive. For lead qualification, use OpenAI API to score and route leads, but require human approval before any lead enters the CRM for contract negotiation. This hybrid approach keeps regulated data local while leveraging frontier models for unstructured text tasks.

    4. Build the human-in-the-loop approval flow

    Configure Slack or Microsoft Teams as the approval channel for human-in-the-loop workflows. When the AI extracts data from a document or qualifies a lead, it sends a notification to the responsible person’s Slack or Teams channel with a one-click approve or reject button. The person reviews the extracted data or lead score, approves it, and the system writes the approved data to the CRM or ERP. This keeps the approval step in the tool the team already uses, reducing friction. Log every approval action with timestamp and user ID for GDPR Article 30 accountability records.

    5. Connect the AI layer to existing CRM and ERP

    Integrate the AI pipeline with your existing CRM, ERP, and helpdesk through their APIs rather than replacing them. For a professional services firm, this usually means connecting to Salesforce, HubSpot, or Microsoft Dynamics for CRM data, and to Xero, QuickBooks, or SAP for ERP data. The AI layer sits on top of these systems, reading from and writing to them via API calls. This preserves the firm’s existing data architecture and avoids the cost and risk of migrating to a new platform. The integration sprint should deliver working API connections by day 10 of the 8-week timeline.

    6. Document GDPR Article 30 accountability records

    Document the AI’s decision logic in your GDPR Article 30 records. For each automated decision, record what data the AI used, what model made the decision, and what human approved it. This satisfies GDPR Article 22’s requirement for meaningful human intervention in automated decision-making. For lead qualification, document that the AI scores leads but a human reviews any lead flagged for contract negotiation. For document extraction, document that the AI parses documents but a person verifies extracted data before it enters the ERP. These records protect the firm if a data subject requests an explanation of an automated decision.

    7. Measure pilot results and plan departmental scaling

    After the 4-week pilot, compare the AI’s cycle time and error rate against the baseline you recorded in step 2. If the AI reduced cycle time by 50% or more and cut error rates by 70% or more, the pilot succeeded. Present these numbers to the firm’s leadership with a clear recommendation to scale the solution to other departments. For a 201-500 employee firm, scaling usually means applying the same AI pipeline to additional document types, lead sources, or customer-facing channels. The 8-week sprint should end with a working pilot, measured results, and a documented plan for rollout.

  • How a 120-Person UK Advisory Firm Cut Contract First-Response Time to 38 Minutes

    Background: A 120-Person UK Advisory Firm

    This case study is a composite based on patterns Forfis has observed across multiple engagements in the UK professional services sector. No named client is represented; the figures are drawn from real pilot baselines and post-rollout measurements. The company in this story is a 120-person firm providing legal and financial advisory services to mid-market clients in London and Manchester. It runs on Microsoft 365, a mid-tier CRM, and a document management system that predates the current team. The firm sits in the 51-200 employee band, which means it has the volume to justify automation but not the headcount to run a dedicated AI team.

    The Challenge: 4.2-Hour First Response and a 14-Week Deadline

    The firm’s contract review process was the bottleneck. Clients sent contracts via email; a paralegal or junior associate extracted key clauses, flagged risks, and drafted a response. First-response time averaged 4.2 hours, with a peak of 11 hours during quarter-end. The error rate on clause extraction was 6.1%, meaning roughly one in sixteen contracts required a second pass. GDPR Article 22 required that no automated system make a decision solely on the basis of profiling without human oversight. The firm also faced a deadline: a major client contract was due in 14 weeks, and the existing team could not absorb the volume without hiring two additional paralegals at a cost of approximately GBP 78,000 per year.

    Approach: Audit, Pilot Sprint, and Model-Agnostic Integration

    Forfis began with a two-week process audit. The team mapped every step of the contract review workflow, measured cycle time and error rate on a sample of 200 contracts, and scored each sub-task by volume, error rate, and regulatory exposure. The audit produced a phased roadmap: a fixed-scope pilot on clause extraction and risk flagging, followed by rollout to the financial advisory team. The pilot used the Anthropic Claude API for extraction and classification, with a human-in-the-loop approval gate for anything touching contract terms. The integration sprint ran five weeks: Forfis built the extraction pipeline, connected it to the firm’s CRM and Microsoft Teams, and shipped a Slack channel where flagged clauses appeared as threaded messages with confidence scores. The model-agnostic architecture meant the firm could swap to an open-weight model on its own hardware if data residency requirements tightened.

    Outcome: 38-Minute First Response and a 1.4% Error Rate

    After the five-week pilot, the firm measured the new baseline. First-response time dropped from 4.2 hours to 38 minutes. Extraction error rate fell from 6.1% to 1.4%. The paralegal team redirected its time from manual extraction to higher-value risk analysis. The firm did not hire the two additional paralegals. Rollout to the financial advisory team took three additional weeks, extending the total engagement to three months. The managed operation phase began in week 13, with Forfis monitoring model performance, handling edge cases, and tuning the extraction prompts. The client retained ownership of the integration code and the Teams/Slack configuration, so it could extend the workflow internally without a new engagement.

    Lessons for Similar Teams

    • Baseline before you build. The audit’s 200-contract sample gave the firm a defensible before/after metric. Without it, the pilot’s success would have been anecdotal. Teams that skip the baseline struggle to justify scaling to stakeholders.
    • Human-in-the-loop is not a compromise. The approval gate for contract terms kept the firm compliant with GDPR Article 22 while still cutting manual effort. The gate added 12 seconds per clause but prevented a single high-risk auto-approval that would have required a client call.
    • Model-agnosticism is a risk hedge. The firm’s data residency requirements could have shifted mid-engagement. Because the architecture supported open-weight models on local hardware, Forfis could swap the backend without rewriting the integration layer.
    • Integration over replacement. Plugging into the existing CRM and Teams meant the team did not have to learn a new tool. Adoption was near-complete in the first week because the workflow appeared in the channel they already checked every morning.
    • Fixed-scope pilots reduce scope creep. The five-week sprint had a defined set of document types and a defined approval gate. Adding new document types was a separate decision, not a mid-sprint change request.
  • Voice Agent for Lead Qualification in a UK Fintech: A 4-Week Pilot

    The Problem: Inbound Calls and Back-Office Errors in a UK Fintech

    A UK fintech with 2,000+ employees is drowning in inbound calls. Sales reps spend 40% of their day on the phone, qualifying leads that are often unqualified. The back office spends 30% of its time manually entering data from these calls into Salesforce, with an error rate of 8%. The cost per support ticket is £12, and the company is losing deals because reps are not available to follow up on qualified leads. The problem is not a lack of tools; it is a lack of automation. The company needs a system that can handle the first 60 seconds of a call, extract the relevant data, and update the CRM without human intervention. The constraint is PCI DSS: the system cannot store or process card numbers. The solution is a voice agent that runs on an on-premise open-weight model, integrated with Salesforce, and approved by a human before any data is committed.

    The Mechanism: A Three-Stage Voice Agent Pipeline

    The voice agent uses a three-stage pipeline. First, a speech-to-text engine (Whisper or Deepgram) transcribes the call in real time. Second, an on-premise open-weight model (Llama 3 70B or Mistral 7B) processes the transcript. The model is prompted to extract specific fields: company name, job title, budget range, and timeline. The model outputs a structured JSON object. Third, the JSON is mapped to the corresponding fields in Salesforce via the REST API. If the model is uncertain about a field, it flags it for human review. The human agent sees the transcript, the extracted fields, and a confidence score, and can approve, edit, or reject the entry before it is committed to the CRM. The entire pipeline runs in under 2 seconds, so the agent can respond to the lead in real time. The on-premise model ensures that no data leaves the building, which is critical for PCI DSS compliance.

    Trade-offs: API vs. On-Premise, Automation vs. Human-in-the-Loop

    The architect faces three key trade-offs. First, the choice between an API-based LLM and an on-premise open-weight model. The API is faster to deploy and cheaper for low volume, but it sends data to a third party, which is a PCI DSS risk. The on-premise model is more expensive to set up (around £20,000 for hardware) but keeps data in-house. Second, the choice between a fully automated system and a human-in-the-loop system. Full automation is faster but riskier; a human-in-the-loop system is slower but safer. For a fintech, the human-in-the-loop approach is non-negotiable. Third, the choice between a narrow use case and a broad one. A narrow use case (lead qualification) is easier to scope and deliver in 4 weeks, but it does not address the back-office error rate. A broad use case (all inbound calls) is more valuable but harder to deliver in 4 weeks. The recommendation is to start with a narrow use case and expand from there.

    Recommendation: A 4-Week Pilot for Lead Qualification

    The recommendation is to run a 4-week pilot focused on lead qualification. Week 1: process audit and baseline measurement. The team measures the current error rate (8%) and cycle time (15 minutes) for lead qualification. Week 2: build the voice agent, integrate with Salesforce, and set up the human-in-the-loop approval workflow. Week 3: closed beta with a small group of real leads. The team tunes the model and fixes edge cases. Week 4: full rollout to the sales department, with daily monitoring of error rates and cycle times. The success criteria are a 20% reduction in error rate and a 30% reduction in cycle time. If the pilot meets these criteria, the team moves to rollout, which involves scaling the solution to other departments and integrating it with additional systems. The pilot is scoped to a single department to keep the timeline realistic and the risk manageable.

  • UK Fintech Cuts Invoice Errors to 0.9% in 8 Weeks with n8n and a Local LLM

    Background: A UK Fintech’s Back-Office Bottleneck

    This case study is a composite based on patterns observed across multiple engagements. We do not name real clients. The company described here is a mid-size UK fintech operating a payments platform for B2B clients, with 1,200 employees across London and Manchester. The back-office operations team handled supplier invoices, payment reconciliation, and vendor onboarding. The stack included a UK-hosted ERP, a Zendesk helpdesk, a custom payments gateway, and a mix of spreadsheets and manual data entry for invoice processing. The company had already deployed a basic RAG assistant over its internal documentation but had not touched invoice processing. The operations director flagged that the back-office error rate had crept to 3.8% over the prior two quarters, driven by data-entry mistakes in vendor codes, tax fields, and payment terms. Each error triggered a reconciliation cycle that averaged 6.5 business days. The board had set a target: reduce the cost per support ticket and the back-office error rate within one fiscal quarter, without adding headcount. The compliance team confirmed that any solution touching invoice data had to satisfy PCI DSS Requirement 3.5.1 (no full PAN storage) and the client’s internal data-residency policy, which prohibited sending invoice data to any third-party API outside the UK.

    Challenge: PCI DSS, Data Residency, and an 8-Week Deadline

    The operations director’s brief was specific: cut the back-office error rate from 3.8% to under 1% within 8 weeks, without adding headcount, and without sending invoice data to any third-party API. The compliance team added a hard constraint: PCI DSS Requirement 3.5.1 prohibited storing the full Primary Account Number on any system, and the client’s internal data-residency policy meant no invoice data could leave the building. The timeline was fixed by the board’s fiscal-quarter deadline. The team had 12 back-office staff processing roughly 4,200 supplier invoices per month across three departments. The manual process involved scanning PDFs, keying data into the ERP, and flagging discrepancies for review. The error rate was not uniform: vendor-code mismatches accounted for 40% of errors, tax-field mistakes for 30%, and payment-term misclassification for the remaining 30%. The operations director also wanted a measured before/after baseline on cycle time and error rate, not just a qualitative improvement. The challenge was not whether an LLM could read an invoice; it was whether the system could do so inside a PCI DSS boundary, on the client’s own hardware, with a human approval step for anything touching a payment amount.

    Approach: n8n Orchestration with a Local LLM and Human-in-the-Loop Approval

    The engagement started with a two-week process audit. We mapped the invoice lifecycle from receipt to payment, identified the three error-prone steps (data entry, classification, and discrepancy flagging), and measured the baseline: median cycle time of 4.2 days, error rate of 3.8%, and an average of 11 minutes of manual work per invoice. The architecture was model-agnostic by design. The n8n workflow ran on the client’s own VPS in a UK region, orchestrating the pipeline: pull invoice from the ERP via a custom REST endpoint, strip any PAN fields before the document reached the model, call a local Llama 3 70B on the client’s A100 GPU, validate the output against a JSON schema, and push the structured data back to the ERP via webhook. The helpdesk integration used Zendesk’s REST API to create a ticket when a human approval was needed. The human-in-the-loop step was non-negotiable: any field touching a payment amount above GBP 5,000 or a contract clause required a reviewer’s sign-off. The n8n workflow logged every approval action with a timestamp, so the team could measure reviewer latency and field-level changes. The pilot covered one invoice category (supplier invoices in GBP, under GBP 25,000) and one department (AP).

    Outcome: 0.9% Error Rate, 1.1-Day Cycle Time, PCI DSS Sign-Off

    The pilot ran in shadow mode for six weeks: the model processed every invoice in parallel with the manual process, and the team compared outputs. After shadow mode, the system went live with human-in-the-loop approval for the first two weeks, then gradual autonomy. The measured outcomes: median cycle time dropped from 4.2 days to 1.1 days; the error rate fell from 3.8% to 0.9%; and the approval queue shrank to 12% of volume after six weeks. The cost per support ticket in the back-office context (reconciliation time plus late-payment penalties) dropped from an estimated GBP 180-240 per error to under GBP 40. The 12 back-office staff were not laid off; they were redeployed to handle the 12% of invoices that still required human review, plus new vendor onboarding tasks that had been backlogged. The n8n workflow handled 88% of invoices end-to-end without human intervention. The model never saw a full PAN; the n8n workflow stripped PAN fields before the document reached the model, and the output schema rejected any field containing a 13- to 19-digit numeric string. The client’s PCI DSS assessor signed off on the architecture in the final week of the pilot.

    Lessons for Teams Scaling AI Across Departments

    Five lessons from this engagement generalize to similar teams scaling AI across departments in regulated environments. First, the process audit is not optional. The two-week audit identified that 40% of errors came from vendor-code mismatches, which a generic OCR solution would have missed. The n8n workflow included a vendor-code validation step that cross-referenced the ERP’s vendor master before the model even ran. Second, model-agnosticism is a risk hedge, not a buzzword. The team swapped from Llama 3 70B to a smaller 8B model for a specific document type (credit notes) where the 70B was overkill and the 8B was 3x faster on the client’s hardware. The n8n workflow logic did not change. Third, the human-in-the-loop step must be measurable. Logging every approval action with a timestamp let the team prove that reviewer latency dropped from 11 minutes to 2.3 minutes per invoice as the model’s accuracy improved. Fourth, the 8-week timeline was only achievable because the pilot scope was fixed to one invoice category and one department. Trying to cover all three departments in 8 weeks would have pushed the timeline to 14 weeks. Fifth, the managed operations contract was not an afterthought. The 12-month post-rollout contract covered model monitoring, prompt tuning, and n8n workflow maintenance, which kept the error rate at 0.9% rather than drifting back to 2% as invoice formats changed.

  • GDPR-Compliant AI Candidate Screening Pilot for UK Logistics Firms

    The Problem: Routine Screening Consumes Senior Recruiter Hours

    Your recruiting team spends 15-25 minutes per CV screening 50-200 applications weekly. That is 12-40 hours of senior recruiter time consumed by routine extraction and matching. The problem is not volume alone; it is that the work is repetitive, rule-based, and error-prone. Missed qualifications, inconsistent scoring, and slow cycle times delay hiring in a logistics market where driver and warehouse roles turn over at 30-40% annually. You need to free senior staff from routine work while keeping the process compliant with GDPR, particularly Article 22 on automated decision-making. The solution is a fixed-scope pilot: one workflow, one document type, one integration, and a measured before/after baseline on cycle time and error rate.

    Prerequisites Before You Start

    • Process audit completed: You have documented the current screening workflow, including cycle time (minutes per CV), error rate (missed qualifications per 100 screened), and cost per screened candidate. These are your before/after baselines.
    • API access provisioned: REST API credentials for your ATS (e.g., Greenhouse, Lever, or Workable) and Google Workspace (Drive, Gmail, or Chat). Confirm the ATS supports webhook or polling for new applications.
    • Named human approver: A recruiter or hiring manager available for at least 30 minutes per day to review AI recommendations and approve or reject shortlists.
    • Data protection documentation: Your Data Protection Impact Assessment (DPIA) updated to include automated screening. Your records of processing activities (Article 30) list the AI system, data flows, and retention period.
    • Model access: API keys for OpenAI or Anthropic, or a self-hosted open-weight model (Llama 3 70B, Mistral 8x22B) on your own hardware if CVs contain special category data.
    • LangChain and LangGraph environment: Python 3.10+, LangChain 0.1+, LangGraph 0.0.5+, and a vector store (ChromaDB or Pinecone) for document retrieval.

    Step 1: Define the Pilot Scope and Success Criteria

    Define the exact scope: one document type (CVs), one job family (e.g., warehouse operatives), one integration (Google Workspace), and one approval gate. Write a one-page scope document specifying the input (PDF or DOCX CVs from the ATS), the output (structured JSON with skills, experience, location, and a 0-100 score), and the success criteria (cycle time under 5 minutes per CV, error rate under 5%). This prevents scope creep during the two-week pilot. If the audit reveals more than two distinct CV formats or the ATS lacks a REST API, narrow the scope to one format or extend the timeline to three weeks. The scope document is your contract with the pilot: anything outside it is a separate engagement.

    Step 2: Build the Document Extraction Pipeline

    Build the extraction pipeline using LangChain’s document loaders. For PDFs, use PyPDFLoader or UnstructuredPDFLoader to handle scanned and digital documents. For DOCX, use Docx2txtLoader. Store extracted text in a vector store (ChromaDB for local, Pinecone for cloud) with metadata: candidate name, job applied, upload timestamp, and source file ID. The extraction node in your LangGraph workflow outputs structured JSON. Use a prompt template that specifies the exact fields to extract: skills (array), years_experience (integer), location (string), education (string), and availability (string). Test the pipeline on 20 sample CVs from your ATS before moving to the next step. Measure extraction accuracy: compare extracted fields against the raw document for each sample.

    Step 3: Orchestrate the Workflow with LangGraph

    Define the LangGraph state machine with four nodes: extract, score, approve, and notify. The extract node calls the extraction pipeline. The score node calls the LLM with a rubric prompt: “Score this candidate 0-100 based on the following criteria: minimum 2 years warehouse experience (40 points), valid driving license (30 points), availability for shift work (20 points), location within 20 miles of depot (10 points).” The approve node pauses the graph and sends a notification to the human approver via Google Workspace API (email or Chat message) with the AI’s recommendation, extracted data, and confidence score. The notify node updates the ATS with the screening status. Use LangGraph’s checkpointing to persist state: if the approver takes 24 hours to respond, the graph resumes from the approve node without re-running extraction or scoring.

    Step 4: Integrate with Google Workspace for Notifications and Storage

    Integrate with Google Workspace using the Google API client library. For notifications, use the Gmail API to send an email to the approver with the AI’s recommendation in the body and a link to the candidate’s profile in the ATS. For document storage, use the Drive API to store CVs in a restricted folder with no external sharing. Set folder permissions to “Only specific people” and add the approver and IT admin. For audit logging, use the Chat API to post a summary of each screening decision to a private channel: candidate name, score, approver decision, and timestamp. This creates a tamper-evident audit trail that satisfies GDPR Article 30 and supports your DPIA. Test the integration with a test account before connecting to production data.

    Step 5: Run the Pilot and Measure Before/After Metrics

    Run the pilot on 50 real CVs from your ATS over five business days. Measure three metrics: cycle time (minutes from CV upload to approver decision), error rate (number of missed qualifications or incorrect shortlists per 50 screened), and approver time (minutes spent reviewing each AI recommendation). Compare against your baseline from the process audit. If cycle time drops from 15 minutes to under 5 minutes and error rate stays under 5%, the pilot meets success criteria. If error rate exceeds 5%, review the scoring rubric: it may be too vague or the extraction pipeline may be missing fields. If approver time exceeds 10 minutes per CV, the AI’s recommendation may be unclear: add a confidence score and a one-sentence justification to the notification. Document all findings in a pilot report with before/after metrics.

  • AI Ticket Triage Glossary for UK Professional Services Firms

    Scope and Conventions

    The following terms are defined in the context of a UK professional services firm with 501 to 2,000 employees that is deploying an AI ticket triage and routing system. The firm uses the OpenAI API for classification, integrates with Google Workspace for internal notifications, and operates under ISO 27001. Each entry gives a concise definition and a one- or two-sentence example drawn from the firm’s specific use case. The glossary is alphabetized and covers the full delivery cycle from process audit through managed operation.

    A through D

    Before/After Baseline is the measured comparison of cycle time and error rate before and after the AI layer goes live. In the firm’s pilot, the baseline is a 200-ticket sample scored for misrouting and a 10-business-day window tracking median time from ticket creation to first human action. The delta between the two measurements is the primary metric the firm uses to justify rollout to additional departments.

    Data Enrichment and Cleanup refers to the automated step where the AI model fills in missing fields on a ticket, such as client name, service line, or urgency level, by extracting them from the ticket body and cross-referencing the CRM. In the firm’s workflow, this step reduces the time a senior associate spends re-keying information from a client email into the helpdesk, freeing roughly 12 minutes per ticket for higher-value work.

    Dedicated AI Team is a fixed group of engineers and a product owner assigned to the firm for the duration of the engagement. The team handles the process audit, builds the pipeline, runs the pilot, and manages the system after go-live. The firm’s internal IT team retains ownership of the helpdesk and Google Workspace configurations, so the AI team’s role is additive rather than replacing existing staff.

    Document and Data Extraction Pipelines are the automated workflows that pull structured data from unstructured inputs such as client emails, PDFs, and ticket bodies. In the firm’s case, the pipeline extracts the client’s name, the service requested, and the deadline from a free-text ticket, then writes those fields into the helpdesk record. The pipeline runs on every new ticket and takes under 2 seconds to complete.

    H through O

    Human-in-the-Loop is the default operating mode where the model drafts or classifies, and a person approves anything that touches money, health data, or a contract. In the firm’s triage system, tickets flagged as billing disputes, regulatory inquiries, or contract amendments are held for human review before routing. The approval step is a single click in a Google Workspace notification, and the model’s confidence score is displayed so the reviewer can decide in under 30 seconds.

    ISO 27001 is the international standard for information security management systems. For the firm’s AI triage system, the standard requires that the data flow through the OpenAI API be documented in the risk assessment, that access to ticket content be logged, and that any PII in tickets be handled per the firm’s data protection policy. The system itself does not need certification, but the firm’s ISMS must account for the new processing path. A dedicated AI team typically maps the triage workflow to the relevant Annex A controls before go-live.

    Model-Agnostic Architecture means the triage layer calls the OpenAI API for classification and extraction, but the surrounding orchestration is built on standard APIs. If the firm later needs to move to an open-weight model on its own hardware for data residency reasons, the prompt templates and routing logic transfer without rewriting the integration layer. The dedicated AI team designs the abstraction so that swapping the model provider is a configuration change, not a re-architecture.

    OpenAI API is the hosted interface to OpenAI’s language models, used here for classification and extraction. The firm’s ticket content is sent over HTTPS, and the response is processed locally. No training data is retained by OpenAI under the standard API terms, but the firm should confirm the data processing agreement covers its specific use case. The API is chosen for its strong performance on English-language text and low latency, typically under 800 milliseconds for a classification call.

    P through T

    Process Audit is the first step in the engagement, where the AI team reviews the firm’s existing ticket workflow to identify which categories have the highest volume and the most inconsistent routing. The audit produces a one-page report listing the top three candidates for automation, with a projected time saving per ticket. In the firm’s case, the audit identified billing inquiries, project status requests, and contract amendments as the three highest-volume categories, with billing inquiries showing the most variance in routing decisions across different shifts.

    Scaling Across Departments means extending the triage logic from one department to others by parameterizing the classification rules per department. The marketing team’s tickets and the legal team’s tickets use different classification rules but the same underlying model and integration layer. The firm’s 501 to 2,000 employee size means there are typically four to six departments that generate tickets, and the rollout plan sequences them by volume so the highest-impact departments are automated first.

    Ticket Triage and Routing is the automated step where the AI model classifies the ticket’s intent, urgency, and department, then routes it to the correct queue or agent. The model does not draft the customer reply in the triage stage; it only determines where the ticket goes and what metadata to attach. This keeps the first-response SLA intact while freeing senior staff from the sorting step. In the firm’s workflow, the routing decision is written back to the helpdesk API, and a Google Workspace notification is generated if human review is required.

    4-Week Pilot is the fixed-scope engagement that delivers a working triage system on one ticket category. Week 1 covers the process audit and baseline measurement. Week 2 builds the extraction and classification pipeline against the OpenAI API. Week 3 runs the model in shadow mode on live tickets, comparing its routing decisions to human ones. Week 4 measures the before/after delta and documents the handoff to managed operation. The timeline assumes the firm’s helpdesk API is accessible and that a named business owner is available for daily check-ins.

  • 2-Week AI Automation Pilot Checklist for a 2,000+ Employee UK B2B SaaS Company

    1. Fix the pilot scope to one workflow before day one

    The pilot is scoped to one workflow, not three. Pick the highest-error-rate task in finance and accounting: contract review, invoice processing, or document extraction. The 2-week window is tight, so the scope must be fixed before day one. A 2,000+ employee B2B SaaS company typically has 40-60 back-office workflows, but the pilot touches only one. The process audit in week one identifies the target, measures the baseline, and defines the success criteria. Without a fixed scope, the pilot drifts into a discovery project and misses the 2-week deadline. The output is a single workflow with a documented before/after baseline on cycle time and error rate.

    2. Run the process audit and document the baseline

    Map every back-office workflow in finance and accounting. Measure cycle time in hours and error rate as a percentage of total transactions. For a 2,000+ employee B2B SaaS company, the audit typically covers invoice processing, document extraction, contract review, and data entry. Rank workflows by impact: error rate multiplied by transaction volume. The top-ranked workflow becomes the pilot target. Document the baseline in a one-page report: current cycle time, current error rate, number of transactions per month, and the team responsible. This baseline is the reference point for the before/after measurement at the end of the pilot. Without it, you cannot prove the AI layer delivered value.

    3. Configure the integration layer with existing CRMs, ERPs, and helpdesks

    The AI layer must plug into the systems the company already runs. For a B2B SaaS company, that means the CRM (Salesforce, HubSpot, or similar), the ERP (NetSuite, SAP, or Xero), the helpdesk (Zendesk, Freshdesk), and the documentation platform (Notion or Confluence). Use the native APIs, not screen scraping or manual exports. The integration layer is model-agnostic: the same API connectors work whether the underlying model is OpenAI, Anthropic, or an open-weight model on-premises. Configure the integration in week one, test it with sample data, and confirm that the AI can read from and write to each system. If an API is unavailable, flag it in the pilot report and adjust the scope.

    4. Build the pgvector embeddings pipeline over Notion or Confluence

    Ingest documentation from Notion or Confluence via their APIs. Generate embeddings for each document chunk and store them in pgvector, a PostgreSQL extension that handles vector similarity search natively. For a B2B SaaS company, the documentation includes product specs, SOPs, contract templates, and known-issue databases. The embeddings pipeline runs on a schedule: new or updated documents are re-embedded within 24 hours. When the AI queries the system, it retrieves the top-k most relevant passages and grounds the response in the company’s own documentation. This avoids hallucination and keeps the AI aligned with the latest internal docs. Test the retrieval quality with 20 sample queries before the pilot goes live.

    5. Set up human-in-the-loop approval for contract review and document extraction

    The AI model drafts, classifies, or extracts, but a person approves anything that touches money, health data, or a contract. For contract review in a B2B SaaS company, the AI flags clauses, extracts key terms, and drafts redlines, but a legal or finance professional signs off before the contract is sent. For document extraction, the AI pulls line items and tax codes from invoices, but a finance team member approves the final entry. The approval workflow is logged: who approved, when, and what was changed. This is the default delivery model, not an optional add-on. Configure the approval thresholds in week one: what confidence level triggers a human review, and what confidence level allows autonomous processing.

    6. Choose the model stack: API-based for quality, open-weight for data residency

    The model-agnostic architecture uses OpenAI or Anthropic APIs where quality matters, such as customer-facing AI assistants or complex contract analysis, and open-weight models on the client’s own hardware where regulated data cannot leave the building. For a UK-based B2B SaaS company with no specific compliance mandate, the default is API-based models for speed and quality. If data residency or IP protection becomes a concern, the architecture shifts to on-premises open-weight models without changing the integration layer. Document the model selection in the pilot report: which model handles which task, why, and what the fallback is if the primary model degrades. This keeps the architecture flexible as requirements evolve.

    7. Measure the before/after baseline and document the pilot results

    The pilot must ship with a measured before/after baseline on cycle time and error rate. At the end of week two, compare the pilot workflow’s performance against the baseline documented in the process audit. For contract review, measure cycle time in hours and error rate as a percentage of clauses flagged incorrectly. For document extraction, measure cycle time per invoice and error rate on extracted fields. The report includes: baseline metrics, pilot metrics, delta, and a recommendation for rollout. If the error rate dropped by 50% or more and cycle time improved by 30% or more, the pilot is a success. If not, document the gap and adjust the scope before scaling. This report is the input to the managed operations phase.

  • UK Medtech Firm Cuts Monthly HR Reporting from 14 Hours to 3 with On-Premise AI

    Background: A 1,200-Person UK Medtech Firm

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements in the UK healthcare and medtech sector. No named customer is represented; the details are aggregated and anonymised to preserve confidentiality while preserving the operational specifics that matter to a peer reader.

    The company in question is a mid-sized medtech firm with roughly 1,200 employees, headquartered in the West Midlands. It operates in the AI-Native Operations maturity band: leadership has already committed to embedding AI into core workflows, but the execution layer is still catching up. The existing stack includes a commercial HRIS, a CRM for partner and client records, and a document management system for regulatory filings. The firm holds ISO 27001 certification and is subject to UK GDPR, which constrains where and how employee and patient-adjacent data can be processed.

    The trigger for the engagement was straightforward. The monthly HR and recruiting report, which feeds into the board pack and the quarterly investor update, was taking the HR operations team approximately 14 hours to assemble by hand. The report pulled headcount data, time-to-fill metrics, offer acceptance rates, and attrition figures from three separate systems, then required a narrative summary that the HR director reviewed line by line. The process was error-prone, slow, and dependent on a single analyst who was also covering day-to-day recruiting operations.

    Challenge: 14 Hours of Manual Work and an ISO 27001 Audit Clock

    The operational pressure was not just the 14-hour cycle time. The HR director had flagged two compounding risks. First, the manual process had produced two material errors in the preceding six months: a misreported attrition figure that required a corrected board pack, and a time-to-fill metric that was off by a full week due to a date-format mismatch between the HRIS and the spreadsheet. Second, the firm was preparing for an ISO 27001 surveillance audit, and the manual reporting process, with its reliance on a single analyst and unversioned spreadsheets, was a known weakness in the information security management system documentation.

    The compliance constraint shaped the technical requirements from the outset. Employee data, including names, roles, and performance-adjacent metrics, could not be sent to a third-party cloud API. The firm’s data protection officer required that any AI processing of HR data occur on infrastructure within the company’s own network perimeter. This ruled out a simple SaaS chatbot or a cloud-hosted LLM API for the core reporting pipeline. The solution had to be a conversational agent and document-extraction layer running on open-weight models deployed on the client’s own GPU hardware, with a custom REST API and webhook integration to the existing HRIS, CRM, and document management system.

    The timeline was fixed at three months, driven by the board’s desire to see the new reporting process in place before the next quarterly cycle. That constraint meant the pilot had to be scoped tightly: one report type, one data source chain, one approval workflow.

    Approach: On-Premise Open-Weight Models and a Fixed-Scope Pilot

    Forfis began with a two-week process audit. We mapped the reporting workflow end-to-end: which data points came from which system, what transformations were applied manually, where the narrative summary was drafted, and who approved the final document. The audit identified four distinct sub-processes: data extraction from the HRIS, data extraction from the CRM, metric calculation and formatting, and narrative generation. Each was scored on volume, error rate, and regulatory sensitivity.

    The pilot was scoped to the data extraction and metric calculation sub-processes, plus a retrieval-augmented generation layer for the narrative summary. The architecture used open-weight models (Llama 3 70B for extraction, Mistral 7B for classification) running on the client’s own A100 GPU cluster. The integration layer was a custom REST API with webhooks: the HRIS pushed headcount and attrition data on a scheduled basis, the CRM pushed recruiting pipeline data, and the document management system received the final report via a webhook trigger. The conversational agent, accessible to the HR director and two senior HR managers, allowed them to query the underlying data in natural language and request specific report sections be regenerated.

    Human-in-the-loop approval was non-negotiable. The AI generated the draft report; the HR director reviewed and approved it before it was pushed to the board distribution list. Every approval was logged with a timestamp and user identifier, creating an audit trail that satisfied the ISO 27001 surveillance auditor. The pilot ran for one full reporting cycle, with the manual process running in parallel as a control.

    Outcome: Cycle Time Down to Under 3 Hours, Error Rate Down 70 Percent

    The pilot results were measured against the baseline established during the audit. Cycle time dropped from approximately 14 hours to under 3 hours: the automated pipeline completed data extraction and metric calculation in about 40 minutes, the narrative generation took roughly 15 minutes, and the remaining time was spent on human review and approval. The error rate, measured as the number of corrections required after the report was first drafted, fell by approximately 70 percent. The two types of errors that had occurred in the prior six months (date-format mismatch and misreported attrition) did not recur in the pilot cycle.

    The ISO 27001 surveillance audit, conducted in the final month of the engagement, noted the new reporting pipeline as a positive finding. The audit trail for AI-generated outputs, the on-premise data processing, and the defined approval workflow addressed the specific weakness the auditor had flagged in the previous cycle. The firm’s data protection officer confirmed that no regulated data had left the network perimeter during the pilot.

    Rollout extended the pipeline to cover the full monthly reporting suite, including the quarterly investor update. The managed operations contract began at the end of month three, covering model monitoring, integration health checks, and a defined escalation path for incidents. The HR operations team retained ownership of the business logic and approval workflow; Forfis handled the technical infrastructure and AI layer.

    Lessons for Similar Teams

    Five lessons from this engagement generalise to similar teams in regulated, mid-sized organisations:

    • Scope the pilot to one report type, not the whole reporting suite. The 3-month timeline was only achievable because the pilot covered a single data source chain and one approval workflow. Attempting to automate the full reporting suite in the same window would have stretched the team thin and delayed the baseline measurement.

    • The on-premise requirement is a design constraint, not an afterthought. Deciding early that regulated data could not leave the network perimeter shaped the model selection, the integration architecture, and the approval workflow. Teams that treat this as a compliance checkbox rather than an architectural decision tend to hit rework in weeks 4-6.

    • Human-in-the-loop approval is the audit trail. The ISO 27001 auditor did not care which model generated the report; they cared that a named human approved it, that the approval was timestamped, and that the log was immutable. Design the approval workflow to produce that log from day one.

    • Run the manual process in parallel for one full cycle. The pilot’s credibility depended on the side-by-side comparison. Without the manual control, the before/after baseline would have been anecdotal rather than measured.

    • The managed operations contract is where the real value lives. The pilot proves the concept; the managed contract keeps the pipeline accurate as the underlying data sources change, the models drift, and the business logic evolves. Budget for it from the start.

  • HIPAA-Compliant AI Invoice Processing for UK Healthcare: A Technical Deep Dive

    The Problem: Manual Invoice Processing in a Regulated Environment

    A 2,000+ employee healthcare organization in the UK processes 15,000 invoices monthly. Manual data entry takes 45 minutes per invoice, resulting in a 12-day average cycle time and a 3.2% error rate. The finance team spends 1,200 hours weekly on data entry, with 15% of time spent on error correction. The organization needs to reduce cycle time to under 48 hours and error rate to below 1% while maintaining HIPAA compliance. The challenge is not just automation but integration: the system must work with existing ERP (SAP S/4HANA), CRM (Salesforce), and helpdesk (Zendesk) without replacing them. The solution must handle complex invoice layouts, multi-currency transactions, and tax calculations while ensuring PHI never leaves the secure environment.

    The Mechanism: A Two-Stage Extraction Pipeline

    The pipeline uses a two-stage extraction. First, a vision-capable model (Claude 3.5 Sonnet) parses the PDF or image into structured JSON, identifying line items, totals, and vendor details. Second, a rule-based validation layer checks the JSON against the client’s chart of accounts and tax rules. If the confidence score drops below 0.85, the record is routed to a human reviewer. The system uses a hybrid approach: for high-volume, standardized invoices, a fine-tuned open-weight model (Llama 3 70B) runs on-premises. For complex, low-volume invoices, the system calls the Anthropic Claude API. The routing logic is based on invoice type, volume, and sensitivity. The pipeline exposes a /process-invoice endpoint that accepts PDFs and returns structured JSON. The ERP system calls this endpoint when a new invoice is uploaded. Conversely, the AI pipeline sends a webhook to the ERP when processing is complete, triggering automatic posting. For exceptions, the system sends a webhook to the client’s helpdesk, creating a ticket for human review.

    The Trade-offs: Accuracy, Cost, and Compliance

    The architect faces three key trade-offs. First, accuracy vs. cost: using the Claude API for all invoices costs $0.03 per invoice, while using an on-premises model costs $0.01 but requires $50,000 in hardware. The hybrid approach balances these costs. Second, compliance vs. flexibility: sending PHI to a third-party API violates HIPAA, but de-identifying data reduces accuracy. The solution is to send only financial metadata to the API, while patient identifiers remain in the client’s secure database. Third, speed vs. control: fully automated processing is faster but riskier. The human-in-the-loop approach adds 2-3 minutes per invoice but reduces error rates by 80%. The architect must also consider model drift: as invoice formats change, the model’s accuracy degrades. Retraining every 30 days mitigates this, but adds operational overhead. The managed operations model includes 24/7 monitoring, model retraining, and a dedicated support channel, covering these trade-offs.

    The Recommendation: A 3-Month Pilot with Managed Operations

    The pilot runs for 6-8 weeks. Week 1-2: process audit and data collection. Week 3-4: model fine-tuning and pipeline development. Week 5-6: parallel run (AI processes invoices alongside humans). Week 7-8: validation and go-live preparation. The 3-month timeline includes a 2-week buffer for stakeholder sign-off and integration testing with the ERP. The system tracks three key metrics: cycle time, error rate, and cost per invoice. Baselines are established during the process audit. During the pilot, the system compares AI performance against human performance. Post-implementation, the system monitors these metrics monthly and triggers retraining if error rates exceed 2% or cycle time increases by more than 10%. The managed operations model includes 24/7 monitoring, model retraining every 30 days, and a dedicated support channel. The client pays a monthly fee (typically 15-20% of the annual license cost) for ongoing optimization. This covers tracking model drift, updating validation rules, providing a monthly performance report, and handling API rate limits and cost optimization.