Tag: UK

  • Forfis AI Automation Audit: Cutting Error Rates in UK Medtech Back Offices

    1. Audit Before You Automate

    A 30-person UK medtech company processes 200 support tickets a week. Forty percent involve retrieving the same 12 clinical trial documents from Confluence. The median cycle time is 4.2 hours per ticket, and 11% require rework because the wrong document version was sent. The audit identifies this as the highest-impact workflow: high volume, repetitive, and error-prone. The fix is a RAG assistant over Confluence that retrieves the correct document version and drafts a response. A human approves anything touching patient data. The pilot runs for two weeks with a measured baseline. Cycle time drops to 1.8 hours. Error rate falls to 3%. The client now has a concrete ROI figure to justify rollout across the remaining 60% of tickets.

    2. Route PHI to On-Prem, Everything Else to Claude

    HIPAA requires that PHI never leaves the client’s controlled environment. Forfis runs open-weight models on the client’s own hardware for any workflow touching PHI, while using Anthropic Claude API for non-PHI tasks like ticket classification or document summarization where data can be de-identified. The architecture is model-agnostic by design. The same workflow routes PHI-sensitive calls to on-prem models and non-sensitive calls to the API. This keeps both speed and compliance intact. A 30-person medtech firm does not need to choose between a fast API and a compliant on-prem model. It uses both, in the same pipeline, with a routing layer that checks whether the input contains PHI before dispatching the call.

    3. Plug Into Confluence and the Helpdesk, Not Around Them

    The AI layer plugs into existing systems through their native APIs. A RAG assistant over Confluence reads from Confluence’s REST API. A ticket triage system writes classifications back to the helpdesk via its webhook. The client’s existing data model, access controls, and audit logs remain untouched. The AI layer is a thin, reversible addition rather than a platform migration. For a 30-person firm, this means no data migration, no retraining on a new tool, and no disruption to the existing workflow. The integration work takes 3 to 5 days per system, which fits inside the 4-week pilot timeline. The client keeps its Confluence, its helpdesk, and its CRM. The AI layer sits on top.

    4. Score Tickets Before a Human Reads Them

    Predictive scoring assigns a probability to each incoming ticket indicating likely resolution path, expected handling time, or risk of escalation. For a medtech company, this flags tickets mentioning adverse event language for immediate human review while routing routine dosage questions to a first-response agent. The scores are generated by the LLM and validated against historical ticket outcomes during the pilot. A human approves any action that touches patient data or contractual commitments. The model drafts the classification and the score. The person decides whether to act on it. This human-in-the-loop default is non-negotiable for any workflow touching money, health data, or a contract. It is the reason the pilot ships with a measured error rate baseline.

    5. Ship a Measured Baseline, Not a Demo

    The pilot ships with a measured before/after baseline on two metrics: cycle time and error rate. For a typical 30-person healthcare firm, Forfis has seen cycle time drop from 4.2 hours to 1.8 hours and error rate fall from 11% to 3% on document-heavy support workflows. These numbers are captured in a one-page report delivered at the end of week 4. The client gets a concrete ROI figure to justify rollout. The report also includes a list of edge cases the model handled poorly, which becomes the input for the next iteration. Without this baseline, the client cannot prove ROI or identify which workflow actually has the highest error rate. The audit and the measured pilot are the two things that separate a working deployment from a demo.

    6. Three Mistakes That Kill a 4-Week Pilot

    The most common failure is skipping the audit and jumping straight to a demo. Without a measured baseline, the client cannot prove ROI or identify which workflow actually has the highest error rate. The second pitfall is assuming a single model handles all tasks. A 30-person medtech firm might need Claude API for nuanced clinical document summarization but an open-weight model on-prem for PHI-tagged ticket routing. The third is underestimating integration work: connecting to Confluence, the helpdesk, and the CRM through their APIs takes real engineering time that a 4-week timeline must account for. The audit, the model routing, and the integration scope are the three things that determine whether a 4-week pilot delivers a measurable result or a slide deck.

  • Fixed-Scope AI Pilot vs. Full Rollout: A Fintech’s 6-Month Decision

    What Is Being Compared: Fixed-Scope Pilot vs. Full-Scale Rollout

    The two options under comparison are a fixed-scope pilot and a full-scale rollout of AI automation across a 2,000+ employee fintech firm in the UK. The pilot targets one workflow — in this case, monthly reporting compilation and internal knowledge search over Notion and Confluence — with a 6-week delivery window, a measured before/after baseline on cycle time and error rate, and a go/no-go decision at the end. The full-scale rollout deploys AI process automation across multiple departments simultaneously: invoice processing, ticket triage for round-the-clock customer response, HR and recruiting workflow orchestration, and a retrieval-augmented assistant over the company’s documentation. Both options use the same underlying architecture: n8n for workflow orchestration, a model-agnostic AI layer (OpenAI or Anthropic APIs for non-regulated data, open-weight models on the client’s hardware for PCI DSS-sensitive data), and human-in-the-loop approval for anything touching money, contracts, or health data. The difference is scope, timeline, and risk exposure.

    Criteria for the Comparison

    The following criteria determine which option fits a fintech firm’s constraints. PCI DSS compliance is the hard gate: any workflow that touches cardholder data must run on-premises or in a PCI-compliant enclave, which rules out cloud-only model APIs for those specific flows. Cycle time reduction is measured in hours per report or per ticket, not in vague efficiency gains. Error rate is tracked as a percentage of transactions requiring manual correction. Integration depth counts the number of existing systems (CRM, ERP, helpdesk, Notion, Confluence) that the automation must connect to without replacing them. Vendor lock-in is assessed by whether the architecture can swap models or orchestration tools without rework. Timeline is the calendar duration from kickoff to managed operation. Cost is the total engagement fee plus ongoing managed operation, expressed in GBP. Scalability is the number of additional workflows or departments that can be added without rebuilding the core architecture.

    Comparison Table

    Criterion Fixed-Scope Pilot Full-Scale Rollout
    PCI DSS compliance One workflow isolated; open-weight model on-premises for cardholder data Multiple workflows; requires a PCI-compliant enclave for all payment-related flows
    Cycle time reduction Measured on one workflow (e.g., monthly reporting: 14 hrs → 2 hrs) Measured across 4-6 workflows; aggregate reduction depends on each workflow’s baseline
    Error rate Baseline established in week 1; target <2% by week 6 Baselines established per department; target <3% aggregate by month 4
    Integration depth 2-3 systems (Notion, Confluence, one CRM) 6-10 systems (CRM, ERP, helpdesk, Notion, Confluence, HRIS, payment gateway)
    Vendor lock-in Low; n8n workflows are portable; model can be swapped Moderate; more integrations increase switching cost, but n8n remains the orchestration layer
    Timeline 6 weeks to pilot completion; 2 weeks to decision 6 months to full managed operation across departments
    Cost (GBP) £18,000–£35,000 for the pilot £120,000–£250,000 for the full engagement plus £4,000–£8,000/month managed operation
    Scalability One workflow; scaling requires a new pilot per department Multi-department from day one; new workflows added to the existing n8n architecture

    When the Fixed-Scope Pilot Wins

    The fixed-scope pilot wins when the firm has not yet established a baseline for AI automation and needs to prove value before committing to a multi-department rollout. For a 2,000+ employee fintech in the UK, the pilot on monthly reporting and internal knowledge search over Notion and Confluence delivers a measurable result in 6 weeks: cycle time drops from 14 hours to 2 hours per report, and the error rate on data extraction falls from 8% to under 2%. The go/no-go decision is based on these numbers, not on a qualitative assessment. The pilot also validates the n8n orchestration layer and the human-in-the-loop approval gates without exposing the entire back office to change. If the pilot meets its targets, the firm has a proven template for the next workflow.

    The full-scale rollout wins when the firm has already completed a process audit, has identified 4-6 high-impact workflows, and has the IT capacity to manage parallel integrations. For a fintech with PCI DSS obligations, the rollout must include an on-premises open-weight model for any workflow that touches cardholder data, while non-regulated workflows (ticket triage, HR recruiting, knowledge search) can use OpenAI or Anthropic APIs. The 6-month timeline assumes that the process audit is complete, that the n8n environment is provisioned, and that each department has a named owner for the integration work. The rollout delivers aggregate cycle time reduction across the firm, but it requires a managed operation team from month 3 onward to handle model updates, integration drift, and new workflow requests.

    When the Full-Scale Rollout Wins

    The full-scale rollout is the right choice when the firm’s process audit has already identified multiple workflows with high impact and low integration complexity, and when the IT team can support parallel workstreams. For a 2,000+ employee fintech in the UK, this means the audit has scored invoice processing, ticket triage, HR and recruiting workflow orchestration, and internal knowledge search as the top four candidates. The rollout deploys all four within 6 months, with the PCI DSS-sensitive workflows (invoice processing, payment-related ticket triage) running on open-weight models on the client’s hardware, and the non-regulated workflows (HR recruiting, knowledge search) using OpenAI or Anthropic APIs. The n8n orchestration layer is shared across all workflows, so a change to one integration (e.g., a CRM API update) is applied once, not four times. The managed operation team, staffed from month 3, handles model retraining, integration monitoring, and new workflow requests. The cost is higher — £120,000 to £250,000 for the engagement plus £4,000 to £8,000 per month for managed operation — but the aggregate cycle time reduction across four workflows justifies the investment within 12 months for a firm of this size.

    Recommendation for the Scenario

    For a 2,000+ employee fintech in the UK with PCI DSS obligations, the recommendation is a fixed-scope pilot first, followed by a phased rollout. The pilot targets monthly reporting compilation and internal knowledge search over Notion and Confluence, with a 6-week delivery window and a measured baseline on cycle time and error rate. The pilot validates the n8n orchestration layer, the human-in-the-loop approval gates, and the model-agnostic architecture without exposing the payment processing workflows to change. If the pilot meets its targets — cycle time reduced from 14 hours to under 3 hours, error rate below 2% — the firm proceeds to a phased rollout over the remaining 4 months of the 6-month timeline. The rollout adds invoice processing, ticket triage for round-the-clock customer response, and HR and recruiting workflow orchestration, with PCI DSS-sensitive workflows running on open-weight models on the client’s hardware. The total engagement cost is £150,000 to £280,000, with managed operation at £5,000 to £8,000 per month from month 4 onward. This approach limits risk, delivers a measurable result in 6 weeks, and scales the architecture across departments without rebuilding it.

  • AI Contract Review Glossary: UK Healthcare, ISO 27001, and Managed Operations

    AI Process Audit

    The AI process audit is the foundational step that determines which workflows are worth automating. For a 51-200 person healthcare company in the UK, the audit maps the contract review process, measures baseline cycle time (e.g., 12 hours per contract) and error rate (e.g., 8% missed clauses), and selects the highest-impact workflow for a 4-week pilot. This ensures the AI investment targets a measurable bottleneck rather than a low-value task. The audit also identifies integration points with existing systems, such as the document management system and CRM, to ensure the AI assistant plugs into the company’s current infrastructure rather than replacing it. By grounding the pilot in concrete metrics, the audit provides a clear baseline against which the AI’s performance can be measured, which is critical for demonstrating ROI to stakeholders and ensuring the project aligns with the company’s ISO 27001 compliance requirements.

    Anthropic Claude API

    The Anthropic Claude API is a large language model service that Forfis uses for high-quality text generation and classification tasks. In the contract review scenario, Claude handles the semantic analysis of legal clauses and drafting of redlines. Because the data is sensitive, the API calls are routed through a custom REST gateway that enforces ISO 27001 logging and access controls, ensuring that no raw contract data is stored on Anthropic’s servers beyond the inference window. The model-agnostic architecture allows Forfis to switch to open-weight models on the client’s own hardware if the data cannot leave the building, but for most contract review tasks, the Claude API provides the best balance of quality and cost. The API’s context window of 200,000 tokens allows the model to process entire contracts in a single pass, which is critical for maintaining context across complex legal documents.

    Retrieval-Augmented Knowledge Assistant

    A retrieval-augmented knowledge assistant retrieves relevant passages from a company’s internal documents—contracts, SOPs, CRM records—and uses them to ground an LLM’s response. In a UK healthcare contract review, the assistant pulls the specific liability clause from a 2023 supplier agreement and flags it against the current ISO 27001 Annex A.8.25 requirements, reducing manual search time from 45 minutes to under 3 minutes per clause. The retrieval index is built from the company’s document management system and updated weekly to include new contracts and policy changes. This approach ensures that the AI’s responses are grounded in the company’s actual data rather than general knowledge, which is critical for legal and compliance tasks where accuracy is paramount. The assistant also logs every retrieval and response, providing an audit trail that satisfies ISO 27001 A.8.15 logging requirements.

    ISO 27001

    ISO 27001 is an international standard for information security management systems. For a 51-200 person UK healthcare company, it mandates risk-based controls for data handling, access, and incident response. When deploying an AI contract assistant, the company must ensure the model’s data pipeline complies with Annex A.8.25 (secure development) and A.8.15 (logging), which is why the architecture uses custom REST APIs to keep PHI and contract data within the client’s VPC rather than sending it to a third-party SaaS. The standard also requires that the company maintains a risk assessment that includes the AI system, which means the AI’s data flow, access controls, and incident response procedures must be documented and reviewed annually. For a company in the healthcare sector, ISO 27001 compliance is not optional—it is a prerequisite for many contracts with NHS trusts and private healthcare providers, making it a critical consideration in the AI deployment strategy.

    Custom REST API and Webhooks

    A custom REST API and webhooks integration allows the AI assistant to pull contract data from the company’s existing document management system and push reviewed drafts back to the legal team’s workflow. Webhooks trigger the AI review when a new contract is uploaded, and the REST API returns the annotated PDF and a JSON summary of flagged clauses. This avoids replacing the existing DMS and keeps the integration within the company’s ISO 27001 scope. The API is designed to be idempotent, meaning that if a webhook is retried, the AI review is not duplicated, which is critical for maintaining data integrity. The integration also includes rate limiting and authentication to ensure that the AI system is not abused or overwhelmed by a sudden spike in contract uploads. By using the company’s existing APIs rather than building a new system, the integration reduces the risk of data loss and ensures that the AI assistant fits seamlessly into the company’s current workflow.

    Managed AI Operations

    Managed AI operations is a delivery model where the vendor handles ongoing monitoring, model updates, and performance tuning after the initial pilot. For a healthcare company, this means Forfis tracks the contract assistant’s accuracy weekly, adjusts the retrieval index when new contract templates are added, and ensures the system remains compliant with ISO 27001 as the company’s security posture evolves. This removes the need for the client to hire a dedicated AI engineer, which is critical for a 51-200 person company that may not have the budget or expertise to maintain an AI system in-house. The managed operations contract includes a service level agreement (SLA) that specifies the maximum downtime (e.g., 4 hours per month) and the response time for critical issues (e.g., 2 hours). By outsourcing the ongoing maintenance, the company can focus on its core business while ensuring that the AI system continues to deliver value and remain compliant.

    Human-in-the-Loop

    Human-in-the-loop (HITL) is a design pattern where the AI drafts or classifies, but a human approves any action that touches money, health data, or contracts. In the contract review scenario, the AI flags clauses and suggests redlines, but a legal reviewer must approve the final version before it is sent to the counterparty. This ensures that the AI’s output is auditable and that the company retains legal accountability, which is critical for ISO 27001 compliance. The HITL workflow is designed to minimize the time the human spends on the task—the AI pre-filters the contract and highlights only the clauses that require attention, reducing the reviewer’s workload from 12 hours to 3 hours per contract. The system also logs every human decision, providing an audit trail that can be used for compliance reporting and continuous improvement. By keeping the human in the loop, the company ensures that the AI is a tool that augments human expertise rather than replacing it, which is essential for maintaining trust and accountability in a regulated industry.

  • AI Candidate Screening and HR Reporting for a UK Insurance Firm: A 3-Month Pilot

    The Problem: Scaling HR Operations Without New Hires

    A 2,000+ employee insurance firm in the UK faces a familiar constraint: HR and recruiting teams are stretched thin, and the volume of candidate applications and monthly reporting cycles keeps growing without a corresponding increase in headcount. The firm needs to process more applications, produce more reports, and maintain compliance with GDPR Article 22 on automated decision-making, all within a 3-month window. The solution is not a new HR platform or a full AI transformation. It is a fixed-scope pilot that automates one or two specific workflows, measures the impact, and establishes a foundation for scaling across departments. The pilot targets candidate screening and monthly reporting, using a retrieval-augmented knowledge assistant that reads from the firm’s existing Confluence or Notion workspace. The architecture is model-agnostic: OpenAI or Anthropic APIs for tasks where output quality matters, and open-weight models on the firm’s own hardware for any data that cannot leave the building. The pilot ships with a measured before/after baseline on cycle time and error rate, so the business case is quantified, not assumed.

    Pilot Scope: Candidate Screening and Monthly Reporting

    The pilot begins with a process audit that maps the current candidate screening workflow end to end. The team identifies where manual effort concentrates: parsing application PDFs, matching candidates against job descriptions, flagging compliance issues, and drafting initial feedback. The same audit covers the monthly reporting cycle, which typically involves pulling data from the HR system, formatting it into a template, and writing narrative summaries. The data sources are the firm’s existing Confluence or Notion workspace, which holds job descriptions, screening criteria, and reporting templates. The assistant connects to these platforms through their public APIs, so the HR team continues to maintain content where it already lives. The architecture uses pgvector for embeddings search, storing vector representations of the source documents in a PostgreSQL instance on the firm’s own infrastructure. This keeps the data within the firm’s control, which matters for an insurance company handling regulated data. The model layer is deliberately model-agnostic: the pilot uses OpenAI or Anthropic APIs for drafting and classification tasks, and open-weight models on the firm’s hardware for any step that touches sensitive candidate data.

    Human-in-the-Loop and GDPR Compliance

    The assistant does not make final decisions on candidates. It classifies applications against the screening criteria stored in Confluence, ranks them, and drafts a summary for the recruiter to review. A human recruiter approves or overrides every screening decision before it reaches the candidate. This human-in-the-loop design satisfies GDPR Article 22, which requires human involvement in automated decisions with legal or similarly significant effects. The same principle applies to monthly reporting: the assistant assembles the data, formats the report, and drafts the narrative sections, but a human analyst reviews and approves the final document before distribution. Every pilot ships with a measured before/after baseline. The baseline captures cycle time, the time from application receipt to screening decision, and error rate, the percentage of screening decisions that a human reviewer would overturn. The baseline is measured during the first two weeks of the pilot, before the AI is fully active, so the comparison is direct. The firm gets a quantified picture of the impact, not a qualitative impression.

    3-Month Timeline and Delivery Phases

    The 3-month timeline breaks into three phases. Weeks 1 to 4 cover the process audit and data mapping: the team interviews HR and recruiting staff, maps the current workflow, identifies the data sources in Confluence or Notion, and defines the success metrics. Weeks 5 to 8 are development and integration: the team builds the retrieval-augmented assistant, connects it to the HR system and the documentation platform, and configures the model layer. Weeks 9 to 12 are user testing and measurement: the HR team uses the assistant in a live environment, the team captures the before/after baseline, and the firm makes a go/no-go decision on broader rollout. The pilot covers one or two workflows, not the entire HR function. The output is a working system, a measured baseline, and a clear picture of what scaling across departments would look like. The architecture is designed so that the next department, whether it is claims processing or customer service, plugs into the same stack without rebuilding from scratch.

    Scaling Across Departments After the Pilot

    The pilot is not the end of the engagement. It is the first step in scaling AI across departments. The architecture established in the pilot, the model-agnostic layer, the human-approval workflow, the pgvector embeddings search, and the measurement framework, is reusable. When the firm decides to extend the assistant to claims processing or customer service, the team reuses the same integration patterns and the same compliance controls. The marginal cost and time for each new use case is lower than the initial pilot because the foundational work is already done. The firm also gets a managed operation model: the team monitors the assistant, handles model updates, and maintains the integration with the HR system and documentation platform. This is not a one-off project; it is a managed service that scales with the firm’s needs. The 3-month pilot gives the firm a quantified business case, a working system, and a clear path to scaling without new hires.

  • UK Medtech Cuts Invoice First-Response Time to 6 Hours with On-Premise AI

    Background: A UK Medtech Distributor at 1,200 Headcount

    This case study is a composite drawn from patterns Forfis has observed across multiple engagements. We do not name real clients. The company described here is a UK-based medtech distributor with roughly 1,200 employees, operating in the 501-2000 band. It handles procurement, supply-chain coordination, and customer-facing service for hospital and clinic clients across the UK and Ireland. The existing stack includes a mid-market ERP, a CRM for customer records, and Microsoft Teams as the primary internal messaging channel. The finance and operations teams were running on a mix of spreadsheets, email threads, and a legacy invoice portal that had not been updated since 2019. The company had no dedicated AI team and had not previously deployed any machine-learning system in production.

    Challenge: 48-Hour Invoice Response, Zero New Hires, GDPR in the Loop

    The trigger was a 40 percent increase in supplier invoice volume over eighteen months, driven by a new product line and expanded distribution contracts. The finance team of eleven was processing invoices manually: extracting line items, matching them against purchase orders, flagging discrepancies, and posting to the ERP. Average first-response time to a supplier query about a disputed invoice was 48 hours. The operations director had a hard constraint: no new headcount in the current fiscal year, and GDPR compliance was non-negotiable because invoice metadata occasionally contained patient-identifiable information from hospital procurement orders. The deadline was six months to show a measurable reduction in cycle time before the next board review. The team needed to cut first-response time without adding a single FTE and without sending regulated data to a third-party cloud API.

    Approach: On-Premise Open-Weight Models, Predictive Scoring, and a Fixed-Scope Pilot

    Forfis ran a two-week process audit across the finance and operations workflows. The audit identified invoice processing as the highest-impact target: high volume, repetitive extraction, and a clear before/after metric. The pilot scope was fixed: one invoice category (supplier purchase orders with line-item extraction), one integration point (the existing ERP API), and one notification channel (Microsoft Teams). The architecture used open-weight models on the client’s own hardware, so no regulated data left the building. A retrieval-augmented layer pulled context from the client’s own procurement documentation and CRM records to improve extraction accuracy. Predictive scoring assigned a confidence value to each extracted field; items above 95 percent auto-posted, items below routed to a human reviewer in Teams. The dedicated AI team of four engineers and one product designer worked on-site for the first four weeks, then shifted to remote with weekly syncs. The pilot ran for eight weeks with a measured baseline captured in week one.

    Outcome: 48 Hours to Under 6, Error Rate Below 2 Percent

    The pilot cleared its threshold. Average first-response time for supplier invoice queries dropped from 48 hours to under 6 hours. Extraction error rate on line items fell from 11 percent to under 2 percent. The finance team’s manual review volume dropped by roughly 60 percent, because the predictive scoring layer auto-approved the high-confidence items. The remaining 40 percent of invoices still required human eyes, but the reviewers now worked from a pre-drafted, context-enriched queue in Teams rather than a blank spreadsheet. The ERP integration held: no data left the client’s infrastructure, and the GDPR data-processing record was updated to reflect the on-premise model deployment. The operations director reported that the team absorbed the 40 percent invoice volume increase without a single new hire. The six-month timeline was met, and the board review proceeded on the strength of the measured baseline.

    Lessons for Similar Teams

    • Baseline first, always. The pilot did not start until the team had a measured before/after baseline on cycle time and error rate. Without that number, the board review would have been a conversation about impressions rather than data. Every similar team should capture the baseline in week one, not after the pilot ends.
    • Model-agnostic architecture pays off. The client started with open-weight models on-premise for GDPR reasons. If a future use case requires a frontier API for a non-regulated workflow, the integration layer does not need to be rebuilt. Teams that hard-code a single vendor API into their architecture will face this problem.
    • Predictive scoring is the human-in-the-loop mechanism. The confidence threshold is not a suggestion; it is the architectural gate. Items above 95 percent auto-approve, items below route to a human. This is what makes GDPR Article 22 compliance operational rather than theoretical.
    • Integration through existing APIs, not replacement. The ERP, CRM, and Teams stack stayed intact. The AI layer sat on top. For a 1,200-person operation, a rip-and-replace project would have taken two years and a budget the company did not have.
    • Dedicated team beats rotating contractors. The four engineers and one product designer stayed on the engagement from audit through rollout. Consistency in the team meant the client’s internal stakeholders had a single point of contact and a shared context that did not reset every sprint.
  • Ticket Triage Automation for UK E-commerce: A 3-Month Fixed-Scope Pilot

    The Problem: Senior Staff Buried in Routine Ticket Triage

    You run a 51-200 person e-commerce operation in the UK. Your support team handles 800 to 1,500 tickets per day across order status, delivery issues, returns, and product questions. Senior staff spend 40-60% of their time on routine triage: reading the ticket, classifying it, routing it to the right queue, and drafting a first response. This work is repetitive, error-prone, and it pulls your most experienced people away from the complex cases that actually need their judgment. The goal is not to replace your support team; it is to free senior staff from routine work so they can focus on escalations, customer retention, and process improvement. The constraint is GDPR: ticket data contains customer names, order numbers, and delivery addresses, so any automation must comply with UK GDPR and the Data Protection Act 2018. The delivery model is a fixed-scope pilot: one ticket category, one helpdesk, one ERP integration, 3 months, measured before/after baselines on cycle time and error rate.

    Prerequisites: What You Need Before Step 1

    Before you write a single line of integration code, you need five things in place. First, a documented list of your top 20 ticket categories with their current routing rules, SLA targets, and escalation paths. This list is your ground truth; without it, the model has no reference for what ‘correct’ routing looks like. Second, API access to your helpdesk (Zendesk, Freshdesk, or similar) and your SAP or Microsoft Dynamics ERP instance. You need read access to order data and write access to ticket status fields. Third, a named GDPR Data Protection Officer or privacy lead who can sign off on the Data Protection Impact Assessment (DPIA). Fourth, a fixed-scope pilot agreement that defines success metrics (cycle time reduction, error rate, cost per ticket), data handling boundaries, and a 3-month timeline. Fifth, a human-in-the-loop approval workflow in your helpdesk UI where agents can accept, edit, or reject the model’s routing suggestion. If any of these are missing, the pilot will stall in week 2 or 3, and you will not have the measured baselines needed to justify scaling.

    Step 1: Audit the Ticket Flow and Define the Baseline

    Run a 2-week process audit on your top 3 ticket categories. Export 500 historical tickets from your helpdesk, tag each one with its final routing destination, cycle time, and error rate (did it go to the wrong queue, get escalated unnecessarily, or take longer than the SLA?). This gives you a baseline: for example, ‘order status’ tickets average 14 minutes from receipt to first response, with a 7% error rate. The audit also reveals which categories are worth automating. If a category has a 90%+ routing accuracy already, the ROI on automation is low. If it has a 40% error rate and a 22-minute cycle time, it is a strong candidate. The output of this step is a one-page brief per category: current metrics, routing rules, and the target metrics for the pilot. This brief becomes the acceptance criteria for the fixed-scope pilot agreement.

    Step 2: Complete the GDPR DPIA and Data Processing Agreement

    Complete a Data Protection Impact Assessment (DPIA) before any ticket data flows through the OpenAI API. The DPIA must document the lawful basis for processing (typically legitimate interest under GDPR Article 6(1)(f)), the categories of personal data involved (names, order numbers, delivery addresses), the retention policy (delete or anonymise ticket payloads after the routing decision is logged), and the security measures (encryption in transit via TLS 1.3, access controls on the API keys). You must also ensure the OpenAI API is covered by a Data Processing Agreement (DPA) with UK Standard Contractual Clauses. If tickets contain health data (e.g., a customer reporting a product caused an injury), Article 9 applies and you need explicit consent or another specific exception. The DPIA is not a one-time document; it must be updated if you change the model, the data flow, or the retention policy. Your DPO signs off on the DPIA before the pilot goes live.

    Step 3: Build the Model-Agnostic Triage Layer

    Build the triage layer as a model-agnostic abstraction. The integration layer calls your helpdesk’s REST API to fetch new tickets, and your SAP or Dynamics ERP’s OData or SOAP endpoints to enrich the ticket with order data (order status, delivery ETA, return eligibility). The enriched ticket payload is sent to the model endpoint, which returns a classification (category, priority, routing destination) and a suggested first-response template. For the pilot, use OpenAI’s GPT-4o API because it requires no GPU infrastructure and provides high-accuracy classification. The model-agnostic design means the routing logic is decoupled from the model: if GDPR or client contracts later demand on-prem inference, you can swap the model endpoint to a locally hosted open-weight model (e.g., Llama 3 70B) without changing the integration layer. The output is written back to the helpdesk via the API, with the model’s confidence score logged for audit.

    Step 4: Integrate with Helpdesk and ERP via API

    Wire the triage layer into your helpdesk and ERP. The helpdesk integration uses the REST API to create a new ticket, update its status, and log the model’s routing decision. The ERP integration uses OData (for Dynamics) or the SAP Business Technology Platform API to fetch order data and update the ticket with order-specific context. The human-in-the-loop approval workflow is critical: the model’s output appears in the helpdesk UI as a suggestion, and a human agent must accept, edit, or reject it before the ticket is routed. For tickets touching money (refunds, chargebacks), health data, or contract terms, the human approval is mandatory and the model’s output is treated as a suggestion only. The approval log feeds back into the model’s prompt engineering in the next sprint, so the system improves over time. This is not a limitation; it is the compliance mechanism that keeps the system within GDPR and internal audit boundaries.

    Step 5: Run the 8-Week Pilot in Three Phases

    Run the pilot in three phases. Phase 1 (weeks 1-2): shadow mode. The model classifies and routes, but a human approves every action. You measure the model’s accuracy against the human-approved outcomes. Phase 2 (weeks 3-6): semi-automated mode. The model handles low-risk categories (e.g., ‘where is my order’, ‘change delivery address’) and escalates the rest to a human. You measure cycle time and error rate for the automated categories. Phase 3 (weeks 7-8): full automation for approved categories with a 5% random sample still routed to a human for quality checks. The pilot ends with a measured before/after report: cycle time reduction (e.g., from 14 minutes to 3 minutes), error rate (e.g., from 7% to 2%), and cost per ticket (e.g., from £4.20 to £2.10). This report becomes the business case for scaling to other departments and ticket categories.

  • n8n Ticket Triage and Monthly Reporting for a 20-Person B2B SaaS Team in the UK

    The Problem: Manual Triage and Reporting at 20 People

    A 20-person B2B SaaS company in the UK runs its support operation on a single helpdesk, a CRM, and a Slack channel where engineers and support agents triage tickets by hand. The operations lead spends four to six hours every month pulling ticket volume, resolution times, and CSAT scores from three systems and formatting a report for the board. Support agents classify and route every incoming ticket manually, and the median first-response time sits at 4.2 hours. The company has no compliance mandate—no GDPR data residency requirement beyond standard UK law, no sector-specific regulation—but it has a hard constraint: it cannot hire another support agent this quarter. The problem is not a lack of tools. The helpdesk and CRM are fine. The problem is that the workflow between them is manual, and the manual steps do not scale with the ticket volume that a 20-person SaaS company generates as it grows from 50 to 200 customers. The fix is not a new platform. It is an orchestration layer that sits on top of the existing systems and automates the classification, routing, and reporting steps that currently consume human hours.

    The Mechanism: n8n Orchestration Over Existing REST and Webhook Surfaces

    The architecture is a single n8n instance running on the client’s own infrastructure, connected to the helpdesk and CRM through their native REST APIs and webhook events. The ticket triage workflow has five nodes. First, a Webhook node receives a ticket.created event from the helpdesk. Second, an HTTP Request node calls the helpdesk’s REST API to fetch the ticket’s subject, body, customer tier, and SLA class. Third, a second HTTP Request node calls the CRM’s REST API to enrich the ticket with account data: annual contract value, support tier, and open cases. Fourth, an AI Agent node calls an LLM API—OpenAI’s GPT-4o or Anthropic’s Claude, depending on which the client’s prompt engineering tests produce the higher classification accuracy on a labeled sample of 200 historical tickets. The prompt includes the ticket text, the account enrichment, and a classification schema with four intent categories (billing, technical, onboarding, escalation) and three urgency levels. Fifth, an IF node checks the model’s confidence score. If confidence is above 0.85, the workflow calls the helpdesk’s REST API to assign the ticket to the correct queue and set the priority. If confidence is below 0.85, the workflow creates an approval task in the helpdesk for a human agent. The agent reviews the AI’s proposed classification, approves or corrects it, and the workflow resumes. The monthly reporting workflow is a separate n8n flow on a cron schedule: it queries the helpdesk and CRM REST APIs for the month’s metrics, assembles a structured report, and delivers it via a Slack webhook or email. No custom middleware. No new database. The n8n instance logs every execution with input, output, duration, and error state, which serves as the audit trail for the human-in-the-loop step and the before/after baseline.

    Trade-offs: Model Choice, Confidence Thresholds, and Fixed Scope

    The first trade-off is model choice. A commercial API like GPT-4o or Claude produces higher classification accuracy on out-of-the-box prompts, but every ticket body and customer name is sent to a third-party endpoint. For a B2B SaaS company with no data residency mandate, this is acceptable. If the company later serves a healthcare or financial-services vertical, the same n8n workflow re-points the AI Agent node to an open-weight model served via Ollama or vLLM on the client’s own hardware. The surrounding orchestration logic—webhook, HTTP Request, IF, approval step—does not change. Only the model endpoint URL and authentication change. The second trade-off is the confidence threshold. Setting it at 0.85 means roughly 10-15% of tickets hit the human approval step in the first month. Lowering it to 0.75 reduces the approval volume to under 5% but increases the misrouting rate. The threshold is not a fixed constant; it is tuned during the parallel run in week 7, where the AI triage runs alongside human triage and both results are logged. The third trade-off is the fixed scope. The pilot covers ticket triage and monthly reporting only. If the audit reveals that invoice processing or document extraction are also candidates, those are separate pilots. The fixed scope is what makes the 8-week timeline credible. Without it, the pilot becomes a platform rebuild and the timeline slips to 16 weeks or more.

    Recommendation: The 8-Week Fixed-Scope Pilot

    The pilot runs on an 8-week timeline with a defined acceptance gate. Weeks 1-2 are the process audit: map every step from ticket creation to resolution, measure cycle time and error rate over a 2-week window, identify the integration surface (which helpdesk, which CRM, what APIs, what webhook events), and produce a one-page scope document. Weeks 3-4 are the n8n build: webhook and HTTP Request nodes for the helpdesk and CRM, the AI Agent node with prompt engineering against a labeled sample of 200 historical tickets, and the IF node with the confidence threshold. Week 5 is the human-in-the-loop approval step and edge-case handling: what happens when the AI Agent returns a classification outside the four intent categories, when the CRM enrichment call times out, when the helpdesk webhook is delayed. Week 6 is the monthly reporting workflow: cron schedule, REST API queries, report template, delivery via Slack webhook. Week 7 is the parallel run: the AI triage runs alongside human triage, both results are logged, and the confidence threshold is tuned. Week 8 is the acceptance gate: the before/after metrics are measured over the same 2-week window as the baseline. The acceptance criteria are: median first-response time reduced by at least 50%, misrouting rate reduced by at least 50 percentage points, and the monthly report generated without manual intervention. The handover includes the n8n workflow export, the prompt engineering documentation, the integration credentials, and a runbook for the operations lead. The company scales its support operation without a new hire. The operations lead gets the monthly report in under 90 seconds instead of four hours. The support agents handle 22% more tickets per day because the classification and routing steps that consumed 40 minutes per agent per hour are now automated.

  • 3-Month AI Ticket Triage Pilot for a UK Fintech: Claude API, Zendesk, GDPR

    The Problem: Misrouted Tickets and Slow First Response in a UK Fintech

    You run a 2,000+ employee fintech in the UK. Your support team handles 50,000+ tickets per month across English, German, and French. First-response time averages 4.2 hours, and 18% of tickets are misrouted to the wrong queue. You need round-the-clock coverage without hiring 200 more agents. The constraint: GDPR Article 22 requires human oversight for automated decisions, and payment data cannot leave your infrastructure without a Transfer Impact Assessment. You are at the “Running Isolated Pilots” maturity stage: you have tested AI in one workflow but have not systematized it. This guide walks you through a 3-month pilot that deploys predictive scoring for ticket triage using Anthropic Claude API, integrated with your existing Zendesk or Intercom instance, delivered by a dedicated AI team.

    Prerequisites: What You Need Before Step 1

    Before you start, confirm these items are in place:

    • Zendesk or Intercom enterprise plan with API access enabled. Verify your API rate limit (100 requests/second for Zendesk enterprise, 50 for Intercom) and webhook endpoint configuration.
    • 6–12 months of historical ticket data exported from your helpdesk. Each record must include: ticket ID, subject, body, category, resolution time, agent ID, customer segment, and language.
    • GDPR Article 30 record of processing activities updated to include AI-assisted triage. Document the data flows, legal basis (legitimate interest or consent), and retention policy.
    • Anthropic Claude API account with billing set up. Confirm you have executed a Standard Contractual Clause (SCC) with Anthropic and completed a Transfer Impact Assessment for UK GDPR compliance.
    • Dedicated AI team of four to six people: one ML engineer, one integration engineer, one product manager, and one data engineer. For multilingual coverage, add a language specialist or localization partner.
    • Baseline metrics measured from your historical data: average first-response time, resolution time, misrouting rate, and ticket volume per category per language.

    Step 1: Extract and Clean Historical Ticket Data

    Export 6–12 months of tickets from Zendesk or Intercom using the REST API. For Zendesk, use the /api/v2/tickets.json endpoint with pagination (100 tickets per page). For Intercom, use the /api/contacts and /api/conversations endpoints. Store the raw data in your data warehouse (Snowflake, BigQuery, or Redshift). Pseudonymize PII per GDPR Article 25: replace customer names with UUIDs, mask card numbers, and hash email addresses. Build a cleaned dataset with columns: ticket_id, subject, body, category, resolution_time_hours, agent_id, customer_segment, language, timestamp. This dataset becomes your training and evaluation set for the predictive scoring model.

    Step 2: Measure the Baseline: Cycle Time and Misrouting Rate

    Calculate your baseline from the cleaned dataset. For each ticket category and language, compute: average first-response time (hours), average resolution time (hours), misrouting rate (percentage of tickets reassigned by a human agent within 24 hours), and ticket volume per month. Store these metrics in a dashboard (Grafana, Looker, or Tableau) with a “pre-pilot” label. This baseline is your before/after reference. For example, if your English “billing inquiries” category has a 4.2-hour average first-response time and an 18% misrouting rate, your pilot success criteria might be: reduce first-response time to 2.5 hours and misrouting rate to 10% within 8 weeks. Document these targets in a one-page pilot charter signed by your support director and CTO.

    Step 3: Define Ticket Categories and Routing Rules

    Define your ticket categories and routing rules. For a fintech, typical categories include: “billing dispute”, “onboarding question”, “security concern”, “transaction inquiry”, and “account closure”. For each category, specify: the target queue, the required agent skill set, and the SLA (e.g., “security concern” routes to the fraud team with a 1-hour SLA). Build a routing matrix in a JSON file: {"category": "billing dispute", "queue": "billing", "sla_hours": 4, "human_review": true}. The human_review flag is critical for GDPR Article 22: any category involving money movement, account closure, or security must require human approval before action. This matrix becomes the logic your AI scoring model will follow.

    Step 4: Build the Predictive Scoring Model with Claude API

    Build the scoring pipeline using Anthropic Claude API. For each incoming ticket, send the ticket body, subject, and customer history to Claude with a system prompt that defines your categories and routing rules. Example system prompt: “You are a ticket triage assistant for a UK fintech. Classify the ticket into one of: billing dispute, onboarding question, security concern, transaction inquiry, account closure. Return a JSON object with ‘category’, ‘confidence_score’ (0.0–1.0), and ‘reasoning’.” Use the claude-3-5-sonnet model for balanced cost and accuracy. Set the temperature to 0.1 for deterministic outputs. Log every request: ticket ID, input tokens, output tokens, model version, timestamp, and output score. Store logs in your data warehouse with a 12-month retention policy.

    Step 5: Integrate with Zendesk or Intercom via Webhooks

    Integrate the scoring pipeline with Zendesk or Intercom. For Zendesk, use the webhook endpoint: when a new ticket is created, Zendesk sends a POST request to your integration server. Your server calls the Claude API, receives the score, and updates the ticket’s tags and group assignment via the /api/v2/tickets/{id}.json endpoint. For Intercom, use the conversation.created webhook and the update_conversation API. Handle rate limits: if Zendesk returns a 429 status, implement exponential backoff (1s, 2s, 4s, 8s). Set a confidence threshold: if the score is above 0.85, auto-route the ticket; if below 0.60, flag it for human review; between 0.60 and 0.85, route it but add a “low confidence” tag. This human-in-the-loop design satisfies GDPR Article 22.

  • Cutting First-Response Time in UK Logistics: A 4-Week AI Ticket Triage Pilot

    The Problem: Slow First-Response Time in UK Logistics Support

    You run a 500-to-2,000-person logistics or supply chain operation in the UK. Your customer support team handles 800 to 3,000 tickets per week across email, web forms, and a helpdesk portal. First-response time sits at 4 to 12 hours, and 30 to 50 percent of tickets are misrouted to the wrong team, forcing manual reassignment. You have run isolated AI pilots before — perhaps a document extraction proof-of-concept or a chatbot experiment — but none have moved into production. Your ISO 27001 certification requires that any new system touching customer data passes a formal risk assessment, and your operations team needs a measured before/after baseline on cycle time and error rate before approving rollout. The goal is not to replace your support staff but to cut first-response time by 30 to 50 percent within four weeks, using predictive scoring to route tickets to the correct team before a human ever opens them.

    Prerequisites Before You Start

    Before you write a single line of integration code, confirm these items are in place:

    • Process map: A documented flow of how tickets currently move from intake to resolution, including which teams handle which categories (delivery delays, billing disputes, customs queries, returns).
    • API credentials: Read/write access to your helpdesk (Zendesk, Freshdesk, Jira Service Management) and CRM via their REST APIs. You will need webhook endpoints for real-time ticket events.
    • ISO 27001 owner: A named compliance lead who can sign off on the risk assessment for using OpenAI API with customer data. This person must be involved from Day 1, not after the pilot is built.
    • Pilot budget: £1,500 to £4,000 for OpenAI API costs over four weeks, plus £8,000 to £15,000 for fixed-scope integration work. Confirm this with finance before Week 1 starts.
    • Operations lead: One person with 5 to 10 hours per week to review model outputs, approve routing rules, and flag misrouted tickets during the pilot.
    • Data samples: 200 to 500 historical tickets with metadata (sender, category, resolution time, team assigned) to train and validate the scoring model.

    Step-by-Step: Build the Pilot in Four Weeks

    Step 1: Run the process audit and capture baselines. Map every ticket category, the team that handles it, and the average time from intake to first response. Export 200 to 500 historical tickets from your helpdesk with fields: ticket_id, sender_email, subject, body, assigned_team, first_response_time_hours, resolution_time_hours, category. Store this in a CSV or database table. This is your before-state. Without it, you cannot prove the pilot worked.

    Step 2: Define routing categories and scoring thresholds. List 5 to 8 ticket categories your support team actually uses (e.g., delivery_delay, billing_dispute, customs_query, return_request, account_issue). For each, define what a correct routing looks like. Set a confidence threshold: tickets scoring 0.85 or above are auto-routed; below 0.85 go to a human queue. Document this in a one-page routing spec that your ISO 27001 owner signs off.

    Step 3: Build the OpenAI API integration via REST and webhooks. Create a webhook listener in your helpdesk that fires on ticket.created. The listener sends the ticket body and metadata to a lightweight service (Node.js or Python) that calls the OpenAI API using the gpt-4o-mini model. The prompt instructs the model to return a JSON object: {"category": "delivery_delay", "confidence": 0.92, "suggested_team": "dispatch"}. Log every API call with timestamp, ticket ID, and response in your SIEM to satisfy ISO 27001 Annex A.12.3.1.

    Step-by-Step: Run the Pilot and Hand Over

    Step 4: Implement human-in-the-loop approval. Any ticket with a confidence score below 0.85, or any ticket mentioning financial amounts, health data, or contract terms, is flagged for human review. Build a simple approval screen in your helpdesk or a lightweight web app where the operations lead sees the AI’s suggested routing, can accept or override it, and logs the reason for any override. This is not optional under ISO 27001 — you must demonstrate that a human controls decisions touching money or regulated data.

    Step 5: Run the pilot on live tickets for two weeks. Enable the webhook on 100 to 200 live tickets per day. The AI scores and routes; the operations lead reviews every ticket for the first three days, then samples 20 percent after that. Track daily: first-response time, routing accuracy (correct team vs. AI suggestion), override rate, and API cost. If the override rate exceeds 15 percent in any 7-day window, pause the pilot and recalibrate the prompt or scoring thresholds.

    Step 6: Measure before/after and document findings. In Week 4, compare the pilot metrics against your Week 1 baselines. You should see first-response time drop by 30 to 50 percent and routing accuracy at 85 percent or above. Write a two-page report: what worked, what failed, API costs, and a recommendation for rollout. This report is your input to the ISO 27001 management review and your business case for scaling to additional teams or channels.

    Step 7: Hand over to managed operations. If the pilot meets targets, transition to a managed operations model. Forfis continues to monitor model performance, tune routing thresholds monthly, update prompts as new ticket patterns emerge, and handle API cost management. You retain ownership of the data and the integration; Forfis operates the AI layer under a service-level agreement with defined accuracy and latency targets.

    Common Pitfalls and How to Detect Them

    • Overfitting on historical patterns: The model learns routing rules from last year’s ticket mix, but your operations have changed (new routes, new clients, new service levels). Detect this by tracking the override rate weekly. If it climbs above 15 percent, the model is misrouting. Recalibrate by retraining on the last 30 days of tickets, not the full historical set.

    • Skipping the human-in-the-loop step for high-value tickets: You auto-route a billing dispute because the confidence score is 0.87, but the ticket involves a £50,000 claim. This is an ISO 27001 breach. Detect this by auditing the approval log monthly. Any ticket with a financial amount above your defined threshold (e.g., £1,000) must have a human approval record.

    • Ignoring API cost creep: GPT-4o-mini costs roughly £0.15 per 1,000 input tokens and £0.60 per 1,000 output tokens. A 500-word ticket with a 200-word response costs about £0.05. At 2,000 tickets per week, that is £500 per week. If you do not set a monthly API budget cap in your OpenAI dashboard, costs can double if ticket volume spikes during peak season. Detect this by reviewing API spend weekly against your pilot budget.

    • Not logging API calls for ISO 27001 audit: If you do not log every OpenAI API call with timestamp, ticket ID, and response, you cannot demonstrate compliance during an ISO 27001 surveillance audit. Detect this by running a monthly audit of your SIEM logs. If any ticket ID is missing from the log, the integration is not compliant.

    What Comes After the Pilot

    The pilot is not the end state. Once you have a measured before/after baseline and a signed-off ISO 27001 risk assessment, the next logical step is to extend the triage layer to additional channels — voice, chat, or email — and to add document extraction for attached invoices, customs forms, or proof-of-delivery images. The same predictive scoring architecture applies: the model classifies the document type, extracts key fields, and routes the data to your ERP or accounting system. The human-in-the-loop control remains for anything touching money or regulated data. Your four-week pilot gives you the data, the compliance sign-off, and the operational muscle to justify that next phase to your board or investors. The integration is already built; the next step is scaling it.

  • Four-Week AI Pilot: Automating Order-Status Data Entry in a UK Medtech Firm

    The Problem: Manual Order-Status Data Entry in a Regulated UK Medtech Firm

    A 51-200 person UK medtech company handling order and shipment status updates for customer support is drowning in manual data entry. Every time a customer emails or calls about an order, an operator opens the CRM, searches for the order reference, checks the logistics provider’s tracking page, types the status back into the ticket, and logs the interaction. At 12-18 minutes per request and 3-5 percent transcription error rate, this single workflow consumes 15-25 percent of the support team’s capacity. The problem is not the volume alone; it is that the data is unstructured (email bodies, PDF attachments, voice notes) and the regulatory environment (ISO 27001, UK GDPR) means you cannot simply pipe customer emails into a third-party API without a documented risk assessment. The pilot targets this one process, automates the extraction and classification, and ships with a measured before/after baseline that proves the case for rollout.

    Prerequisites Before Week 1

    Before the dedicated AI team begins the four-week pilot, you need the following in place:

    • One named process owner from the customer support team who can answer questions about the current workflow and approve the pilot scope.
    • Access to historical documents: at least 200-500 examples of customer emails, PDFs, or spreadsheets containing order and shipment status requests, exported from Google Workspace or the CRM.
    • CRM API credentials with read/write permissions for the order and ticket objects, scoped to the pilot’s data set.
    • Google Workspace API access: Gmail API and Google Drive API scopes for the pilot mailbox, with data residency set to the UK or EU region.
    • A GPU server or cloud instance with at least 80 GB of VRAM (e.g., an A100 or H100) for running the open-weight model on-premise, or a confirmed decision to use a cloud GPU for the pilot phase only.
    • ISO 27001 documentation access: the client’s current statement of applicability and any existing risk assessments covering customer data handling, so the pilot’s controls align with the existing certification scope.

    Step 1: Run the Process Audit and Capture the Baseline

    The dedicated AI team maps every manual step in the order-status workflow and captures the baseline metrics. You export 200-500 historical requests from Google Workspace and the CRM, and the team tags each one with cycle time (from email receipt to ticket closure), error type (wrong order reference, missed shipment detail, incorrect status), and number of human touches. The output is a one-page scorecard: for a typical UK medtech firm, the baseline shows 14 minutes average cycle time, 4.2 percent error rate, and 3.1 human touches per request. This scorecard becomes the denominator for the before/after report and the justification for the pilot’s scope. The team also identifies which fields in the extracted data touch money, health data, or contracts, because those fields will require human-in-the-loop approval in the next step.

    Step 2: Select and Fine-Tune the Open-Weight Model On-Premise

    The team selects an open-weight model that fits the client’s GPU and data constraints. For a UK medtech firm where patient identifiers and order details cannot leave the building, the default is Llama 3 70B or Mistral 8x7B running on the client’s on-premise A100 server. The model is fine-tuned on the 200-500 historical documents from Step 1, using a supervised fine-tuning (SFT) dataset where each example pairs the raw email or PDF with the correctly extracted fields (order reference, shipment ID, status, date, customer name). The fine-tuning runs for 2-3 epochs on the client’s GPU, taking 4-8 hours. The team evaluates the fine-tuned model on a held-out set of 50 documents, targeting a field-level accuracy of 95 percent or higher before moving to integration. If accuracy falls below 95 percent, the team iterates on the SFT dataset or switches to a larger model variant.

    Step 3: Build the Google Workspace and CRM Integration

    The pipeline connects to Google Workspace through the Gmail API and Google Drive API. Incoming emails to the pilot mailbox trigger a push notification; the pipeline fetches the message body and any attached PDFs or spreadsheets, passes them to the on-premise inference endpoint, and receives structured JSON output containing the extracted fields. The pipeline then calls the CRM’s REST API to look up the order by reference, pulls the current shipment status from the logistics provider’s API (DHL, DPD, or the 3PL system), and merges the two data sets. The output is a draft customer-facing update and a structured record for the CRM. All API calls are logged with timestamps, request IDs, and data classification tags, feeding directly into the client’s ISO 27001 audit trail. The integration uses the client’s existing service accounts, not new credentials, to minimize the attack surface.

    Step 4: Configure the Human-in-the-Loop Approval Gate

    The approval interface is a simple web dashboard where the support operator sees a diff view: the source document on the left, the model’s extracted fields on the right, and a highlight on any field classified as touching money, health data, or a contract. The operator can approve, edit, or reject each field. In practice, 70-85 percent of routine order-status updates pass without human intervention because the model’s confidence score exceeds the threshold (typically 0.92) and no sensitive fields are present. The remaining 15-30 percent route to the approval queue with a 4-hour SLA. The queue is monitored by the process owner, and any rejection is logged with a reason code that feeds back into the SFT dataset for the next model iteration. This loop ensures the model improves with each week of live operation.

    Step 5: Run the Pilot and Produce the Before/After Report

    The pilot runs on a controlled sample of 50-100 live requests over two weeks. The measurement harness captures the same metrics as the baseline: cycle time, error rate, and human touches per request. The team compares the pilot results against the Step 1 scorecard and produces a before/after report. A typical result for a UK medtech firm is a 65 percent reduction in cycle time (from 14 minutes to 5 minutes) and a 50 percent drop in transcription errors (from 4.2 percent to 2.1 percent). The report also documents the ISO 27001 controls in place: on-premise data residency, access controls on the inference server, audit logging, and the human-in-the-loop gate for sensitive fields. This report becomes the business case for rollout to additional workflows, such as invoice processing or document extraction for clinical trial records.